Skip to main content

CMS Technical Reference Architecture

The CMS Technical Reference Architecture (TRA) provides CMS-approved technical standards and architecture guidance for designing, integrating, modernizing, and maintaining information systems across all CMS processing environments, ensuring consistency, interoperab

Last Reviewed: 10/1/2026

Contact: TRA Team | EnterpriseArchitecture@cms.hhs.gov

FOUNDATION 

Introduction to the TRA 

The CMS Technical Reference Architecture (hereafter referred to as the “CMS TRA”) articulates the technical architecture of all Centers for Medicare & Medicaid Services (CMS) processing environments (hereafter simply the “CMS Processing Environments”). A CMS Processing Environment is defined as: 

Any computing environment (e.g., CMS data center, virtual computing environment, or cloud computing including Infrastructure as a Service (IaaS), Software as a Service (SaaS), and Platform as a Service (PaaS)) that creates, consumes, and/or stores CMS-related data. CMS data includes sensitive, non-sensitive, and security information and event management-related information used to provide CMS services to the public and internal CMS users. 

The CMS TRA represents the Agency’s policy guidance to all Agency business partners wishing to develop, transition, and maintain information systems that interact with the CMS Processing Environments. The CMS TRA is approved and authorized by the CMS Chief Information Officer (CIO) and Chief Enterprise Architect (CEA). 

Adherence to the CMS TRA supports the Agency’s healthcare mission by providing: 

  • A secure CMS Processing Environment that protects sensitive information, including CMS, beneficiary, provider, and partner information 
  • Appropriate disaster recovery and business continuity capabilities 
  • Timely and economic transition of CMS applications into new processing environments 
  • An enterprise computing solution that responds to CMS’s evolving mission and business needs 

Toward these ends, the CMS TRA defines a common set of terms and definitions to support CMS’s architecture approach and ensure an effective operating environment. It conveys required design considerations, including security policies and controls (for confidentiality, integrity, and availability), reusability, scalability, and sustainability. The common framework of the CMS TRA supports future application designs as well as architecting and engineering CMS applications. It encourages the use and creation of enterprise shared services and presents clear guidance on their appropriate use and interaction among CMS data centers. By promoting a technical reference standard for future CMS task orders and acquisitions, the CMS TRA clarifies the decision-making process and target technical environments, which helps Agency contractors develop sound and acceptable transition approaches. 

Purpose 

This topic introduces the CMS TRA and is the keystone of CMS TRA guidance, as shown in Essential Guidance for the Entire CMS TRA. It articulates the CMS Architectural Vision and provides guidance relevant to all stakeholders. It establishes the guiding principles of the CMS architecture that informs all guidance in subsequent topics and presents overviews of CMS Services Framework and Multi-Zone Architectures. 

TRA Overview Diagram (page 7) 

 

Essential Guidance for the Entire CMS TRA 

 

This topic also summarizes how to request a change to the CMS TRA through the Architecture Change Request (ACR) process. 

Stakeholders interested in specific architectural topics can find detailed guidance organized in topics as follows: 

  • Network Services – focuses on network infrastructure including security-related appliances that monitor network traffic 
  • Infrastructure Services – focuses on physical or virtual infrastructure supporting CMS applications 
  • Application Development – focuses on application-level concepts and methodology 
  • Data Management – focuses on data management and data lakes 

The guidance in each CMS TRA topic includes narrative providing context; Business Rules (BR), which are requirements for TRA compliance; and Recommended Practices (RP), which are strongly encouraged for use within the CMS Processing Environments but not required. 

Intended Audience 

The CMS TRA is intended to guide system architects, business owners, system maintainers, and security auditors in working with CMS’s technical environment. A publicly available version is located at CMS TRA. CMS restricts access to the complete CMS TRA to the following authorized users: 

  • CMS staff 
  • CMS Processing Environment contractors 
  • Operator of the CMS Alliance to Modernize Healthcare Federally Funded Research and Development Center (the Health FFRDC) 

CMS executive or management approval is required to provide this document to other entities pursuant to business need. 

Complementary Documents 

The CMS TRA complements CMS standards documentation. The CMS TRA supersedes and takes precedence over other existing CMS standards documentation, with the following exceptions: 

  • CMS Information Security (IS) Acceptable Risk Safeguards (ARS, hereafter simply the “CMS ARS”) 
  • All volumes of the CMS Risk Management Handbook (RMH) 
  • CMS Information System Security and Privacy Policy (IS2P2) 
  • CMS Section 508 Policy 

The CMS ARS is located at: CMS Acceptable Risk Safeguards (ARS) on the CMS CyberGeek website. 

CMS provides an orientation briefing for CMS TRA Document Development (MS PowerPoint) and the ACR Form (PDF) for initiating an Architecture Change Request. 

A mailing list entitled “CIO Resource Library Communications” is available to notify subscribers when new or revised information technology (IT)-related policies, technical or TRA standards, directives, or guidelines are available. To subscribe to the new list, please select the following URL and enter your email address: 

https://public.govdelivery.com/accounts/USCMS/subscriber/new?topic_id=USCMS_12066 

Please direct all questions, comments, suggestions, or requests for further information to the CMS CIO Policy Officer at IT_Policy@cms.hhs.gov.  

Guiding Principles 

The CMS TRA informs the development of the Agency’s Common Enterprise Infrastructure (CEI) and is the foundation for the design and implementation of processing components and computing environments across the CMS Enterprise. Ensuring the secure operation of all CMS Processing Environments (as defined in CMS Processing Environments) is a primary architectural consideration. Adhering to the CMS TRA improves interoperability of services across the enterprise, supporting the goal of providing high-value services at lower cost, with higher availability, reduced implementation times, and improved maintainability. 

The CMS TRA describes the technical baseline for all CMS Processing Environments. The CMS TRA cites proven, effective, and tested common controls that meet legislatively mandated mission, security, and privacy requirements. The Agency updates the CMS TRA as required to respond to CMS’s changing needs and direction and to reflect current technical standards. 

As CMS’s repository of technical reference standards, the CMS TRA plays a critical role in all contracts. Contracting Officers refer to the CMS TRA topics in task orders and solicitations to ensure that contractors understand and apply technical consistency across data centers and contract efforts. 

The TRA focuses on services, outcomes, and effects instead of physical structures, specific products, or system components. Unless explicitly stated, references to specific products and/or vendors within the CMS TRA provide information about the existing technical environment rather than express an Agency preference. The CMS TRA supports the Department of Health and Human Services (HHS) Enterprise Architecture principles and satisfies the CMS ARS security control, CM-2, Baseline Configuration. 

CMS relies on the following guiding principles for the CMS TRA: 

  • Defense-in-Depth 
  • Least privilege 
  • Vendor agnostic, amenable to new technologies 
  • Reuse 
  • Service-Oriented Architecture (SOA) and Microservices 
  • Enterprise Services 
  • Common Platform Services 
  • Cloud First 
  • Use of Open Source Software 
  • Automation 
  • Sustainability 

Defense-in-Depth 

Defense-in-Depth employs multiple, coordinated security countermeasures to protect against security issues throughout the infrastructure. 

A Services Framework architecture separates services by their function and encapsulates them with their own layers of protection/defense providing adaptable layers of depth. 

A multi-zone architecture, which is a type of Services Framework architecture, places multiple layers of protection/defense on each zone by segregating networks and controlling inter-zone communications. 

Least Privilege 

The CMS TRA adheres to the security concept of least privilege, the security objective of granting users only those accesses they need to perform their official duties. When a user is permitted to “self-elevate” temporarily (e.g., Linux sudo command), this action is logged. 

Vendor Agnostic 

Unless explicitly stated, the CMS TRA does not advocate a preferred vendor or suite of products. One notable exception to this rule is shared services (including security services), where a specific technology has been selected and is mandated for use of or integration with this shared service. CMS encourages innovation of supportable solutions within the parameters of the CMS TRA. 

When selecting software, business owners should evaluate competing product licenses and leverage existing agency purchasing sources before approving new contractor purchases. The focus should be on minimizing costs and promoting consistency across CMS applications and environments. Business owners are required to utilize Inter-Agency Agreements (IAA), CMS Enterprise License Agreements (ELA), and Government Furnished Software (GFS) to the greatest extent possible. These actions will support CIO Directive 25-02 Use of CMS Enterprise License Agreements, issued July 1, 2025. 

CMS maintains enterprise licensing information at: Enterprise Software Licensing 

For more information on IAA, ELA, or GFS agreements, please contact the Software Asset Management Program Management Office (SAM PMO) at SoftwareLicensesCMS@cms.hhs.gov. 

Reuse 

CMS encourages reuse and sharing wherever practicable to avoid creating redundant services/applications. There can be multiple cloud and physical data centers in the CMS environment. Any application can reside in any combination of CMS data centers (including CMS Clouds) and use software services—for example, Identity Management (IDM) and Enterprise Portal—of the other data centers. Application software in one data center may leverage services in another data center. 

The benefits of reuse include: 

  • Reduced implementation times 
  • Lower project risk 
  • Reduced costs 
  • Minimized redundancy 
  • Improved security 
  • Improved interoperability 

CMS employs several strategies for reuse: adoption of a Service-Oriented Architecture, use of Microservices, and the implementation and use of Strategic and Preferred Solutions. 

Service-Oriented Architecture and Microservices 

SOA defines a set of principles for developing software as interoperable, reusable, distributed services to fulfill business functions. For example, the Medicaid Information Technology Architecture (MITA) is a SOA framework that allows CMS and state Medicaid organizations to create, deploy, execute, and manage reusable, modular, and interoperable business, data, and technology services. The Application Development section addresses SOA extensively. 

Microservices (MS) are a modern interpretation of SOAs for building distributed software systems. Each component/ module of an application is developed and deployed separately. This contrasts with a traditional, “monolithic” application in which all components are developed and deployed as one piece. Microservices are well suited to DevOps methods and tools. The Application Development section addresses Microservices, with additional detail in the TRB Research Spotlight, “Microservice Architecture,” published April 11, 2023. 

Strategic and Preferred Solutions 

CMS strongly encourages the use of Strategic and Preferred Solutions. These are solutions which have been designed, tested, and approved for broad use within CMS. Utilizing these existing solutions, which include formally designated Enterprise Shared Services (ESS), shortens the Time-to-Delivery (TTD), facilitates efficient infrastructure use, reduces redundant capabilities, improves access, and leads to more consistent cybersecurity posture. The Office of Management and Budget (OMB) Memorandum M-19-16, Centralized Mission Support Capabilities for the Federal Government, April 26 2019, defines federal IT shared services and policy. (Also see Solutions to Improve Agency Management Efficiency (ussm.gsa.gov) and cio.gov/policies-and-priorities/shared-services) 

Details and guidance for the use of these services are described in the following section, CMS Strategic Guidance and Preferred Solutions. 

Common Platform Services 

Common platform services are technology services provided in support of all hosted applications. The CMS ecosystem now supports platform services across both cloud and on-premises environments. These capabilities are designed to standardize and simplify support and management. 

Common platform services may include (but are not limited to): 

  • Network capabilities such as load balancing, domain name services (DNS), dynamic host control protocol (DHCP) 
  • Event logging (Syslog, SIEM) 
  • Database Administration 
  • Operating System (OS) Administration 

Application developers should use enterprise shared services and common platform services where possible rather than implementing specialized or single-use software. 

Cloud First 

The CMS TRA supports the Federal Cloud Computing Strategy’s (OMB/Federal CIO Kundra, February 14, 2011) Cloud First Policy. For details, please refer to the Network Services and Infrastructure Services sections. 

Pursuant to this objective, CMS is closing its on-premises data centers and is migrating all on-premise applications to cloud environments. Any residual on-premises functionality is migrating to cloud-connected collocation facilities. This strategy is aligned with OMB M-19-19 regarding the Data Center Optimization Initiative. “Agencies may not budget any funds or resources toward initiating a new agency-owned data center or significantly expanding an existing agency-owned data center without approval from OMB.” See the full OMB M-19-19 directive for more information. The objective is to drive meaningful improvement to IT infrastructure, achieve cost savings through optimizations and closures, and foster IT modernization. 

CMS has a robust set of recommended cloud services and capabilities available to enable rapid migration to and deployment within the cloud. See the CMS Strategic Guidance and Preferred Solutions section for further detail. 

Use of Supportable Open Source Software 

The CMS TRA permits the use of supportable Open Source software in CMS Processing Environments. For details, please refer to the chapter on Open Source Software in the Application Services Section. Unsupported software in any capacity cannot be patched in a timely manner; accordingly, any Open Source software without a dedicated support matrix violates the CMS ARS. 

Development Methodology Agnostic 

The CMS TRA applies to all CMS Processing Environments regardless of the approach used to develop and maintain CMS Federal Information Security Modernization Act (FISMA) applications (e.g., using Waterfall, Prototyping, Incremental, Spiral, Rapid Application Development, and Agile). 

Automation 

The CMS TRA encourages the appropriate use of automation to reduce variability and increase speed. In code development and deployment process, continuous integration automation (code checkout/build, static code analysis, dependency check, unit testing), continuous deployment release orchestration, Infrastructure as code can be used to avoid manual errors and manage complexity. In production, operational monitoring can be automated by vulnerability scanning, system health, alerts, etc. 

Sustainability 

The CMS TRA encourages projects to consider sustainability as a critical element of architectural solutions. Projects should apply appropriate attention to future scale, cost, security, operational efficiency, technology, and vendor support. The long-term viability of applications and data products must be considered when integrating any solution into the CMS Processing Environments. Risk and mitigation planning for sustainability should address personnel (technical expertise), operational processes (automation), adherence to Service Level Agreements, and total cost of ownership. 

In cloud environment the sustainability for cost, scalability, service level agreements can be achieved by adoption of cloud native features like dynamic scaling up/down of resources based on usage, use of shared instances vs reserved instance and optimal sizing of resources. 

Accessibility and Section 508 Compliance 

Inaccessible technology interferes with an individual’s ability to obtain and use information quickly and easily. Section 508 of the Rehabilitation Act of 1973 was enacted to eliminate barriers in IT, to make available new opportunities for people with disabilities, and to encourage development of technologies that will help achieve these goals. CMS operates many programs that are critically important to people with disabilities. IT is often the vehicle for delivering these programs. The most effective and least costly approach to technology accessibility is for CMS to incorporate accessibility into its IT governance, development, and procurement processes. 

Under Section 508, federal agencies must give federal employees and members of the public with disabilities access to Information Communication Technology (ICT) and information that is comparable to the access available to individuals without disabilities. For CMS Processing Environments and applications, adherence to the Agency’s Section 508 policies include: 

  • External and internal user interfaces 
  • Hardware and software products, including Commercial Off-the-Shelf (COTS) and custom products 
  • Development and maintenance tools 
  • External and internal reports or other artifacts created by CMS applications or tools 
  • Electronic documentation 
  • Support services 

CMS Guidance 

CMS provides resources for Section 508 compliance, including policies, validation procedures, and a list of CMS Section 508 Clearance Officers, at: Section 508 and CMS. 

PREFERRED: Additional CMS guidance and resources are found in Application Development, Web-based UI Services, as well as in the CMS Design System. 

The CMS Section 508 Policy provides a clear rationale and compliance road map for CMS Processing Environments and applications and takes precedence over the CMS TRA. The following paragraphs summarize the basis for these requirements, which is explained further in the CMS Section 508 Policy. 

In 1986, Congress added Section 508 to the Rehabilitation Act of 1973. Section 508 established non-binding guidelines for IT accessibility. On August 7, 1998, the president signed into law the Workforce Investment Act of 1998, which included amendments to the Rehabilitation Act. These amendments significantly expanded and strengthened the IT accessibility requirements in Section 508 and made them binding on federal agencies. 

Section 508 applies to Information Communication Technology, previously referred to as Electronic and Information Technology (EIT). Section 508, as amended, specifically requires that, when federal agencies develop, procure, maintain, or use ICT: 

  1. Individuals with disabilities who are federal employees shall have access to and use of information and data that is comparable to the access to and use of the information and data by federal employees who are not individuals with disabilities; and 
  2. Individuals with disabilities who are members of the public seeking information or services from a federal department or agency shall have access to and use of information and data that is comparable to the access to and use of the information and data by such members of the public who are not individuals with disabilities (FAR 39.201 and 36 CFR 1194.1). 

Section 501 of the Act requires CMS to be a model employer of people with disabilities. Section 504 requires CMS to ensure equally effective access to its programs and services. CMS can reach both objectives only if it adheres to the requirements of Section 508. 

Relevant Section 508 Standards for CMS IT 

The CMS Section 508 policy is based on the standards described in the following subtopics. 

U.S. Access Board 508 Standards for Information Communication Technology 

The U.S. Access Board’s Information and Communication Technology Revised 508 Standards and 255 Guidelines apply to a wide range of products and services, including but not limited to, the following: 

  • Computers 
  • Telecommunications equipment 
  • Multifunction office machines (e.g., copiers/printers, software, websites, information kiosks, and transaction machines) 
  • Electronic documents 

The final rule jointly updates requirements for information and communications technology covered by Section 508 of the Rehabilitation Act and Section 255 of the Telecommunications Act. The Section 508 Standards apply to electronic and information technology procured by the federal government. The Section 255 Guidelines address access to telecommunications products and services and apply to manufacturers of telecommunication equipment. 

Web Content Accessibility Guidelines 2.1 

Web Content Accessibility Guidelines (WCAG) 2.1 define how to make web content more accessible to people with disabilities. WCAG 2.0 is developed through the World Wide Web Consortium (W3C) process in cooperation with individuals and organizations around the world. The goal is to provide a shared standard for web content accessibility that meets the needs of individuals, organizations, and governments internationally. WCAG 2.0 is designed to apply broadly to different web technologies now and in the future, and to be testable with a combination of automated testing and human evaluation. WCAG 2.0 has three conformance levels, including A (lowest), AA, and AAA (highest). CMS requires that the WCAG 2.0 standards be met at both the Level A and Level AA conformance level. 

Web accessibility depends on accessible content as well as accessible web browsers and other user agents. Authoring tools also have an important role in web accessibility. For an overview of how these components of web development and interaction work together, please refer to the following links: 

PDF / UA 

The Portable Document Format/Universal Accessibility (PDF/UA-2) standard, defined by ISO 14289-2:2024, describes technical requirements for universally accessible PDF documents by identifying a set of relevant PDF functions (including text content, images, form fields, comments, bookmarks, and metadata) based on ISO 32000-2:2020 (PDF 2.0) and specifies how they should be used in PDF/UA-compliant documents. Successful access to content within PDFs depends both on compliant documents and compliant PDF programs and assistive technology. PDF/UA therefore also specifies requirements for compliant assistive technology. 

For assistive technologies to work properly with PDF/UA, they must meet the following requirements: 

  • They must be able to recognize all structural elements, attributes, and key values used in the specification and output them for the user of a PDF document. 
  • They must allow the user to navigate through the document by page number, through the structure tree, or by using bookmarks. 
  • They must allow the user to easily set and change the magnification of a PDF document at any time. 

Document Accessibility at CMS and HHS Section 508 Compliance Checklist 

For information about designing accessible documents that adhere to HHS’s requirements, please refer to the HHS Section 508 Accessibility Conformance Checklists. 

HHS Guidance 

HHS Web Standards apply to all HHS Office of the Secretary (OS) websites—including all Operating Divisions (OPDIVs) / Staff Divisions (STAFFDIVs) and Secretarial Priority websites, whether aimed at internal or external audiences. These standards are required for the design and development of all HHS/OS websites and may be adopted by OPDIVs. 

The Framework covers all HHS web sites, internal or external, that are owned, managed, or funded by Operating and Staff Divisions, including CMS, whether developed by staff or acquired through contracts, cooperative agreements, grants, and/or formally established partnerships with other government entities and/or the private sector. The Framework covers traditional web content, including all attached or linked HHS files, and all news and social media, including videos, podcasts, blogs, Wikis, and associated applications on all HHS web sites. 

CMS Testing for 508 Compliance 

CMS uses the following tools and methods to test IT products and artifacts for Section 508 compliance: 

  • Manual code review 
  • Review of system development test plans 
  • JAWS screen reading software 
  • Dragon translation software speech-to-text 
  • Zoom Text screen magnification software 

     

CMS Strategic Guidance and Preferred Solutions 

Introduction 

The architectural guidelines described by the CMS TRA are designed to provide flexibility to development teams in choosing a technical approach, and yet there are cases where CMS has a strongly preferred approach or solution option. This section provides information about such Strategic Guidance and Preferred Solutions. 

CMS systems and data must be protected from a constantly evolving threat and vulnerability landscape. Aligning to strategic guidance and utilizing preferred solutions strengthens CMS security posture through extending the use of thoroughly validated and tested security controls and procedures. Maintaining CMS data within the CMS security boundary facilitates CMS stewardship of the data. 

The use of preferred solutions also shortens time-to-delivery, reduces redundant development, facilitates efficient use of infrastructure, and leverages CMS strategic investments and economies of scale. This aligns to the TRA Guiding Principle of Reuse as well as broader federal IT policy (see Office of Management and Budget (OMB) Memorandum M-19-16, Centralized Mission Support Capabilities for the Federal Government, April 26 2019, and Federal Shared Services). 

The goal is to speed secure system deployment by utilizing existing solutions that may provide built-in automation, approved security configurations, and pre-authorized infrastructure. Use of these solutions is strongly encouraged unless there is a more compelling business case for an alternative. 

CMS encourages its organizations to maximize the benefits of utilizing the Microsoft 365 Government Community Cloud (GCC) capabilities through the CMS tenant. All CMS email accounts include G5 subscriptions that provide access to Office 365 applications, storage, and other advanced security and data management features. CMS has completed transition from Zoom to Microsoft Teams for all CMS-hosted meetings and is migrating from the on-premises SharePoint 2019 environment to SharePoint Online. 

The CMS TRA will provide high-level information about preferred solutions. Detailed information including their setup and use will be found through the supplementary sources/ links to additional information. 

Note that hyperlinks in this section may refer to CMS internal information sources; a CMS EUA ID and access to the CMS intranet may be required for access. 

Additionally, links to Strategic Guidance and Preferred Services will be included throughout the TRA, aligned to associated TRA guidance. 

PREFERRED - CMS strategic guidance and preferred solution information throughout the CMS TRA will appear like this. 

Enterprise Services 

The preferred solutions in this section are CMS “enterprise services”, based on an expanded definition of the term. The updated CMS definition of enterprise services reflects that these capabilities may be developed and provided by any CMS Center or Office or by a vetted partnership of Centers/Offices. This includes but extends well beyond centrally provided “shared services”. In the new model, enterprise services can be leveraged as a shared capability or provide a pattern for a new implementation if required. 

A CMS enterprise service may be leveraged without onboarding with an established support organization or a central enterprise instance. CMS encourages the use of the most appropriate deployment model depending on business requirements. A project might, for example, choose to deploy a parallel instance but still use the standards and processes of the enterprise instance, potentially enabling them to utilize existing support staff. An existing, validated solution can be used as an architectural pattern, with any refinements shared back with the original. 

The overall group of enterprise services form a federated set of capabilities which development teams can tap into, avoiding duplication and aligning with CMS standards. 

Data Stewardship and Governance 

Sound data stewardship practices are essential to the protection of CMS data. The risk of data compromise is exacerbated when CMS data is moved or copied outside of the CMS security boundary. With the shift to flexible cloud computing environments, CMS now has the capability and capacity to provide for very large and complex data storage and analytics requirements. Keeping CMS data and associated data processing and analysis within the CMS security boundary enables CMS to maintain provenance and proper stewardship of CMS data. 

As such, CMS policy is that all CMS data remain within CMS authorization boundaries, except for public data released by CMS. Business requirements to do otherwise will be reviewed on an exception basis to ensure appropriate controls (which may include a Data Use Agreement) are in place. A new TRA section, Data Sharing and Governance provides additional information on this topic. It also introduces Business Rule BR-DG-1, which reinforces the imperative to keep CMS data inside CMS boundaries, while Recommended Process RP-DG-2 suggests some ways to accomplish this. 

CMS Cloud is the preferred environment for CMS data storage and analytics. 

Preferred Solutions 

CMS Preferred Solutions are presented here in the following categories: 

  • CMS Hybrid Cloud 
  • CMS Enterprise Data 
  • CMS DevSecOps Support 
  • CMS Collaboration Capabilities 

Note that the list of Preferred Solutions will evolve over time. Additional sections and solutions will be added in future TRA releases. 

CMS Hybrid Cloud and CACHE 

The CMS TRA supports the Federal Cloud Computing Strategy’s (OMB/Federal CIO Kundra, February 14, 2011) Cloud First Policy as well as the Federal Cloud Computing Strategy’s (OMB/Federal CIO Kent, June 24, 2019), Cloud Smart Policy. 

CMS Hybrid Cloud is the strongly preferred hosting platform for all CMS developed applications. It provides the most operationally integrated and cost-effective platform solutions. In addition to “self-service” cloud computing with guardrails, CMS Cloud features managed Infrastructure-as-a-Service (IaaS) environments for both Amazon Web Services (AWS) and Microsoft Azure Government (MAG), as well as for CMS on-premises data centers. 

The CMS TRA chapters on Cloud Infrastructure and Virtualization, among others, address CMS Hybrid Cloud services. 

CACHE and DRaaS-CACHE 

The Continuously Available CMS Hosting Environment (CACHE) service offerings provide the primary on-premise infrastructure hosting platform, with seamless integration with the overall CMS hybrid enterprise environment. These include: 

  • Geo-diverse data center and hosting within the DRaaS-CACHE FISMA boundary 
  • Full Infrastructure-as-a-Service (Iaas) virtualized environment 
  • Application Workload hosting across multiple computing platforms, including Mainframe, Open Systems (x86), and hybrid cloud, as well as multiple storage platforms, including disk, tape, and file-level 
  • Comprehensive Disaster Recovery as a Service (DRaaS) with next-gen replication capabilities across multiple technology platforms 
  • Enterprise Services that include robust Application and Network Load Balancing, Domain Name Service (DNS), Directory Services (LDAP), and Time Services (NTP) 

Critical applications that support CMS programs and operations should, if managed by CMS, rely on DRaaS-CACHE to protect both their system availability and their data. 

CMS Enterprise Data 

CMS maintains several preferred solutions for accessing and managing CMS program data: 

  • The Integrated Data Repository (IDR) Cloud is a high-volume cloud data platform integrating Medicare claims with beneficiary and provider data sources, as well as such ancillary data. This robust, integrated data supports mission-essential analytics in CMS and in other government agencies. It also helps to prevent the proliferation of replicated data sources that perpetuate data inefficiency, duplication, inconsistency, inferior quality, and increased costs for associated infrastructure. 
  • The IDR Enterprise Data Product (EDP) now supports the data mesh functions of the previous Enterprise Data Mesh (EDM), which was decommissioned in 2024. Based on the IDR’s Snowflake implementation, the EDP maintains a central point where systems and individual users can find, locate, access, and use CMS enterprise information with “data in place.” This enables data owners and curators to focus on data and data quality while also enabling consumers to bring their own preferred compute resources, analytics, and APIs. Additionally, the EDP enables a wide spectrum of programs and consumers to leverage program data sets with close to zero provisioning time with their choice of tools and technologies optimal for their use case. 
  • CMCS DataConnect is an all-in-one cloud analytics platform for the Center for Medicaid & CHIP Services (CMCS) data. DataConnect provides a range of capabilities to streamline Medicaid and CHIP program analysis, monitoring, and oversight. It includes dashboards, tools, and datasets to enable CMCS staff, researchers, and other data experts to produce meaningful insights from complex data. 
  • CMMI Analysis and Management System (AMS) holds key information about CMMI models and demonstrations. AMS helps stakeholders access information about, and collected by CMMI models. 
  • CMS Master Data Management (MDM) has provided a single point of access to a singular, synchronized, comprehensive and ID-resolved authoritative source of Provider and other data for use by CMS and other external organizations. MDM provided support across multiple CMS business units with a focus on eliminating redundancy, inconsistency, and fragmentation of CMS data. IDR, IDR Enterprise Data Product, and CMCS DataConnect have all relied on this authoritative data. 

MDM is scheduled to be retired in February 2027. Support of this data management functionality is migrating to the CMMI Analysis and Management System (AMS), Integrated Data Repository (IDR) Cloud, IDR Enterprise Data Product (EDP), and CMCS DataConnect. ID Resolution will be performed as needed within these systems, as well as by the Center for Clinical Standards & Quality (CCSQ) and the Center for Program Integrity (CPI). 

  • CMS Research Data Assistance Center (ResDAC) provides technical assistance to researchers interested in CMS Medicare and Medicaid data. Various data sets are available; however, access is restricted to approved research requests. Public use data files (which contain no protected information) are available via data.CMS.gov 

CMS DevSecOps Support 

A key goal of the CMS TRA is to support sound and secure software development practices. The DevSecOps tooling landscape is vast, and it can be challenging to determine which solutions are the best match for project requirements. CMS has preferred platform and tool options available which can help development teams manage and secure development pipelines. 

Enterprise Services are preferred over other project specific services as they offer potential operational and/or cost benefits. Existing systems are encouraged to utilize enterprise services wherever possible. 

  • CMS DevOps Tools 
    • GitHub is a common repository platform, used to distribute open-source software and other content. Thus, it is both a DevOps and a collaboration tool. Premium Enterprise GitHub provides advanced management and security features. CMS has implemented GitHub Copilot, an AI coding assistant (see TRA recommended practices for AI-Assisted Coding). 
      • SaaS GitHub instances include OC Tools-supported CMSgov and theCMCS | MACBIS organization that collaborates with states and other Medicaid and CHIP partners. 
      • CMS Hybrid Cloud supports SaaS-based OIT Enterprise GitHub, which provides enterprise-grade collaboration, security, and administration. 
      • OC GitHub Enterprise is supported by the Web Help Service Desk. It is primarily for support of public-facing medicare.gov, healthcare.gov, cms.gov, and their related applications and ADOs. During the first half of 2026, the OC Web and Emerging Technologies Group (WETG) is migrating these repositories from this internal GitHub to CMSGov. 
      • CCSQ manages an additional GitHub Enterprise instance. 
      • The Open Source Program Office provides an additional GitHub repository containing additional tools, templates, and other information about using and share open source software. 
    • CloudBees CI is a continuous integration, deployment, and delivery (CI/CD) server solution 
    • JFrog Platform is comprised of Artifactory, a binary repository manager and XRay, an add-on to the Artifactory product used to enhance the security posture 
  • CMS Testing 
    CMS provides Testing as a Service (TaaS), a CMS Cloud offering that provides a suite of products support application teams with test case management, functional and regression test execution, and performance test execution. This includes: 
  • CMS Security Scanning and Analysis 
    These CMS Cloud inspection and analysis tools support application security and vulnerability management for both developers and the security team: 
    • SonarQube, which provides static code analysis. Static code analysis attempts to highlight possible vulnerabilities within ‘static’ (non-running) source code by using analytic techniques. 
    • Snyk (“sneak”), a SaaS-based system composition analysis (SCA) tool that enables applications to be developed and built securely. SCA identifies the open-source software in the codebase. This analysis is performed to evaluate security, license compliance, and code quality. It proactively finds and fixes vulnerabilities in codes, open-source dependencies, container images and Infrastructure as Code (IaC) configurations, and offers context, prioritization, and remediation. 

CMS Collaboration Capabilities 

A number of enterprise collaboration tools are available within the CMS environment to facilitate data sharing and team communications, supported by OIT and the Office of Communications (OC). These include: 

  • Jira and Confluence: 
    • Atlassian Jira is an application lifecycle management solution that helps teams plan, manage, and report on their work. It is used for bug tracking, issue tracking and project management. 
    • Atlassian Confluence is a team workspace providing teams a place to create, capture, and collaborate on project information including text, tables, images, and other content. 

The CMS Cloud Agile Tools Team supports Enterprise Confluence and Enterprise Jira, with integrated TestRail. The Office of Communications provides OC Jira and OC Confluence, primarily for support of public-facing medicare.gov, healthcare.gov, cms.gov and their related applications. 

  • SharePoint is an application platform that allows organizations to store and organize any content and information. That includes documents, images, videos, news, links, lists of data, web pages, and tasks. The CMS Enterprise SharePoint Support (ESS) team supports the SharePoint Online (SPO) environment. 
    • CMS SharePoint Online (SPO) is the new home of CMS-wide enterprise information. In addition to new sites being created for all CMS Centers and Offices, the entire contents of the decommissioned On-Premise SharePoint have been migrated to cmsgovonline.sharepoint.com. These include: 
      • share.cms.gov 
      • capms.cms.gov 
      • cmsintranet.share.cms.gov 
      • mysite.share.cms.gov 
      • pbrs.cms.gov 
    • SharePoint Online is part of the CMS Microsoft 365 SaaS enclave. 
  • Slack is a workplace messaging tool through which CMS employees and contractors send messages and files, defining and using specialized and general channels that contain message threads. Some chat and meeting functions will migrate to Microsoft Teams. 
  • Microsoft Teams is also part of the CMS Microsoft 365 Government subscription. Microsoft Teams is a collaboration tool designed to enhance productivity by integrating with other Microsoft products and providing a centralized space for communication, file sharing, chat, and meetings. Preliminary guidance is available. Transition to Teams from Zoom Workplace took place during 2025. Partial transition from Slack is also planned. 
  • Another feature of Microsoft 365 is OneDrive. OneDrive provides backing cloud storage for M365 tools such as Word, Excel, PowerPoint, Teams, and SharePoint. Note that CMS Enterprise SharePoint Online is managed separately from SharePoint sites for Teams and Channels. CMS is migrating CMS Box users to OneDrive, primarily for secure sharing within CMS. Box remains available, primarily for external sharing. 

Please also see CISO Memorandum 25-01: Updated Best Practices and Guidance for the Use of Approved CMS Collaboration Tools. 

   

Artificial Intelligence Guidance 

Introduction 

Artificial Intelligence (AI) tools and technology are rapidly evolving with the promise of enabling increased efficiency, insight, and analysis across a broad range of use cases. CMS is embracing AI to improve health care administration and delivery. At the same time, policies, processes, and mechanisms are being established to ensure that AI-enabled systems at CMS are developed, deployed, and used responsibly and ethically, mitigating risks and protecting sensitive information while also maximizing benefits for CMS and beneficiaries. 

Aligned to federal guidelines (see ai.cms.gov), this section provides current AI-related policies and recommended practices at CMS. Given the dynamic nature of the AI domain, frequent updates and changes are expected. It is not intended as a primer on AI technologies or approaches. Additional information on those topics can be found within the CMS AI Playbook and Technical Application Resources. 

More policy guidance can be found in the HHS Intranet: “Reminder of Existing HHS IT User Policies Relevant for Third-Party Generative AI Tools,” “HHS Policy for Securing Artificial Intelligence (AI) Technology,” and “HHS Policy for Rules of Behavior for Use of Information and IT Resources.” 

Additional Considerations 

Accountability and Oversight 

All individuals and teams employing AI or Machine Learning (ML) at CMS are responsible for the system's output, regardless of the model or tools used. The “owner(s)” of this output are also required to use the appropriate mandated security controls for “sensitive” and especially “privacy data” assets. Continuous human oversight is required to ensure CMS and federal guidelines are met. 

Risk Management and Security 

AI Model Supply Chain and Provenance 

In the modern information security landscape, AI-driven supply chain attacks are a concern. These can introduce malicious code through compromised AI models, data, or code repositories. Teams deploying AI models should ensure the provenance and integrity of AI components used in production. 

Example: The “Model Namespace Reuse” attack, demonstrated against Google and Microsoft products, highlights the risk of trusting a model based on its name alone. This form of attack can enable a threat actor to gain code execution permissions and gain access to underlying infrastructure. 

Recommendation: Implement System Composition Analysis to verify the origin and integrity of AI dependencies. 

Embracing a Zero-Trust Architecture for AI 

Using AI technology that can perform outbound internet requests can introduce risk to CMS environments, including sensitive code and data that could be transmitted outside the network. 

Recommendation: Teams should perform threat-modeling and risk assessments to ensure proper network segmentation and limit an AI tool’s access to only the necessary resources to contain the blast radius of any potential breach. Implement Data Minimization principles to ensure AI tools access only the information strictly necessary for their function. See The TRA section Zero Trust Architecture. 

Monitoring, Versioning and Observability for AI and Production Operations 

System maintainers must implement observability to understand how their AI systems are behaving over time. This requirement supports debugging, security, and continuous improvement, and ensures alignment with OMB M-25-21 and M-25-22. 

Recommendation: Track traces, evaluations (EVALs), prompt management or versioning, and key metrics for production systems using AI. 

Recommendation: Prompts — AI system prompts used for production at CMS systems should be versioned (See BR-SCM-1). Store prompts (or reusable “recipes”) and review them and measure their performance over time. 

Data Handling, Data Retention and Privacy 

Privacy-Preserving Techniques 

When developing and testing AI models that handle sensitive data, teams should adopt privacy-preserving techniques like federated learning and homomorphic encryption. These methods allow for model training and inference on encrypted or decentralized data without direct sensitive data exposure. 

Use of Synthetic Data for AI Initiatives 

The use of synthetic data is highly recommended to avoid the use of sensitive Personally Identifiable Information (PII) or Protected Health Information (PHI) in development and lower environments entirely. This mitigates the risk of exposure and simplifies compliance. 

Examples: AI techniques like Generative AI, GANs and VAEs, and initiatives like SyntheticMass and Synthea™ [can also be used to avoid the use of sensitive PII or PHI from development and lower environments entirely at CMS. 

Review Data Rights, Retention Policies prior to the use of External AI Tools and Services 

The use of AI tools, services or infrastructure external to CMS may introduce risks, including companies that train or fine-tune AI models using CMS non-public code, data or metadata. 

Recommendation: Application teams, CMS system owner or business owners introducing new AI technology into their systems should ensure compliance with CMS security and privacy requirements, and that appropriate data use agreements are in place before incorporating external AI systems that can access non-public CMS data . Teams should be especially careful with Terms of Use and Agreements that enable the training, share or selling of CMS data. For inquiries, consult with the CMS Privacy Office. 

AI Services and Data Retention: Content generated by your AI system may be subject to Data Retention policies and FOIA. For inquiries, consult with Records_Retention@cms.hhs.gov. . 

Overall AI Business Rules 

BR-AI-1: AI Tools and Services Must Meet Federal AI, Cybersecurity and Privacy Standards in the Handling of Sensitive Data 

Sensitive data, including Protected Health Information (PHI), Sensitive Personally Identifiable Information (SPII), classified, export controlled, trade secret and other confidential information may only be used with AI tools and services that meet HHS and CMS Cybersecurity Standards. 

To ensure compliance and protect privacy, consult with the CMS Privacy Office in relation to AI tools and Services that could handle sensitive data. 

Rationale: 

All HHS and CMS policies regarding AI use, Personally Identifiable Information (PII) protection, and data security (including storage, transmission, and sharing) must be adhered to in order to ensure proper implementation of AI risk management practices per the HHS AI Strategy and HHS Compliance Plan for OMB Memorandum M-25-21. 

Also see in the TRA: BR-F-5: Any System That Processes CMS Data Must Be Covered by a CMS ATO; BR-SQ-6: De-Identification of Production Data Is Required in Non-Production Environments; NIST SP 800-122, “Guide to Protecting the Confidentiality of Personally Identifiable Information (PII);”NIST SP 800-18, ”De-Identifying Government Datasets: Techniques and Governance.” 

BR-AI-2: High-Impact AI Use Cases Must Meet Minimum Risk Management Practices 

A high-impact use case, as defined in the Office of Management and Budget (OMB) Memorandum M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, is one where the AI’s output serves as a principal basis for decisions or actions that have a legal, material, binding, or significant effect on specific critical areas. 

A CMS AI Governance risk assessment will determine if an AI use case is high-impact. All CMS high-impact use cases must apply minimum risk management practices in accordance with M-25-21. 

Contact CMS IT_Governance for an AI Governance Risk Assessment. 

Rationale: 

Non-approved AI tools may retain or use submitted information in ways that could expose CMS sensitive data. Attackers may be able to steal sensitive data by manipulating prompts or exploiting vulnerabilities in AI systems. Non-approved AI tools may not be compliant with evolving federal regulations regarding data protection, records retention, consumer rights, and algorithmic transparency. Additionally, these tools may lack the security controls and oversight mechanisms required for handling federal government data within established security boundaries. 

BR-AI-3: Foreign Entity AI Tools May Only Be Used if Deployed on CMS Infrastructure 

With the rapid introduction of new AI capabilities into the market, it can be difficult to determine which are aligned to federal government requirements. AI technology from foreign entities can be used if it’s deployed on CMS infrastructure and does not send data to the internet. This aligns with the principle of CMS data staying within the United States (see BR-SAAS-8: CMS Data Must Always Reside in the U.S.). Such deployments must utilize normal CMS and AI governance processes, including the CMS Risk Management Framework and the NIST AI Risk Management Framework. 

Rationale: 

The use of unapproved or foreign AI tools and models could compromise CMS information. Foreign AI models, especially those from countries with broad government access to data, could be used for surveillance and the collection of sensitive personal information. Third-party AI models may lack sufficient security controls, potentially exposing the systems to data breaches. 

BR-AI-4: Human Review Must Follow Use of AI Tools to Write CMS Policies 

Individual CMS employees are accountable for official CMS policies. AI tools are only to be used in support of and not as a substitute for human decisions and oversight of CMS strategic or compliance-related activities. Humans must be in the loop. 

Rationale: 

AI tools lack the capacity for human judgment and critical thinking, which is crucial for interpreting complex situations, considering ethical implications, and crafting policies that fully align with CMS’ values and objectives. AI tools may generate false or misleading information as factual and may reflect biases present in their training data. 

BR-AI-5: Do Not Rely on AI for Final Decisions for “High Impact” Cases 

AI should provide advice or recommendations; final decisions must be made by qualified staff with documented oversight. AI use cases for these purposes are considered potentially “High-Impact AI,” and actions they support may risk compliance with privacy and civil rights requirements. 

Rationale: 

OMB memorandum M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust defines “High-Impact AI” as “AI with an output that serves as a principal basis for decisions or actions with legal, material, binding, or significant effect on: 

  1. “an individual or entity’s civil rights, civil liberties, or privacy; or 
  2. “an individual or entity’s access to education, housing, insurance, credit, employment, and other programs; 
  3. “an individual or entity’s access to critical government resources or services; 
  4. “human health and safety; 
  5. “critical infrastructure or public safety; or 
  6. “strategic assets or resources, including high-value property and information marked as sensitive or classified by the Federal Government.” 

AI tools lack the capacity for human judgment and the nuanced perspective necessary to make complicated decisions. High-Impact AI risks need to be managed. 

BR-AI-6: AI-Supported Official Actions Are Subject to Records Retention Requirements 

When AI provides work products in support of official CMS actions that would be subject to records retention, those work products become part of the record and must be retained according to the applicable records retention schedule. Those AI-created work products must include the AI model version and the prompt(s) used. 

Rationale: 

“Adequate and proper documentation” (per 44 U.S.C. Chapter 31 §3101) supporting official CMS actions are part of the action. This supports transparency of AI use. For more information, consult with the OSORA Issuances, Records & Information Systems Group. 

AI-Assisted Coding Guidance 

Threat Modeling for AI Systems 

See TRA Recommended Practice AD-SS-8: Perform Threat Modeling During the Design Phase to Identify Potential System Threats. Additional recommended practices regarding performing Threat Modeling for AI systems can be found in A Practical Understanding of Threat Modeling for AI Systems 

AI Recommended Practices 

This section provides Recommended Practices information for AI overall and for specific AI concepts. 

RP-AI-1: Release and Maintain AI Code as Shareable Open Source Software 

CMS projects must prioritize sharing AI code, models, and data government-wide, consistent with the Open, Public, Electronic and Necessary (OPEN) Government Data Act. Custom-developed AI code is also covered by the SHARE-IT Act and OMB M-25-21. More information is found in the TRA Application Development section on Open Source Software. 

AI Prompt Crafting 

Prompt Crafting is the practice of writing clear and helpful directions that a large language model (LLM) can use to generate more accurate, relevant, and useful outputs for a given task or question. This is like providing a person the right level of contextual detail and specific instructions to complete a task in the manner it is to be completed. For the best results using LLMs, it is important to build instructions and relevant details in a similar manner- by providing the task to accomplish, the role that the AI is taking on to complete it, any relevant reference materials as context, and the desired response format. 

Prompt Crafting is the bridge between human intent and LLM capability. It is the heart of what determines the quality and usefulness of AI-generated responses from LLMs. In the context of CMS Chat, effective prompting enables users to interact with documents, synthesize information across sources, and generate new content effectively. 

For CMS, effective Prompt Crafting matters because it: 

  • Improves the accuracy and reliability of AI-generated outputs 
  • Increases consistency in AI interactions across the organization 
  • Reduces the likelihood of AI hallucinations or incorrect responses 
  • Enhances the user experience of staff using CMS Chat 
  • Increases the efficiency of the AI and reduces time and compute costs 

The following best practices can be applied to effectively use chatbots like CMS Chat to enhance the user experience and ensure more accurate and useful AI interactions 

RP-AI-2: Clearly define the context of the prompt 

Be specific about which aspects the AI is to focus on and what type of analysis is needed. Provide relevant background information, specify limitations, include domain-specific terminology, and define the scope of the response 

Example: “Within the context of the 2024 Medicare Physician Fee Schedule final rule, focusing specifically on telehealth provisions, analyze the requirements for audio-only services.” 

RP-AI-3: Clearly define the role the AI should adopt 

This approach provides clear context for the AI’s responses, helps maintain consistent tone and expertise level, and enables more targeted and relevant outputs. 

Example: Instead of “Tell me about Medicare Part B coverage,” use “Within the context of the 2024 Medicare Physician Fee Schedule final rule, focusing specifically on telehealth provisions, analyze the requirements for audio-only services.” 

It is also possible to provide roles that the AI can use to tailor the generated output. 

Example: “Act as an experienced Medicare benefits counselor with 15 years of experience explaining coverage to beneficiaries. Explain Medicare Part B coverage in simple terms.” 

RP-AI-4: Break down complex requests into clear, sequential steps 

This approach helps AI systems provide more organized and concise responses. When analyzing documents or synthesizing information, structure prompts to guide the AI through the process. 

Example: “First, analyze the key points of this policy document. Then, identify any changes from the previous version. Finally, summarize the potential operational impacts for the policy team.” 

Sometimes, it’s best to break these steps into separate prompts for each step. AI have a maximum amount of response length they can provide, so allowing them to focus on one step at a time, with a full response at each stage, may provide better end results. This is also true when you have large documents. The AI can be asked to review the first chunk of text, then the next, and so on until the needed details are derived from the entire document. 

Example: 

  • Prompt 1: “Analyze the key points of this policy document.” 
  • Prompt 2: “Review the key points provided and identify any changes from the previous version.” 
  • Prompt 3: “Based on the output in the previous response, summarize the potential operational impacts for the policy team.” 

RP-AI-5: Provide Guidelines 

Outline any scoping, rules, writing styles, or specify any specific format requirements for the response, such as using headers, bullet points, or specific citation formats. Clearly communicate how the information should be presented. 

Example: “Present your findings in a structured format using headers to break up the text, followed by a few sentences that explain the header, and then bullet points to breakdown the main points.” 

Sometimes, more complex guidelines and response formats are necessary. It’s possible to instruct many AIs to respond in structured formats while also demanding it utilize a specific approach, style, or scope. 

Example: “Present your findings in a structured table format. I’d like columns A, B, C, and in A, please create categories for the text, followed by a brief description of each in B, then bullet points to breakdown the main points in C. Write in the style of a CMS Policy Expert but be concise and clear. Avoid jargon or complex topics. Instead, focus on clarity and simplicity.” 

RP-AI-6: Create New Chats or Sessions When Switching Topics 

It is best to create a new chat or discussion session when switching topics. The context of conversations in generative AI prompting sessions is typically stored and available throughout a single chat. Unless the context of the entire chat is needed for the next query, it is recommended to create a new conversation, so the AI does not hallucinate or conflate topics. For example, if the AI was asked in one chat to clean up a project presentation, then instead of asking how to write a policy document in the same chat, start a new one with the right level of context provided to ensure the AI has no distractions from the desired intent. 

RP-AI-7: Use Structured Prompts 

The best prompts are structured with all the details covered in the previous best practices in mind. A recommended structure template to get the best results is: 

[Task] + [Role] + [Guidelines] + [Reference Context (Reference Materials)] 

Each are further defined below: 

  • [Task] – The action(s) you wish the AI to take based on the [Role], [Guidelines], [Format Requirements], and [Reference Context (Reference Materials)] 
  • [Role] – The role that the AI is to take on to complete the [Task] 
  • [Guidelines] – Rules or guidance for the AI to ensure it accomplishes the task as expected. This may include the expected output structure/format or instructions on how to write (perspective, style, etc.), and other guidance for the AI 
  • [Reference Context (Reference Materials)] – Relevant attachments, details, documentation, resource materials, etc. that can act as the source of truth for the AI to accomplish its [Task] 

Example: 

Task: “Please create a categorized list of Medicare Part B rules relevant to a retiree from Tennessee.” 

Role: Act as an experienced Medicare benefits counselor with 15 years of experience explaining coverage to beneficiaries. Explain Medicare Part B coverage in simple terms. You complete TASKS based on the GUIDELINES and REFERENCE CONTEXT and respond as outlined by the FORMAT REQUIREMENTS 

Guidelines: Please respond in a bulleted or tabular format. Responses should be concise, but fully explained in a style most suitable to the intended audience. Note if any relevant details may be missed and what additional actions should be taken to ensure the [Task] can be completed. 

Reference Context: *Attached Document* “OR [Copy/Paste text from the appropriate documentation]” 

RP-AI-8: Use Iterative Refinement 

Prompt Crafting is not a one-time process but rather an iterative approach that involves starting with a basic prompt, evaluating the response, refining the prompt based on the output, and then testing and validating any improvements. It is also a creative process — one that benefits from continuous experimentation and adjustment to ensure the AI’s responses align with evolving requirements and user needs. By reviewing each new output, identifying gaps or inaccuracies, and incorporating feedback into the next version of the prompt, steady improvements will be achieved in both the clarity and effectiveness of AI interactions. 

RP-AI-9: Recommended AI-Assisted Development Methodologies 

Recommended Workflow Steps: 

  1. Specify: Create a formal, human-readable specification that defines the goals in accordance with best-practices. 
  2. Plan: Use AI to generate a technical plan based on the spec and enterprise standards. 
  3. Implement: Use AI to break down the plan into small, testable tasks and generate code. 
  4. Review and Verify: A developer reviews each code change for accuracy, security, and quality. Automated testing and security checks run in CI/CD. 
  5. Iterate: If issues are found, refine the specification or prompts and re-run the cycle.  

 

CMS Processing Environments 

A CMS Processing Environment is defined as any computing environment that creates, consumes, and/or stores CMS data, as defined in the CMS Data and CMS Sensitive Information definitions. The CMS TRA applies to all CMS Processing Environments that CMS controls. This includes all production environments having a CMS Authorization To Operate (ATO) and any development or testing environments residing in data centers that host CMS workloads. 

The CMS TRA refers to production environments as ATO(ed) environments. CMS TRA guidance for production environments also applies to any non-production environment with an active CMS ATO, as such environments must be managed and secured like a production environment to maintain its ATO. 

By convention, there are three types of processing environments: Development Environments, Validation Environments (sometimes known as Implementation), and Production Environments (also known as Operational environments). These environments provide a set of high-level stages corresponding to a software promotion model, which follows the baseline progression as follows: 

  • Development environments support multiple business application baselines, including the execution of functional testing, unit testing, application integration testing, and regression testing. When there is a significant defect in a baseline in one of the other environments, the baseline returns to the Development environment for repairs and other corrective actions. 
  • Validation environments provide a staging area for production releases and evaluation of the final release versions of production code, database, and related packages for business applications. Many non-functional requirements (security, privacy, performance, and various “ilities” such as scalability and reliability) are tested here. CMS intends the Implementation environment for performance and stress testing, system and user acceptance testing, final integration testing, and security assessment. The Validation environment should be comparable to the production environment to promote CMS business owner confidence in the performance and functionality of the system. 
  • Production environments provide stability and security for the final release versions of production code, database, and related packages. All production environments are required to have a CMS ATO. 

Projects may select different naming conventions for their lower environments. CMS recommends, however, using Development and Validation because of their frequent use in CMS IT. Projects may choose to create additional lower environments as required to facilitate testing (e.g. Test environment; Integration environment). Projects may choose to perform activities such as Independent Validation and Verification (IV&V) or User Acceptance Testing (UAT) in whichever environment makes the most sense for the project. 

Each environment should emulate the production environment; accordingly, the infrastructure must mirror production in versions, patch levels, configurations, and network. Development and testing services may be provided, such as performance testing, functional compatibility testing, and performance monitoring. 

Business application code and database changes, including break-fix, must be tested successfully in a development or test environment before promotion to the next higher environment. Related changes to infrastructure services (such as firewalls and switches) must be tested in a development or test environment before deployment to the implementation environment. 

Non-ATO(ed) environments may not contain sensitive information. As described in Business Rule (BR) BR-F-5, CMS prescribes the removal or replacement of personal identifiers, other personalizing information, and other CMS sensitive information to render information anonymous before it may be loaded into an environment not covered under an ATO. Redaction of sensitive data must occur in an environment covered under an ATO. Existence of Personally Identifiable Information (PII), Protected Health Information (PHI), or other CMS sensitive information in any environment causes that environment to require production-level security, privacy controls, and coverage by ATO, regardless of its name. 

Development Environments 

A development environment supports business application baselines during code and database development, and for development testing functions. The data center may support multiple instances of a development environment to accommodate development of multiple variants of a business application. 

Validation Environments 

The validation environment provides a secure, locked-down staging area for production baselines and for evaluating the final release versions of the production code, database, and related packages for the business applications. It is used to perform implementation testing and to complete business application testing before release into the production environment. This includes validating system performance, events, alerts, and notifications. CMS may provide support from this environment for end-user training on system usage and functionality as well as for access to training data. 

There may be multiple instances of a validation environment to support the testing of a business application (or multiple versions) and its preparation for promotion to the production environment. The validation environment must maintain an image of the last successfully tested version of a business application code and database. 

The validation database should be approximately the same size as the production database (for realistic performance testing) and, if it has been issued an ATO, may contain sensitive information. If it contains sensitive information, the implementation database must be protected to the same standards as the production database (please refer to BR-F-5.) 

All changes to the validation environment must be subject to CMS change management. Additional requirements and constraints include, but are not limited to: 

  • Scheduling of operations in the implementation environment must include adequate time to roll back any changes, for example, changes to business application code and databases, and any infrastructure changes. 
  • Troubleshooting of business application code and database errors will be limited to Root Cause Analysis (RCA) only. 
  • Business application code and database changes, including break-fix changes, must be tested successfully before their promotion from the implementation environment to the production environment. 
  • Changes related to infrastructure services must be tested successfully in the implementation environment before promotion to the production environment. 

CMS ATO(ed) and Production Environments 

CMS production environments are required to be ATO(ed) environments. CMS TRA guidance for production environments also applies to any non-production environment with an active CMS ATO, because such an environment must be managed and secured like a production environment to maintain its ATO. 

An ATO(ed) environment provides a stable, secure, locked-down operational environment to complete the business application production baseline. This environment is inaccessible except to authorized operations and security personnel. Access to business applications will be provided via approved user interfaces only. Only the Administrator role will be granted server-level access to the CMS business applications in the ATO(ed) environment. (and must follow a Software Development Life Cycle (SDLC) driven promotion process). 

All changes to the ATO(ed) environment must be subject to CMS change management. Additional requirements and constraints include, but are not limited to: 

  • All production environments are required to have a CMS ATO. 
  • All business application code and database changes, including break-fix, must be tested successfully in the implementation environment before promotion into a CMS production environment. 
  • All related changes to infrastructure services must be tested successfully in the implementation environment before deployment to the CMS production environment. 
  • Production-readiness testing is required for all initial deployment of business applications. Troubleshooting of business application errors will be limited to RCA only. 
  • Monitoring & Reliability Testing is required for all initial deployment of business applications. New business application code, databases, and infrastructure migrated into the ATO(ed) environment will become part of the production baseline. 

The quality of CMS’s production environments reflects on the organization’s reputation and public trust. Although the production environment may be used for quality assurance and testing, the CMS TRA generally discourages such use, suggesting instead that the validation environment be used for such tasks. This is true if the code or data tested could lead to incorrect or misleading results or behavior. The production environment may also be used for community verification of quality assurance of data. 

Non-CMS Processing Environments 

The CMS Processing Environments may interact with non-CMS processing environments. 

Third-Party Websites and Applications 

The HHS Office of the Chief Information Officer (OCIO) Policy for Social Media Technologies, March 07, 2012, Policy 2010-0003.1 includes managing the use of Third-Party Websites and Applications (TPWA) and establishes requirements for Department access to web-based technologies that are not exclusively operated or controlled by HHS. This policy is consistent with federal guidelines and regulations. 

HHS-OCIO directs that TPWAs must be addressed in a Privacy Impact Assessment (PIA) and may require a Security Impact Analysis (SIA). 

OMB Memorandum M-10-23, Guidance for Agency Use of Third-Party Websites and Applications, states: 

The term “third-party websites or applications” refers to web-based technologies that are not exclusively operated or controlled by a government entity, or web-based technologies that involve significant participation of a nongovernment entity. Often these technologies are located on a “.com” website or other location that is not part of an official government domain. However, third-party applications can also be embedded or incorporated on an agency’s official website. 

Social Media 

The use of social media at CMS is governed by the Guidelines for Secure Use of Social Media issued by the CIO Council, and by HHS-OCIO Policy for Social Media Technologies. Also, see HHS Social Media Policies.  

 

Services Framework Architecture 

Services Framework - Concept 

The CMS Services Framework provides a standardized template for implementing systems at CMS. This template or pattern details security and interoperability requirements. The goal of the framework is to maintain the integrity of CMS applications and data. 

CMS Services Framework - Principles 

CMS utilizes a Defense-in-Depth (DID) and a least privilege strategy to protect its assets. Defense-in-depth is a security strategy that uses a series of layered, redundant defensive measures to protect sensitive data, personally identifiable information (PII) and information technology assets. If one security control fails, the next security layer thwarts the potential cyber attack. Least privilege is the concept of restricting access rights of users to only those resources that are required for performing their legitimate functions. CMS’ services framework provides a structure for implementing security challenges to ensure CMS resources are protected. 

The goal is to protect CMS assets and reputation. CMS data must be protected from unauthorized access. To minimize the impact of any breach, the egress of data from CMS must also be protected. Too many security implementations concentrate on protecting from unauthorized access but do very little in protecting the egress of data if a breach was to occur. The implementor must address the complete security picture, prohibiting unauthorized access to the data, as well as minimizing the ability to obtain (download) data when a breach has occurred. Security implementations are always a trade-off between risk and cost. The implementor should work with their ISSO (Information System Security Officer) to validate appropriateness of security controls for any implementation. 

A key guidance to protect CMS data is to ensure that the overall architectural design is consistently putting CMS’ valuable data at least three independent legitimacy tests away from the open Internet. A legitimacy test is defined as the challenge, filtering or transformation of the data to obfuscate the technical details regarding the sensitive data. For more information regarding challenges, see the information on Mediation Principles. Within the CMS network, services may interface with any other available service, given that the proper authorization is in place. 

The following topics will define the CMS services framework from a services perspective to detail the zoned architecture and the security obligations required to protect CMS assets. The TRA Multi-Zone architecture will be depicted as a services framework defined by requirements and risks rather than network routers, and will consider the role of each service, its interfaces, its parent services, and its dependent services in supporting TRA compliance.  

 

Services Framework Architecture 

CMS Services Framework 

The architecture will be described by the set of services within the architecture, detailing the functions these services provide and the objectives for protecting these services. The following section of the TRA will discuss how these services can be applied. 

The services framework is centered around core data, application, and edge services, integrated with management and security services. See the Elements of the Services Framework diagram below. Management and security services both enable administrative access to the system. Security services assist in cybersecurity threat detection and protection, as well as the compliance with CMS security policies. The management and security services are housed in different subnets, accessible only through a virtual private network (VPN). 

Elements of the Services Framework  (page 7)

Core Services 

Core services include data, application, and edge services. These services are detailed below. 

Data Services 

Data services are fundamental and the most basic component protecting the data. They are the components of the architecture that reside in front of and provide access to the data. 

Services Framework Data Services (page 4) 

The data services must: 

  • Restrict access to the data 
  • Validate the caller (service consumer) has the appropriate access rights to the data services (producer) 
  • Obfuscate implementation details from the consumer, not exposing data formats, credentials, etc. to the consumer. 
  • CMS expects non-data services to front CMS data services. Data services are not directly accessible from outside a CMS data center. 

Application Services 

Application services are the components of the architecture that are generally responsible for housing the business logic. They are the consumer of data services and the producer of edge services. 

Services Framework Application Services (page 3) 

The application services must: 

  • Validate the caller (service consumer) has the appropriate access rights to the application services (producer) 
  • Provide access to data services 
  • Act as the mid layer between data and edge services, obfuscating the access details of the data from services which are accessing that data 

Edge Services 

Edge services interface with external components from non-CMS systems. 

Services Framework Edge Services (page 5)

The edge services: 

  • Must restrict access to the application services 
  • May, depending on architecture, provide authentication services. 

Mediation Principles 

Mediation principles are implemented as part of the data, application, and edge services to enhance the security by obfuscating the access details to the data. As the request proceeds through the services framework, the request must pass legitimacy tests and be filtered or transformed to provide the necessary obfuscation of critical system components and data. An inbound data request, in most use cases, passes through edge, application and data services. Therefore, at each level, validations and transformations are performed to protect the underlying service. Some systems need to support outbound connections, and these must also follow the path out through the edge services. As outbound calls provide a ‘bad actor’ with a pathway for extracting data from a CMS data center, this pathway is restricted out to only those specific components which are required to be traversed. The outbound calls to a specific destination are also limited. 

When implementing the data, application and edge services, the system designer should consider the following mediation principles: 

Challenges -– Using rules to allow authorized access or the routing of messages and network packets. 

Filtering -– Using criteria and rules to select data, permit a request, or check the validity of data, messages, or network packets. 

Transforming -– Changing a structured request, query, result, or response, into a different structure or syntax. Changing only the headers and routing envelopes of messages or network packets does NOT qualify as transforming. 

Management Services 

Management services assist the system administrators in maintaining the system. The tools required are placed in a management zone, accessible only through the VPN, which provides two factor authentication. 

The tools included within the management zone vary by application, but typically include the tools necessary for software releases and patching. 

Security Services 

The security zone houses any security tools for the application. CMS has enterprise logging and monitoring tools that are available to the system maintainers, so not all applications require a separate security zone. 

CMS Services Framework - Summary 

In summary, the CMS Services Framework enables applications to be orchestrated and/or composed of a service-fabric that integrates a set of services which enable access to and provide protection of CMS data. The architecture should implement a service as the authorized channel to access the application data. Users from external networks, CMSNet, or the CMS LAN must be authenticated to access the service fronting the data. Additionally, the file storage service must be configured to deny access to all unauthenticated users and all unauthorized users. Furthermore, users from external networks, CMSNet, or the CMS LAN must not have direct access to the CMS files persisted by the storage services.  

 

Services Framework Architecture 

CMS Multi-Zone Architecture 

Introduction 

The CMS TRA specifies a zone architecture that provides defense against security attacks and implements layers of challenges to ensure only authorized access to CMS resources. The services framework details the functions provided by each of the service types, and these services are represented by zones in the CMS multi-zone architecture. In general, the service names match the zone terminology (e.g., Data Services and Data Zone) except for Edge services which corresponds to the Presentation Zone. As the use of cloud platforms and cloud services (such as content delivery networks) have become more prevalent within CMS, the term “Edge Services” has been adopted to represent the services that front and protect CMS internet-facing resources. 

The multi-zone architecture is focused on protecting CMS assets via security challenges and enforcing defense-in-depth principles. Zones are a way of grouping and sharing resources based upon a shared security posture. Distinct zones generally exist in legacy data centers but are not as prevalent in cloud implementations where virtual resources, such as network, compute, and storage are usually project based implementations and there is tight integration with other cloud service provider (CSP) provided services. CMS does not require distinct zones for protecting assets, but it does require that data is protected by at least three security challenges. This is a minimum requirement, as additional security measures may be helpful in reducing the risk of a security breach. In addition, the security measures should not be housed on the same resource. This is done to ensure that if one resource is compromised, other security measures are in place to thwart an attack. 

The following subtopic will diagram and discuss the multi-zone architecture, detailing the relationship between the zones and the services framework. Note that the depiction is reminiscent of the CMS three-zone architecture as the three-zone model is one instance of the CMS multi-zone architecture. This is a result of the physical nature of the three-zone architecture and its implementation with CMS data centers. The services framework and multi-zone architecture is applicable to cloud environments, where the delineation of the zones is not as obvious. Cloud implementations that utilize a combination of CSP services and CMS services obscure specific zones implemented by the system developer. The specific zone is not as significant to the design as the need to implement challenges to protect CMS assets. 

After the discussion of the multi-zone architecture, examples of typical use cases are provided that demonstrate how the function could be implemented in the three-zone architecture as well as the cloud. These examples will support the reader's understanding of the architecture and security requirements. Please note, since system requirements and design vary greatly, these examples are to demonstrate the principle of protecting system assets, but the specific example may not apply to every system design. It must be stressed that the system developers should work with their ISSO in developing and implementing protections to CMS assets. 

Zone Architecture 

CMS defines a zone as “a portion of the network isolated by firewalls that serves a specific business function”. These firewalls may be physical or virtual, or implemented by other means (e.g. Security Groups in AWS, Network Security Groups in Microsoft Azure, etc). 

CMS data centers have typically utilized a three-zone architecture design. While CMS does not dictate the actual number of zones, in CMS data centers the common practice was the implementation of three zones based around data, application and edge services. As a result of the hierarchical nature of the implementation, using specific zones allowed for ‘like’ services to communicate (e.g., application to application service) in and between data centers as long as the CMS network was utilized. In addition, this implementation method also allowed for services to communicate with other services one layer below their own (e.g., edge services to application). Therefore, data services are not directly accessible from outside a CMS data center. Services on the same level are assumed to have passed through the same levels of security and may be accessed directly, although common practice is to restrict consumers of a service to just those components that require the use of the service. Systems implemented in the cloud do not generally have this distinct segmentation and assumptions about the security posture do not exist. 

In today’s more dynamic, elastic, and transient cloud environments, the multi-zone architecture is still relevant, but the security challenges may be implemented within Cloud Service Provider (CSP) services or within resources implemented by the CMS development organization. Clouds, microservices, software defined networks, and other emerging technologies present new paradigms for application architecture and often favor new design patterns for performance, security, and cost avoidance. Also, applications may resemble a fabric of services running on a dynamically changing cloud of virtual resources connected through networks often more defined by permissions than wires. These newer cloud environments and implementations rely upon CSP’s services which leads to a collaboration between the CSP and system implementor for implementing security. 

CMS Processing Environments are composed of one or more of the following zones to provide Defense-in-Depth, as depicted in the CMS TRA Multi-Zone Architecture figure below: 

  • Presentation Zones (PZ) – Edge services that support the presentation of content. Presentation Zones are accessible to external networks via firewalls through a Trusted Internet Connection (TIC). 
  • Application Zones (AZ) – Application services which support business logic for applications and creating dynamic user presentations. 
  • Data Zones (DZ) – Data services that contain data and data services used by applications. 
  • Management Zone – To support specialized services, such as Public Key Infrastructure (PKI), Domain Name System (DNS) services, and system management services. 
  • Security Zone – A shared mediation service to support security services. 

As the CMS TRA Multi-Zone Architecture diagram shows, the Presentation Zone controls the ingress and egress of all external communications into the CMS Processing Environment. The Application and Data Zones may communicate with corresponding Application Zones and Data Zones in other CMS data centers only. CMS has the capability of communicating from a zone in a CMS data center to a CMS cloud environment. Application and Data Zones would be allowed to communicate to equivalent zones in the cloud. These equivalent zones would be part of a CMS private subnet within the CMS enclave at the service provider. 

In the CMS vernacular, a Management or Security Zone is a network segment whose primary function is in support of security or infrastructure, and typically provides services to all the other zones. These Zones generally follow the same rules as other zones though there are a few additional rules. A data center may choose to group infrastructure functions into more zones than those listed here. Note: CMS TRA Multi-Zone Architecture represents the conceptual connections between zones and does not depict network implementation such as Trusted Internet Connections (TIC) integration and Virtual Routing and Forwarding (VRF), which are covered in more detail in TRA Network Services. 

CMS TRA Multi-Zone Architecture (page 10)

Transactions within a zone are permitted without restriction unless the traffic is firewalled between data centers. Transactions traversing the zones are controlled and protected via firewalls and other security mechanisms. This multi-zone architecture allows CMS to monitor and control business application transactions within and between zones. The Transport and Management Zones provide infrastructure and supervisory services to manage the core zones. 

Presentation Zone 

The Presentation Zone contains the front-end components of applications. The Presentation Zone receives requests from external sources, performs data validation, and proxies the requests to the Application Zone for processing. No business logic or database processing is performed on Presentation Zone servers. The Presentation Zone function is to proxy communication requests to the Application Zone (i.e., application-related connections must originate in the Presentation Zone). It is also the zone where business applications first receive data from external sources and is therefore the first zone to challenge requests for validation, authorization, and malicious content. 

For outbound communications, the Presentation Zone is the last zone traversed before reaching the Internet. 

Application Zone 

The Application Zone contains the business logic components of an application. It receives requests from the Presentation Zone and requests necessary data from Data Zone components. The Application Zone can host proxy services to the Data Zone. 

Data Zone 

The Data Zone contains all data sources, including data stores supporting directory services, authentication functions, data lakes and meshes as well as the applications’ operational databases. All interactions between the Application Zone and the Data Zone must use mediation and data access services. CMS permits the use of database-stored procedures that may reside on database servers in the Data Zone. 

Management Zone 

The Management Zone provides services to all zones and includes such functions as: 

  • DNS servers 
  • Backup servers 
  • Logs 
  • Remote Access Services 
  • System monitoring applications 
  • Asset/vulnerability management 
  • Application deployment functions 
  • System configuration management function 

The specific services provided by the Management Zone will vary from data center to data center. Projects are encouraged to discuss specifics with their data center provider. The Management Zone is the appropriate zone for hosting development and operations automation, such as DevOps deployment components. 

Security Zone 

The Security Zone provides services to all zones and includes such security functions as: 

  • Security monitoring and response 
  • Network and Host Intrusion Detection Systems (NIDS/HIDS) 
  • Intrusion Protection Service (IPS) 
  • Asset/vulnerability management 
  • Recording and monitoring of system, security, and audit logs 
  • Antivirus monitoring and analysis of server based anti-malware agent 
  • Enterprise-level security monitoring by integrating with the CMS Cybersecurity Integration Center (CCIC) 

Applying Zoned Architecture 

As described above, CMS implements defense-in-depth principles in a multi-zoned architecture. In the subtopic above, we talked about defense-in-depth and how each zone within the architecture contributes to protecting CMS resources. 

In the cloud, the definitive lines between zones are not as clearly defined, and utilizing services from the cloud providers further obscures the responsibility for implementing defense-in-depth. Firewalls may not be implemented, but the team may use security groups (in AWS environments) or their equivalent within their respective cloud for controlling access and flow through the system. CMS requires at least three challenges before allowing access to sensitive data. In the cloud environment, these challenges may not be as clearly separated into zones as CMS data centers would support. The reader should take ‘three’ challenges as a minimum to implement security needs. As such, the implementor should strive to deliver the most secure implementation, limiting risk, that resources can support. 

We will detail some common implementations within the cloud environment below and discuss how defense-in-depth can be implemented. Remember these examples are to demonstrate possible implementations and the implementor should work with their ISSO to validate the security adequacy of the design. Note that the current practice is for CMS cloud engineers to provide the implementor with a private and public subnet. The implementor creates the zones or the defense-in- depth using this model. For example, the application zone and data zone could be implemented within the private subnet, controlling flow and access using security groups. 

Example 1 – API Implementation

In this example, the implementor desires to implement an API to access CMS sensitive data for users external to CMS. The API endpoint should be protected by a Web Application Firewall (WAF) which may implement one or more of the security challenges and mediation principles (e.g. Geo fencing). A typical implementation would include: 

  1. Multi-factor authentication prior to accessing the APIs that front sensitive data 
  2. Authorization that the user has adequate permissions/roles to access the API. This could be accomplished using a directory. 
  3. The API endpoint should obfuscate (see Mediation Principles) the data access details to avoid providing the end user with any type of information regarding the type of database, location, etc. 

The diagrams below depict the zonal mapping of the API example to a CMS data center implementation and a cloud model where the implementation is supported with cloud services. Amazon Web Services (AWS) is used in the example, but the example holds for any CMS approved cloud. Please note that in the cloud example, we do not explicitly name the presentation and application zones. In this example, the implementation utilizes services provided by the cloud service provider (CSP). While not explicitly named, the reader can infer the web application firewall is analogous to the presentation zone and the API gateway to the application zone. 

Example 1 - API Implementation (page 1)

Additional security challenges/features, such as data masking, may be applied to further protect CMS sensitive data. The implementation should also prohibit the egress of data using the API or the resources that house the API. 

Example 2 – Business Intelligence Tool Implementation 

In this example, the implementor intends to utilize a cloud service that will use a virtual desktop/BI tool to access CMS data. In this example, due to the nature of the data, multi-factor authentication must be used. 

  1. A typical implementation would utilize multi-factor authentication, either through a CSP or CMS service to validate the user. 
  2. Once authenticated, the user’s authorization must be validated to ensure that access to the data is appropriate. This could be achieved through a sign-in or verification check within the database or CMS directory for roles/permissions. 
  3. The virtual desktop/BI tool should also be restricted to access only the targeted CMS data source, such as for example the CMS Enterprise Data Mesh (EDM). This could be accomplished by using security groups, or locking down the IP/ports that the cloud-based virtual desktop/BI tool is permitted to access within the CMS enclave. 

The diagrams below depict the zonal mapping of the Business Intelligence Tool example to a CMS data center implementation and in the cloud where the implementation is supported with cloud services. AWS is used in the example, but the example holds for any CMS approved cloud. 

Example 2 - Business Intelligence Tool Implementation (page 2)

Additional security challenges/features may be implemented depending on the risk posture of the system. For example, data masking can provide additional security as well as configuring the desktop/tool to limit the ability of the user to download or obtain the data. 

Sharing Components and Capabilities 

The following definitions are critical to understanding the components of the CMS Processing Environments. These definitions apply to all virtual and physical components of information processing systems: 

  • Dedicated components – are those physical resources provided exclusively for CMS Environments. 
  • Shared component – means the physical or virtual component is shared among multiple tenants in a data center, in a virtualization hosting system, or in a community or public cloud environment. Details of CMS cloud shared services are available at CMS Cloud Services. 
  • Independently managed component – means the physical or virtual component is independently usable and managed by CMS, in isolation from other tenants. 
  • Third-Party Websites and Applications – are web-based technologies (covered within the CMS ARS) that are not exclusively operated or controlled by HHS. The CMS TRA covers TPWAs linked to or accessed by CMS applications. 
  • Single-purpose component – means the physical or virtual component is a standalone custom or commercial device or appliance product (including hardware, firmware, software, and/or operating system) designed for a specific purpose, and not to perform general purpose processing. An example would be a firewall appliance that consists of firewall software running on a virtual machine (VM) with a highly customized operating system that is designed to run only the firewall application. 

Network Connectivity and Trust Boundaries 

CMS data centers and the CMSNet connections between them are within the CMS security perimeter and currently operate with an elevated level of trust. The movement towards newer security models such as zero trust will continue to reduce the inherent trust at the network connectivity level, and best practice is to authenticate all network connections. Although individual data centers may connect to the untrusted Internet, connections between CMS data centers should preferably traverse CMSNet. Connections to external entities must be covered by an Interconnection Security Agreement (ISA) approved by the CMS Authorizing Official. 

CMS Data and CMS Sensitive Information 

The CMS TRA uses the term “Sensitive Information” as defined in the NIST Glossary and the Guide to Cyber Threat Information Sharing, SP 800-150 and referenced in Executive Order 13556 -- Controlled Unclassified Information, with guidance in the NARA CUI Registry. It is the responsibility of the business owner, in consultation with their Group Director (GD), Information Systems Officer (ISO), Information Systems Security Officer (ISSO), Cyber Risk Advisor (CRA), Chief Information Security Officer (CISO), Administrative Officer (AO), Data Guardian, and Privacy subject matter expert to determine what data are sensitive. This includes all data that require protection due to the risk and magnitude of loss or harm, such as PII, PHI, and Federal Tax Information (FTI). 

There are some system data elements in ATO(ed) environments that should always be considered sensitive: 

  • Passwords and private keys, including application programming interface (API) keys and other keys used to access or configure sensitive CMS data, services, or IT resources 
  • Final release of code, executables, and configuration files that will be or are deployed in ATO(ed) environments 
  • Configuration “Reference Data” such as configuration files or metadata used for configuring virtual resources and services in ATO(ed) environments 
  • Internal network addresses, and unique addressable device identifiers—such as Media access control (MAC) addresses, Virtual Machine identities (ID), etc. in ATO(ed) environments 

The Privacy Impact Assessment (PIA) is a critical tool for spotting privacy risks and compliance with federal regulations or laws, tracking implementation of privacy controls, identifying instances where CMS collects or handles Personally Identifiable Information (PII) and/or Protected Health Information (PHI) and for identifying CMS systems subject to the Privacy Act of 1974.  PIAs must be conducted as part of the ATO process for CMS IT systems, and must be reviewed at least every three years and/or upon a major change to the IT system or electronic information collection. For additional information, reference the CMS Privacy Impact Assessment (PIA) Standard Operating Procedures and ARS SC-7(24) - Personally Identifiable Information. 

Designating High Value Assets 

A High Value Asset (HVA) is an asset used as a mission-critical information resource supporting infrastructure providers and suppliers or partnering organizations. The unauthorized disclosure, modification, destruction, or disruption of access to this information could be expected to have a severe or catastrophic adverse effect on organizational operations, organizational assets, or individuals. 

The business owner of a CMS information system must categorize the system in accordance with Federal Information Processing Standards (FIPS) 199, and document the system attributes used to identify PII, PHI, and HVAs. The CIO determines whether a system is an HVA. If the CMS CIO identifies a system as a High Value Asset, an HVA Designation Letter must be on file. Please refer to TRA Network Services, Cyber Security Operations / Risk Management for more information.  

Zero Trust Maturity Introduction 

Introduction 

The Federal Government has directed agencies to modernize their approach to cybersecurity. Executive Order 14028, “Improving the Nation’s Cybersecurity”, and OMB Memorandum M-22-09, “Moving the U.S. Government Toward Zero Trust Cybersecurity Principles” direct Federal Civilian Executive Branch (FCEB) agencies to base their enterprise security architecture on Zero Trust principles. While HHS and CMS have not published new policies regarding Zero Trust, the CMS Zero Trust Workgroup is working to evolve the Zero Trust Maturity of all CMS environments through incremental change, with “Advanced” or “Optimal” maturity being the objective for systems, based on their sensitivity (the results of a recent data call for Zero Trust Maturity for AWS for CMS Cloud suggests an overall maturity level of "Advanced" for CMS Cloud). 

CMS is in the midst of defining its Zero Trust strategy, policies, and approach. As such, this TRA section does not introduce any new business rules. However, to assist CMS in preparing for and aligning with Zero Trust objectives, the following topics illustrate how existing TRA business rules and recommended practices align with Zero Trust Maturity capabilities. This information can aid development teams in understanding which areas may need additional focus along the Zero Trust Journey. This is not per se a Zero Trust primer. See below for additional resources. 

Overview 

The Federal Zero Trust Architecture (ZTA) strategy involves migration from existing perimeter-based defenses to a “Zero Trust” approach. Zero Trust is not a single architecture, but a set of guiding principles that can improve the security posture of agency applications and environments. These seven tenets outlined in NIST SP 800-207 guide its implementation: 

  1. All data sources and computing services are considered resources. Including: 
  • Multiple classes of devices 
  • Small footprint devices 
  • Personally owned devices (BYOD) 
  1. All communication is secured regardless of network location. 
  • Both enterprise-owned network infrastructure and any other non-enterprise-owned network 
  • In the most secure manner available 
  • Access requests from inside must meet same requirements as from outside the enterprise 
  1. Access to individual enterprise resources is granted on a per-session basis. 
  • Trust in the requester is evaluated before the access is granted 
  • Access is granted with the least privileges needed to complete the task 
  • Authentication and authorization to one resource will not automatically grant access to a different resource 
  1. Access to resources is determined by dynamic policy, including: 
  • The observable state of client identity, application/service, and the requesting asset 
  • May include other behavioral and environmental attributes, such as security posture 
  • Rules and attributes are based on the needs of the business and acceptable level of risk 
  • Least privilege principles restrict both visibility and accessibility 
  1. The enterprise monitors and measures the integrity and security posture of all owned and associated assets. 
  • No asset is inherently trusted. 
  • Evaluates the security posture of the asset when evaluating a resource request. 
  • Continuous diagnostics and mitigation (CDM) systems monitor the state of devices and others. 
  1. All resource authentication and authorization are dynamic and strictly enforced before access is allowed. 
  2. The enterprise collects as much information as possible about the current state of assets, network infrastructure and communications and uses it to improve its security posture. 

CMS-specific guidance for ISSOs and ADOs can be found at The 7 Tenets of Zero Trust for ISSOs and ADOs. Zero Trust Maturity is the degree to which ZTA principles have been implemented across the agency. CMS assesses this using the CISA Zero Trust Maturity Model, which is organized around a structure that is shown with five pillars: 

Figure 1: Zero Trust Maturity Model Pillars (page 9) 

  • Identity: Federal staff, as well as partners and end users, use enterprise-managed accounts to access everything they need to do their job, protected from phishing and other attacks. 
  • Devices: The devices that Federal staff use are consistently tracked and monitored, with those devices’ security postures used to grant access. 
  • Networks: Agency systems are isolated, with encrypted network traffic flowing between and within them. 
  • Applications and Workloads: Enterprise applications can be made available to staff securely over the internet. 
  • Data: Federal security teams and data teams develop data categories and security rules to automatically detect and ultimately block unauthorized access. 

Below the pillars are three steps: 

  • Visibility and Analytics: Collecting information about the systems to identity how things in the Pillar are working. 
  • Automation and Orchestration: Methods for automatically creating and maintaining the different entities and assets in a Pillar. Manual configurations can introduce errors over time, and automation helps prevent that. 
  • Governance: The set of policies that tell how you control and direct different assets and entities in your environment. This can include how teams create new users, decide on data classifications, and manage servers. 

High-level Summary 

Below are the characteristics of the maturity levels assessed for CMS environments. 

Zero Trust Maturity Capabilities Summary 
Pillar Traditional Initial Advanced Optimal 
Identity 
  • Passwords or MFA 
  • On-premises identity stores 
  • Limited identity risk assessments 
  • Permanent access with periodic review 
  • MFA with passwords 
  • Self-managed and hosted identity stores 
  • Manual identity risk assessments 
  • Access expires with automated review 
  • Phishing-resistant MFA 
  • Consolidation and secure integration of identity stores 
  • Automated identity risk assessments 
  • Need/session-based access 
  • Continuous validation and risk analysis 
  • Enterprise-wide identity integration 
  • Tailored, as needed automated access 
Devices 
  • Manually tracking device inventory 
  • Limited compliance visibility 
  • No device criteria for resource access 
  • Manual deployment of threat protections to some devices 
  • All physical assets tracked 
  • Limited device-based access control and compliance enforcement 
  • Some protections delivered via automation 
  • Most physical and virtual assets are tracked 
  • Enforced compliance implemented with integrated threat protections 
  • Initial resource access depends on device posture 
  • Continuous physical and virtual asset analysis including automated supply chain risk management and integrated threat protections 
  • Resource access depends on real-time device risk analytics 
Networks 
  • Large perimeter / macro-segmentation 
  • Limited resilience and manually managed rulesets and configurations 
  • Minimal traffic encryption with ad hoc management 
  • Initial isolation of critical workloads 
  • Network capabilities manage availability demands for more applications 
  • Dynamic configurations for some portions of the network 
  • Encrypt more traffic and formalize key management policies 
  • Expanded isolation and resilience mechanisms 
  • Configurations adapt based on automated risk-aware application profile assessments 
  • Encrypts applicable network traffic and manages issuance and rotation of keys 
  • Distributed micro-perimeters with just-in-time and just-enough access controls and proportionate resilience 
  • Configurations evolve to meet application profile needs 
  • Integrates best practices for cryptographic agility 
Applications and Workloads 
  • Mission critical applications accessible via private networks 
  • Protections have minimal workflow integration 
  • Ad hoc development, testing, and production environments 
  • Some mission critical workflows have integrated protections and are accessible over public networks to authorized users 
  • Formal code deployment mechanisms through CI/CD pipelines 
  • Static and dynamic security testing prior to deployment 
  • Most mission critical applications available over public networks to authorized users 
  • Protections integrated in all application workflows with context-based access controls 
  • Coordinated teams for development, security, and operations 
  • Applications available over public networks with continuously authorized access 
  • Protections against sophisticated attacks in all workflows 
  • Immutable workloads with security testing integrated throughout lifecycle 
Data 
  • Manually inventory and categorize data 
  • On-prem data stores 
  • Static access controls 
  • Minimal encryption of data at rest and in transit with ad hoc key management 
  • Limited automation to inventory data and control access 
  • Begin to implement a strategy for data categorization 
  • Some highly available data stores 
  • Encrypts data in transit 
  • Initial centralized key management policies 
  • Automated data inventory with tracking 
  • Consistent, tiered, targeted categorization and labeling 
  • Redundant, highly available data stores 
  • Static DLP 
  • Automated context-based access 
  • Encrypts data at rest 
  • Continuous data inventorying 
  • Automated data categorization and labeling enterprise-wide 
  • Optimized data availability DLP exfil blocking 
  • Dynamic access controls 
  • Encrypts data in use 

 

Resources and References 

Find more information about Zero Trust within CMS OIT and the Federal Government at large. Some are out of the scope of the CMS migration. 

Internal 

External 

OMB 

NIST 

CISA 

  

Zero Trust Maturity Model Pillars 

The sections that follow show ways in which the CMS TRA aligns with capabilities required for Zero Trust Maturity. The CMS objective is to implement these capabilities at the Advanced or Optimal maturity level.  

 

Zero Trust Maturity Identity Pillar 

Introduction 

This section covers the capabilities needed for the Zero Trust Maturity Identity Pillar. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

The identity pillar considers how accounts are created and how users log into systems. The ability to identify every user and entity requesting system access is foundational to the concept of zero trust. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Identity Pillar. This includes: 

  • How to store identities of users as well as authenticate them 
  • Consideration of identity for the developers and admins creating the system and the end users of the system like providers or the public 
  • Consideration of non-person entities such as devices, service accounts, and APIs 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Authentication capabilities 
Traditional Initial Advanced Optimal 
Agency authenticates identity using either passwords or multi-factor authentication (MFA) with static access for entity identity. Agency authenticates identity using MFA, which may include passwords as one factor and requires validation of multiple entity attributes (e.g., locale or activity). Agency begins to authenticate all identity using phishing-resistant MFA and attributes, including initial implementation of passwordless MFA. Agency continuously validates identity with phishing-resistant MFA, not just when access is initially granted. 
Zero Trust Identity Stores capabilities 
Traditional Initial Advanced Optimal 
Agency only uses self-managed, on-premises (i.e., planned, deployed, and maintained by agency) identity stores. Agency has a combination of self-managed identity stores and hosted identity store(s) (e.g., cloud or other agency) with minimal integration between the store(s) (e.g., Single Sign-on.). Agency begins to securely consolidate and integrate some self-managed and hosted identity stores. Agency securely integrates their identity stores across all partners and environments as appropriate. 
Zero Trust Risk Assessment capabilities 
Traditional Initial Advanced Optimal 
Agency makes limited determinations for identity risk (i.e., likelihood that an identity is compromised). Agency determines identity risk using manual methods and static rules to support visibility. Agency determines identity risk with some automated analysis and dynamic rules to inform access decisions and response activities. Agency determines identity risk in real time based on continuous analysis and dynamic rules to deliver ongoing protection. 
  • BR-SEC-Gen-22: All Information Systems Must Have a System Risk Assessment in CFACTS 
  • BR-CCIC-02: Assessment of Information Security and Privacy Risks 
Zero Trust Access Management capabilities 
Traditional Initial Advanced Optimal 
Agency authorizes permanent access with periodic review for both privileged and unprivileged accounts Agency authorizes access, including for privileged access requests, that expires with automated review. Agency authorizes need-based and session-based access, including for privileged access request, that is tailored to actions and resources. Agency uses automation to authorize just-in-time and just-enough access tailored to individual actions and individual resource needs. 
  • BR-ACID-1: Valid Purpose Required to Access CMS Information Systems 
  • BR-ACID-11: CMS Business Owners Provide Privilege Administration 
  • BR-ACID-13: OIT Is Responsible for Identity Management of Users with Credentials Provisioned in the CMS Enterprise Directory 

Zero Trust Maturity Devices Pillar 

Introduction 

This section covers the capabilities needed for the Zero Trust Maturity Devices Pillar. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

The devices pillar highlights the importance of consistently tracking and monitoring devices to understand their security posture prior to granting access to systems. For applications, this will include physical servers as well as digital assets such as virtual machines and containers. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Device Pillar. This includes: 

  • How to maintain a complete inventory of every device a system operates through the CDM program 
  • Managing supply chain risks for both devices and the software it runs 
  • Preventing and detecting incidents on those devices 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Policy Enforcement & Compliance Monitoring Capabilities 
Traditional Initial Advanced Optimal 
Agency has limited, if any, visibility (i.e., ability to inspect device behavior) into device compliance with few methods of enforcing policies or managing software, configurations, or vulnerabilities. 

Agency receives self-reported device characteristics (e.g., keys, tokens, users, etc., on the device) but has limited enforcement mechanisms. 

Agency has a preliminary, basic process in place to approve software use and push updates and configuration changes to devices. 

Agency has verified insights (i.e., an administrator can inspect and verify the data on device) on initial access to device and enforces compliance for most devices and virtual assets. 

Agency uses automated methods to manage devices and virtual assets, approve software, and identify vulnerabilities and install patches. 

Agency continuously verifies insights and enforces compliance throughout the lifetime of devices and virtual assets. 

Agency integrates device, software, configuration, and vulnerability management across all agency environments, including for virtual assets. 

Zero Trust Asset & Supply Chain Risk Management Capabilities 
Traditional Initial Advanced Optimal 
Agency does not track physical or virtual assets in an enterprise-wide or cross-vendor manner and manages its own supply chain acquisition of devices and services in ad hoc fashion with a limited view of enterprise risks. Agency tracks all physical and some virtual assets and manages supply chain risks by establishing policies and control baselines according to federal recommendations using a robust framework, (e.g., NIST SCRM.) Agency begins to develop a comprehensive enterprise view of physical and virtual assets via automated processes that can function across multiple vendors to verify acquisitions, track development cycles, and provide third-party assessments. Agency has a comprehensive, at-or near-real-time view of all assets across vendors and service providers, automates its supply chain risk management as applicable, builds operations that tolerate supply chain failures, and incorporates best practices. 
Zero Trust Resource Access (Formerly Data Access) Capabilities 
Traditional Initial Advanced Optimal 
Agency does not require visibility into devices or virtual assets used to access resources. Agency requires some devices or virtual assets to report characteristics then use this information to approve resource access. Agency’s initial resource access considers verified device or virtual asset insights. Agency’s resource access considers real-time risk analytics within devices and virtual assets. 
Zero Trust Device Threat Protection Capabilities 
Traditional Initial Advanced Optimal 
Agency manually deploys threat protection capabilities to some devices. Agency has some automated processes for deploying and updating threat protection capabilities to devices and to virtual assets with limited policy enforcement and compliance monitoring integration. Agency begins to consolidate threat protection capabilities to centralized solutions for devices and virtual assets and integrates most of these capabilities with policy enforcement and compliance monitoring. Agency has a centralized threat protection security solution(s) deployed with advanced capabilities for all devices and virtual assets and a unified approach for device threat protection, policy enforcement, and compliance monitoring. 
  • BR-PMM-3: Monitor Production Environments 
  • BR-SEC-Gen-11: Periodic and Continuous Configuration Compliance Scanning Is Required 
  • BR-SEC-Gen-12: Host Intrusion Detection Capabilities on All IT Components 
  • BR-SEC-Gen-21: Malware and Malicious Code Scanning Results Must Be Sent to the Security Zone 
  • BR-CCIC-09: Local Information Sharing and Cyber Threat Intelligence Support 
  • BR-CCIC-25: Insider Threat Detection 

 

Zero Trust Maturity Networks Pillar 

Introduction 

This section covers the capabilities needed for the Zero Trust Maturity Networks Pillar. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

A network refers to an open communications medium including typical channels such as agency internal networks, wireless networks, and the Internet as well as other potential channels such as cellular and application-level channels used to transport messages. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Network Pillar. This includes: 

  • Encrypting traffic that leaves the CMS boundary, as well as encrypting internal traffic 
  • Focusing on network resilience from both normal use and adversaries 
  • The need to begin executing a plan to break down the perimeters into isolated environments 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Network Segmentation Capabilities 
Traditional Initial Advanced Optimal 

Agency defines their network architecture using large perimeter/macro-segmentation with minimal restrictions on reachability within network segments. 

Agency may also rely on multi-service interconnections (e.g., bulk traffic VPN tunnels). 

Agency begins to deploy network architecture with the isolation of critical workloads, constraining connectivity to least function principles, and a transition toward service-specific interconnections. Agency expands deployment of endpoint and application profile isolation mechanisms to more of their network architecture with ingress/egress micro-perimeters and service-specific interconnections. Agency network architecture consists of fully distributed ingress/egress micro-perimeters and extensive micro-segmentation based around application profiles with dynamic just-in-time and just-enough connectivity for service-specific interconnections. 
Zero Trust Network Traffic Management Capabilities 
Traditional Initial Advanced Optimal 
Agency manually implements static network rules and configurations to manage traffic at service provisioning, with limited monitoring capabilities (e.g., application performance monitoring or anomaly detection) and manual audits and reviews of profile changes for mission critical applications. 

Agency establishes application profiles with distinct traffic management features and begins to map all applications to these profiles. 

Agency expands application of static rules to all applications and performs periodic manual audits of application profile assessments. 

Agency implements dynamic network rules and configurations for resource optimization that are periodically adapted based upon automated risk-aware and risk-responsive application profile assessments and monitoring. Agency implements dynamic network rules and configurations that continuously evolve to meet application profile needs and reprioritize applications based on mission criticality, risk, etc. 
Zero Trust Traffic Encryption Capabilities 
Traditional Initial Advanced Optimal 
Agency encrypts minimal traffic and relies on manual or ad hoc processes to manage and secure encryption keys. Agency begins to encrypt all traffic to internal applications, to prefer encryption for traffic to external applications, to formalize key management policies, and to secure server/service encryption keys. Agency ensures encryption for all applicable internal and external traffic protocols, manages issuance and rotation of keys and certificates, and begins to incorporate best practices for cryptographic agility. Agency continues to encrypt traffic as appropriate, enforces least privilege principles for secure key management enterprise-wide, and incorporates best practices for cryptographic agility as widely as possible. 
  • BR-SA-6: Network Communications Must Meet the TRA Rules for Encryption 
  • BR-SEC-Int-3: HTTPS on CMS Public-Facing Websites and Services on the Internet 
  • BR-CCIC-24: FIPS 140-2 or FIPS 140-3 Validated Encryption Use 
  • BR-WAN-S-0: Use Mutual Authentication and Encrypted Tunnels between Data Centers 
  • BR-WAN-S-1: The WAN Must Implement FIPS 140-2 or FIPS 140-3 Compliant Encryption 
Zero Trust Network Resilience Capabilities 
Traditional Initial Advanced Optimal 
Agency configures network capabilities on a case-by-case basis to only match individual application availability demands with limited resilience mechanisms for workloads not deemed mission critical. Agency begins to configure network capabilities to manage availability demands for additional applications and expand resilience mechanisms for workloads not deemed mission critical. Agency has configured network capabilities to dynamically manage the availability demands and resilience mechanisms for the majority of their applications. Agency integrates holistic delivery and awareness in adapting to changes in availability demands for all workloads and provides proportionate resilience. 

 

Zero Trust Maturity Applications & Workloads 

Introduction 

This section covers the capabilities needed for the Zero Trust Maturity Applications and Workloads Pillar. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

Applications and workloads include systems, computer programs, and services that execute in on-premises and cloud environments. In mature zero trust deployments, users strongly authenticate into applications, not into the underlying networks. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Zero Trust Maturity for AWS for CMS Cloud. This includes: 

  • Application-specific threat protections 
  • Application security testing at all stages of development and deployment 
  • Enabling access to applications based on additional user attributes beyond mere presence on specific networks 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Application Access Capabilities 
Traditional Initial Advanced Optimal 
Agency authorizes access to applications primarily based on local authorization and static attributes. Agency begins to implement authorizing access capabilities to applications that incorporate contextual information (e.g., identity, device compliance, and/or other attributes) per request with expiration. Agency automates application access decisions with expanded contextual information and enforced expiration conditions that adhere to least privilege principles. Agency continuously authorizes application access, incorporating real-time risk analytics and factors such as behavior or usage patterns. 
Zero Trust Application Threat Protections Capabilities 
Traditional Initial Advanced Optimal 
Agency threat protections have minimal integration with application workflows, applying general purpose protections for known threats Agency integrates threat protections into mission critical application workflows, applying protections against known threats and some application-specific threats. Agency integrates threat protections into all application workflows, protecting against some application-specific and targeted threats. Agency integrates advanced threat protections into all application workflows, offering real-time visibility and content-aware protections against sophisticated attacks tailored to applications. 
Zero Trust Accessible Applications Capabilities 
Traditional Initial Advanced Optimal 
Agency makes some mission critical applications available only over private networks and protected public network connections (e.g., VPN) with monitoring. Agency makes some of their applicable mission critical applications available over open public networks to authorized users with need via brokered connections. Agency makes most of their applicable mission critical applications available over open public network connections to authorized users as needed. Agency makes all applicable applications available over open public networks to authorized users and devices, where appropriate, as needed. 
Zero Trust Secure Application Development and Deployment Workflow Capabilities 
Traditional Initial Advanced Optimal 
Agency has ad hoc development, testing, and production environments with non-robust code deployment mechanisms. Agency provides infrastructure for development, testing, and production environments (including automation) with formal code deployment mechanisms through CI/CD pipelines and requisite access controls in support of least privilege principles. Agency uses distinct and coordinated teams for development, security, and operations while removing developer access to production environment for code deployment. Agency leverages immutable workloads where feasible, only allowing changes to take effect through redeployment, and removes administrator access to deployment environments in favor of automated processes for code deployment. 
Zero Trust Application Security Testing Capabilities 
Traditional Initial Advanced Optimal 
Agency performs application security testing prior to deployment, primarily via manual testing methods. Agency begins to use static and dynamic (i.e., application is executing) testing methods to perform security testing, including manual expert analysis, prior to application deployment. Agency integrates application security testing into the application development and deployment process, including the use of periodic dynamic testing methods. Agency integrates application security testing throughout the software development lifecycle across the enterprise with routine automated testing of deployed applications. 

 

Zero Trust Maturity Data Pillar 

Introduction 

This section covers the capabilities needed for the Zero Trust Maturity Data Pillar. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

Data includes all structured and unstructured files and fragments that reside or have resided in federal systems, devices, networks, applications, databases, infrastructure, and backups (including on-premises and virtual environments) as well as the associated metadata. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Data Pillar. This includes: 

  • How data should be protected on devices, in applications, and on networks 
  • How data should be inventoried, categorized, and labeled, as well as protected at rest and in transit 
  • The advantage of cloud security services for monitor access to sensitive data and the preferred practice of implementing enterprise-wide logging and information sharing 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Data Inventory Management Capabilities 
Traditional Initial Advanced Optimal 
Agency manually identifies and inventories some agency data (e.g., mission critical data). Agency begins to automate data inventory processes for both on-premises and in cloud environments, covering most agency data, and begins to incorporate protections against data loss. Agency automates data inventory and tracking enterprise-wide, covering all applicable agency data, with data loss prevention strategies based upon static attributes and/or labels. Agency continuously inventories all applicable agency data and employs robust data loss prevention strategies that dynamically block suspected data exfiltration. 
Zero Trust Data Categorization Capabilities 
Traditional Initial Advanced Optimal 
Agency employs limited and ad hoc data categorization capabilities. Agency begins to implement a data categorization strategy with defined labels and manual enforcement mechanisms. Agency automates some data categorization and labeling processes in a consistent, tiered, targeted manner with simple, structured formats and regular review. Agency automates data categorization and labeling enterprise-wide with robust techniques; granular, structured formats; and mechanisms to address all data types. 
Zero Trust Data Availability Capabilities 
Traditional Initial Advanced Optimal 
Agency primarily makes data available from on-premises data stores with some off-site backups. Agency makes some data available from redundant, highly available data stores (e.g., cloud) and maintains off-site backups for on-premises data. Agency primarily makes data available from redundant, highly available data stores and ensures access to historical data. Agency uses dynamic methods to optimize data availability, including historical data, according to user and entity need. 
Zero Trust Data Access Capabilities 
Traditional Initial Advanced Optimal 
Agency governs user and entity access (e.g., permissions to read, write, copy, grant others access, etc.) to data through static access controls. Agency begins to deploy automated data access controls that incorporate elements of least privilege across the enterprise. Agency automates data access controls that consider various attributes such as identity, device risk, application, data category, etc., and are time limited where applicable. Agency automates dynamic just-in-time and just-enough data access controls enterprise-wide with continuous review of permissions. 
Zero Trust Data Encryption Capabilities 
Traditional Initial Advanced Optimal 
Agency encrypts minimal agency data at rest and in transit and relies on manual or ad hoc processes to manage and secure encryption keys. Agency encrypts all data in transit and, where feasible, data at rest (e.g., mission critical data and data stored in external environments) and begins to formalize key management policies and secure encryption keys. Agency encrypts all data at rest and in transit across the enterprise to the maximum extent possible, begins to incorporate cryptographic agility, and protects encryption keys (i.e., secrets are not hard coded and are rotated on a regular basis). Agency encrypts data in use where appropriate, enforces least privilege principles for secure key management enterprise-wide, and applies encryption using up-to-date standards and cryptographic agility to the extent possible. 

 

Zero Trust Maturity Cross-Cutting Capabilities 

Introduction 

This section covers the cross-cutting capabilities needed for the Zero Trust Maturity Foundation. It illustrates how existing TRA Business Rules and Recommended Practices align with these capabilities. 

CMS Guidance 

Each pillar also implements three cross-cutting capabilities. Visibility and Analytics refers to how we monitor systems. Automation and Orchestration is the process of creating reusable processes that can be automated within our systems. And Governance is the policies we set for the systems as well as how we track how we enforce those policies. 

The CMS Zero Trust Workgroup is developing guidelines for CMS ADOs. Specific guidance can be found in CMS Cloud documentation: Zero Trust Maturity for AWS for CMS Cloud. This includes: 

  • Using existing logging, monitoring, and alerting infrastructure where possible 
  • Centralizing the implementation of the cross-cutting capabilities over all five of the other pillars 
  • Documenting policies and procedures so they can be automated 

Capabilities 

The business rules shown beneath each capability aren’t comprehensive — Other requirements are defined in the ARS and RMH. 

Zero Trust Visibility and Analytics Capabilities 
Traditional Initial Advanced Optimal 
Agency manually collects limited logs across their enterprise with low fidelity and minimal analysis. Agency begins to automate the collection and analysis of logs and events for mission critical functions and regularly assesses processes for gaps in visibility. Agency expands the automated collection of logs and events enterprise-wide (including virtual environments) for centralized analysis that correlates across multiple sources. Agency maintains comprehensive visibility enterprise-wide via centralized dynamic monitoring and advanced analysis of logs and events. 
Zero Trust Automation and Orchestration Capabilities 
Traditional Initial Advanced Optimal 
Agency relies on static and manual processes to orchestrate operations and response activities with limited automation. Agency begins automating orchestration and response activities in support of critical mission functions. Agency automates orchestration and response activities enterprise-wide, leveraging contextual information from multiple sources to inform decisions. Agency orchestration and response activities dynamically respond to enterprise-wide changing requirements and environmental changes. 
Zero Trust Governance Capabilities 
Traditional Initial Advanced Optimal 
Agency implements policies in an ad hoc manner across the enterprise, with policies enforced via manual processes or static technical mechanisms. Agency defines and begins implementing policies for enterprise-wide enforcement with minimal automation and manual updates. Agency implements tiered, tailored policies enterprise-wide and leverages automation where possible to support enforcement. Access policy decisions incorporate contextual information from multiple sources. Agency implements and fully automates enterprise-wide policies that enable tailored local controls with continuous enforcement and dynamic updates. 

 

CMS TRA Business Rules 

Guidance in the CMS TRA includes business rules, which are requirements for TRA compliance, and recommended practices, which are strongly encouraged but not required. The format for BRs and RPs is similar. The business rules consist of a brief, binding requirement and a rationale that provides context on the intent of the rule. Where applicable, the rationale also contains references to relevant policies, standards, and specific CMS ARS controls. The following BRs for this topic (i.e., rule numbers beginning with BR-F) provide high-level guidance applicable to all TRA stakeholders. Other topics contain additional business rules relevant to their respective sections. 

BR-F-1: Any Deviations from the CMS TRA Must Be Requested and Approved 

All projects and applications are required to adhere to the CMS TRA unless covered by a Technology Review Board (TRB)-approved exception. The CMS CEA and CIO has have joint approval authority for the CMS TRA. Only the TRB can approve deviations from the CMS TRA. 

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration and CA-6 - Authorization. 

Rationale: 

ARS control CM-2 requires “Baseline configurations of information systems reflect the current enterprise architecture.” The CMS TRA is the CMS Enterprise Architecture standard. In some cases, applications identify circumstances where the benefits of deviating from the CMS TRA may outweigh the risks. The risk/benefit of any deviation from the CMS TRA must be assessed and approved by the TRB before the CIO can determine if the risk is acceptable and the system should be granted an ATO. 

BR-F-2: The CMS TRA Applies to All CMS Processing Environments 

The CMS TRA provides the standard technical reference for all CMS Processing Environments. This includes both physical and virtual server and end-user (e.g., Virtual Desktop Infrastructure) environments, cloud-based environments, and hybrid environments. These standards are effective on publication. A CMS Processing Environment is defined as any computing environment (e.g., CMS data center, virtual computing environment, or cloud computing) that creates, consumes, and/or stores CMS-related data. CMS data includes sensitive and non-sensitive information, security information, and event management-related information used to provide CMS services to the public and internal CMS users. For existing systems, compliance with CMS TRA updates is required within twenty-four (24) months of publication. When certain changes require immediate compliance, CMS will communicate these changes outside the TRA process via CIO directive. 

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration and SA-8 - Security and Privacy Engineering Principles, Implementation Standard 1: 

The information system must follow system security and privacy engineering principles consistent with: 

  1. The information security steps of the CMS Target Life Cycle (TLC) to incorporate information security and privacy control considerations; 
  2. The information system architecture defined within the Technical Reference Architecture (TRA); and 
  3. The Technical Review Board (TRB) processes defined by CMS. 

Rationale: 

Updates to the CMS TRA require corresponding changes to all CMS information systems. A 24-month compliance window provides sufficient time for data center operators and application maintainers to plan for and execute required CMS TRA changes to bring systems into compliance. 

BR-F-3: The CMS TRA Defines a Zoned Architecture 

At its most fundamental level, the CMS TRA defines a zoned architecture, a type of services framework architecture (see BR-F-22), where each service type provides different functions and has different responsibilities. There are three application zones and two infrastructure zones. The three business application zones are the Presentation/Edge, Application, and Data Zones. The two infrastructure zones are the Management Zone and the Security Zone. See additional information in the CMS Services Framework. 

In the multi-zone architecture, and specifically in the cloud environment, CMS does not specify the number of zones, but does require that access to CMS data resources be protected by at least three security challenges (e.g. multi-factor authentication, security certificates, etc. See Mediation Principles). This ensures that a user or other resource that requires access CMS sensitive data has been properly vetted prior to gaining access to the sensitive resources. This leads to a hierarchy, where zones have bi-directional adjacency. The Presentation/Edge Zone is adjacent to the Application Zone. (Please note, there are restrictions on the storage of data within the Presentation/Edge Zone. See BR-SA-3 for additional information.) The Application Zone is adjacent to both the Presentation/Edge Zone and the Data Zone. The Management and Security Zones are adjacent to all zones. The following rules govern the zoned architecture: 

  1. No other adjacency relationships are permitted. 
  2. Firewalls (virtual or physical) separate adjacent zones. 
  3. The three business application zones (the Presentation, Application, and Data Zones) extend across all CMS Processing Environments. 
  4. There are no other zones. 

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration, PL-8 - Security and Privacy Architectures, RA-9 - Criticality Analysis, SA-8 - Security and Privacy Engineering Principles, SC-2 - Separation of System and User Functionality, and SC-32 - Information System Partitioning. 

Rationale: 

The CMS TRA follows a Defense-in-Depth strategy implemented in part by a multi-zone architecture. Access to the most valuable parts of the architecture (typically the data) requires crossing multiple zones, including multiple firewalls and servers (which may be physical, virtual, or configured via cloud-based services). The CMS Multi-Zone Architecture defines a consistent architecture used across the enterprise. In addition to Defense-in-Depth for each application, this architecture allows inter-application communication while maintaining a secure environment. 

Appendix I, section 4 of Office of Management and Budget (OMB) Circular No. A-130, Managing Information as a Strategic Resource, revised July 28,2016, directs, in pertinent part, under sub-paragraph i (4) that federal agencies shall: 

Isolate sensitive or critical information resources (e.g., information systems, system components, applications, databases, and information) into separate security domains with appropriate levels of protection based on the sensitivity or criticality of those resources; 

BR-F-4: Within a CMS Processing Environment, Communication Must Flow Only between Adjacent Zones or within a Single Zone 

Network communication must flow only between adjacent zones (as described in BR-F-3). Direct communication between non-adjacent zones is prohibited. 

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration, SC-32 - Information System Partitioning, and SC-7 - Boundary Protections. 

Rationale: 

Communication limited to zone adjacency implements the CMS TRA Defense-in-Depth strategy. Network Services provides details. 

BR-F-5: Any System That Processes CMS Data Must Be Covered by a CMS ATO 

A CMS processing system is any environment dedicated to CMS FISMA applications or services (in whole or part) that process and/or store CMS data. A CMS processing system must be covered under a CMS Authorization to Operate (ATO). Additional guidance regarding CMS data is provided in CMS Data and Sensitive Information. 

All systems that process CMS enterprise data, whether called “production,” “test,” “pilot,” “proof-of-concept,” or “other,” must be covered under a CMS ATO in compliance with the Federal Information Security Modernization Act (FISMA)). This rule applies to any environment that stores or processes CMS data, to include disaster recovery (DR) sites and archival storage. Lower environments that contain only test data might not require an ATO. 

See the CMS ATO website for additional information regarding the CMS ATO process and the different types of compliance authorizations provided by CMS to manage agency-wide risk. For further information about test data and “synthetic data generation,” see NIST SP 800-188, “De-Identifying Government Datasets: Techniques and Governance.” 

Related CMS ARS Security Controls include: CA-6 - Authorization, CM-2 - Baseline Configuration, AC-21 - Information Sharing, SC-32 - Information System Partitioning, RA-2 - Security Categorization, andSA-3(2) - Use of Live Operational Data. 

Rationale: 

The “Types of authorizations” section of the CMS ATO website states: 

“Every system that is integrated at CMS — either built in-house or contracted — must get a compliance authorization to operate and access government data. This ensures that the agency is aware of all components interacting with its data, and that each system can be monitored for compliance and risk mitigation. This helps safeguard sensitive personal information, manage the risk to critical infrastructure, and address cybersecurity issues when they arise. 

“If you are introducing a new system at CMS, you must go through the security and compliance process.” 

The emphasis on CMS ATO is important to shared environments such as clouds, which may be covered under the Federal Risk and Authorization Management Program (FedRAMP) or another agency’s ATOs. While other agencies may have different ATO requirements, the compliance with such an ATO or FedRAMP does not ensure that all CMS-prescribed controls have been addressed. 

BR-F-6: Mainframes Must Be Dedicated to CMS 

CMS must be the sole user of the mainframe resources, when those resources are operated with an ATO and access CMS data, as detailed in the following paragraphs. 

  • For Processing Resources: CMS’s IBM Mainframe Logical Partitions (LPAR) may be used on the same machine as other data center tenants (including the data center operator), provided the LPARs processors are dedicated to CMS use only when processing CMS data. 
  • For Storage Resources: If storage is secured via the LPAR, dedicated storage is not required. Storage allocation on the hardware must be dedicated to the CMS LPAR. Thus, storage media may not be shared between CMS and other tenants. Note that this rule applies to online, nearline, and offline storage. 

For all resources, an appropriate Authorization, Auditing and Authentication tool (AAA) for access control must be used to protect both CMS data and processing capability. 

Rationale: 

CMS has two broad concerns with sharing mainframes. The first concern is protecting CMS data. CMS wants to ensure that CMS data is protected and separated from other user workloads. As a result, CMS considers it essential to isolate CMS processing resources from other tenants (including the hosting provider). For example, although storage hardware such as tape silos may be shared with other tenants, the storage medium such as tapes may not be. 

The second concern is sharing infrastructure among non-CMS tenants. Shared infrastructure, including the mainframe that hosts the LPARs and any related networking or other components shared by the LPARs, constitutes both performance and security risks. The performance risk is that shared infrastructure may be unable to meet CMS performance needs. The security risk of shared infrastructure includes, for example, potential for data leakage and lessened availability due to misconfiguration or lack of coordination or oversubscription. 

In addition, CMS needs assurance that it is not subsidizing other tenant workloads on CMS-funded environments. 

Mainframe LPAR processors need not be dedicated to CMS when used for development or test purposes, so long as no CMS data is processed. 

Note that this business rule also applies to CMS mainframe resources used for Disaster Recovery. 

BR-F-7: Cost-Effective Reuse of Data Centers with Established TIC, CMSNet, and CCIC Integration 

Projects must reuse data centers with established TIC, CMSNet, and CCIC integration. 

Related CMS ARS Security Controls include: SA-2 - Allocation of Resources. 

Rationale: 

This rule supports cost savings related to OMB-mandated physical data center consolidation (Update to Data Center Consolidation Initiative (DCOI), OMB/Federal CIO Kent, June 25, 2019) and more efficient use of security monitoring services. As stated in the OMB memorandum, data center consolidation will promote the use of Green IT by reducing the overall energy and real estate footprint of government data centers; reduce the cost of data center hardware, software, and operations; increase the overall IT security posture of the government; and shift IT investments to more efficient computing platforms and technologies. 

BR-F-8: Backup CMS Data 

All CMS data must be backed up regularly, on a documented schedule, regardless of hosting implementation (i.e., physical data center, Cloud, etc.). 

ARS Security Control CP-9 provides specific requirements for backup. Note that Cloud Service Providers have additional requirements under this control. All backups must be secured from unauthorized access and disclosure. 

Related CMS ARS Security Controls include: CP-6 - Alternate Storage Site, AU-2 - Event Logging, CP-9 - System Backup, and CP-9(8) - Cryptographic Protection. 

Rationale: 

Disaster Recovery and Fault Tolerance (FT) is sometimes confused with data backup. Data backup is an essential part of any CMS Processing Environment because it offers the ability to restore data to the most recent copy as well as one of several older copies. This backup capability facilitates comparative analysis and recovery from data corruption, neither of which can be handled by DR or FT. 

The OIT Backups Policy provides details about implementation. The basic requirement is to meet the “return to service” requirements per the system’s ATO. AWS Backups for AWS and Azure Backups for Microsoft Azure for Government (MAG) cloud are preferred, replacing the Cloud Protection Manager (CPM). 

Backups also differ from archives, which are performed for Records Management. 

PREFERRED - The CMS recommended backup tools are AWS Backups and Azure Backups in Microsoft Azure for Government (MAG). 

BR-F-9: Test CMS Backups on a Documented Schedule 

CMS backups and associated procedures must be tested for corruption and completeness on a regular basis to ensure that restoration can occur, if need be. 

Rationale: 

Improper testing of backups is a frequent cause of operational issues. Restoring backups and comparing against a master copy is one way to test backups. System maintainers are encouraged to explore alternatives to select the most appropriate method for their projects. 

Related CMS ARS Controls include: CP-9(1) - Testing for Reliability/Integrity and CM-4(1) - Separate Test Environments. 

BR-F-10: Annual Review and Exercise of Data Center Disaster Recovery Plans 

Disaster recovery plans and their supporting documents must be reviewed and re-evaluated annually or on a significant change to the operating environment. In addition, DR plans must be exercised to ensure that the plans are complete and accurate and to make certain that authorized personnel can execute their training. 

Rationale: 

The review and exercise of DR plans assures that the plans are accurate and up to date and that all parties understand their roles in the event of a disaster. 

Related CMS ARS Controls include: CP-8 - Telecommunications Services and CP-10 - System Recovery and Reconstitution. 

BR-F-11: Annual Review and Exercise of Contingency Plans 

Contingency Plans (CP) and any supporting documents must be reviewed and re-evaluated annually or on a significant change to the operating environment. In addition, the CPs must be exercised to ensure that the plans are complete and accurate and to ensure that personnel can execute their training (see CMS Information System Contingency Plan (ISCP) Handbook). 

Rationale: 

The review and exercise of CPs assures that the plans are accurate and up to date and that all parties understand their roles in the event of a disaster. 

Related CMS ARS Controls include: CP-10 - System Recovery and Reconstitution and CP-4 - Contingency Plan Testing. 

BR-F-12: Role-Based Security AAA Must Be Used for Management and User Roles 

Use Role-based Security Authentication, Authorization, and Accounting (AAA) functions provided by a shared AAA server within the data center to ensure that management and user roles are common throughout a zone. 

Related CMS ARS Controls include: AC-3 - Access Control, AC-5 - Separation of Duties, and AC-6 - Least Privilege. 

Rationale: 

This rule prevents accidental or unauthorized access. 

BR-F-13: Consistent Security Categorization within ARS Security Boundary 

All applications or systems within the same security control assessment boundary must implement security controls that meet the requirements for the highest Federal Information Processing Standards (FIPS)-199 security categorization level rating of any of the applications or systems. 

Related CMS ARS Security Controls include: RA-2 - Security Categorization, CA-6 - Authorization, and SC-32 - Information System Partitioning. 

Rationale: 

Security is only as good as the weakest link. All systems within the same Security Control Assessment (SCA) Boundary must be hardened adequate to prevent compromise of the most sensitive system or data. A compromised system may be used to attack other systems within the boundary. 

BR-F-14: Applications with Disparate FIPS-199 Security Categorization Levels Must Not Be Hosted on the Same Server 

BR-F-13 also establishes requirements related to this recommended practice. 

Related CMS ARS Security Controls include: RA-2 - Security Categorization, CA-6 - Authorization, and SC-32 - Information System Partitioning. 

Rationale: 

Co-locating systems with disparate ratings is not a normal practice because providing additional security controls to bring FIPS Low-rated applications to the FIPS Moderate or a higher grade, as required by BR-F-13, will increase the complexity and costs of the process. 

BR-F-15: Ensure Timely Version, Patch, and Configuration Management Practices 

All organizations responsible for operations and maintenance (O&M) of CMS Processing Environments must develop and employ timely version, patch, and configuration management practices to protect CMS component hardware, software, and data against threats to confidentiality, integrity, and availability while minimizing negative impacts to business operations. These practices must include coordination with other related CMS Processing Environment operators and affected application owners. 

Related CMS ARS Security Controls include: CM-1 - Policies and Procedures, CM-3 - Configuration Change Control, and PL-2 - System Security and Privacy Plan. 

Rationale: 

Delaying deployment of new security features and patches leaves CMS IT assets vulnerable. Untimely and uncoordinated deployment of new security features and patches can also interfere with business operations, disrupt the normal operation of IT systems, or introduce unexpected issues affecting the interoperability, performance, or access to CMS IT systems. Therefore, having effective and efficient version, patch, and configuration management practices are essential to the operations of CMS Processing Environments. 

BR-F-16: All Hosts Must Share a Common, Authenticated Time Server 

Rationale: 

This rule ensures that audit logs can be correlated between systems. Originally, this rule was non-production only, but it applies to all OS instances (such as VM, physical, cloud compute and containers that run their own Network Time Protocol Daemon) as well as networking equipment. 

BR-F-17: The CMS TRA Applies to Custom-Produced as well as COTS Products and Services 

Rationale: 

The CMS TRA is an integrated approach to system and security engineering. Commercial Off-the-Shelf (COTS) products as well as custom-developed software must comply with the CMS TRA to ensure interoperability within and integration into the CMS Processing Environments. This compliance must be validated prior to acquisition. The TRB may allow an exception for COTS products that cannot be made to comply; however, these deviations must be documented, reviewed, and approved by the TRB. 

BR-F-18: Use .Gov Domain Names for All CMS Internet Traffic 

All CMS Internet traffic for Agency business, including all web traffic and email, must use domain names ending in “.gov”. Requests for the use of second-level domain names not previously used for CMS Agency business must be approved by the CMS Office of Communications. 

Rationale: 

Both the OMB and HHS have policies prohibiting the use of most top-level domains for agency business. 

The OMB Memorandum M-23-22, Delivering a Digital-First Public Experience, September 22, 2023, stipulates each agency must use only an approved .gov or .mil domain for its official public-facing websites. The requirement to use only approved government domains does not apply in circumstances where the agency is a user or a customer of a third-party website or service that resides on a non-governmental domain. 

The HHS Internet Domain Names Policy regulates the usage, approval, acquisition, and registration of HHS Internet domain names. The CMS Office of Communications coordinates waiver requests for second-level .Gov domain names, such as healthcare.gov, medicare.gov, and CMS.gov. 

BR-F-19: Communication Initiated to Internet-Based Services from ATO’d Environments Must Be Allowlisted 

Communications initiated to Internet-based services must be allowlisted, either by IP address or by domain name, by a CMS-controlled security service. 

Exception: An application may be granted broader access capabilities if it uses an authenticated security service that keeps detailed usage logs. To obtain such an exception, consult with the TRB before adopting such an architecture. 

Related CMS ARS Security Controls include: AC-3(09) - Supplemental: Controlled Release, AC-4 - Information Flow Enforcement, AC-4(08) - Supplemental: Security and Privacy Policy Filters, and AC-6 - Least Privilege. 

Rationale: 

Environments covered under an ATO typically contain CMS data (not test data) and are targets for attack. To reduce the likelihood of unauthorized or unintended disclosure and exfiltration as well as limit the impact of malware, Internet endpoints for CMS applications must appear in a security service allowlist. 

An associated CMS TRA business rule, BR-SEC-FW-2, defines how interzone communications are restricted (by default). 

BR-F-20: Untrusted Services and Code from Third-Party Websites and Applications 

CMS applications should only use services from Third-Party Websites and Applications (TPWA) where the terms of service are acceptable with regard to CMS business needs, including security and privacy, data use, and service levels. 

Likewise, CMS applications should only reference or initiate the download and execution of code from TPWA providers where the terms of service are acceptable with regard to CMS business needs, including security, privacy, and data use. Examples of TPWA-provided code include, but are not limited to, scripts that run in a browser, apps that install on desktop or mobile devices, or executable software. 

Related CMS ARS Security Controls include: AC-20, Use of External Systems. 

Rationale: 

Using untrusted code or services from a TPWA may lead to potential data leakage and privacy violations. Referencing services or downloading or code from a TPWA is inviting an external source to operate in concert with applications from CMS. This is a dangerous practice because the TPWA provider may change their code at any time and without notice. This may result in malware execution because a TPWA domain name expired, the link was changed, or the code was changed without notice to CMS. 

Some TPWA providers may also embed in their code or content links to additional TPWA providers who may distribute unwanted content or malware. TPWA providers may also share collected data with additional unidentified providers who may not be bound by the TPWA’s terms of service. 

RP-F-21: Limit Data in the Application and Presentation Zones 

While application services (Application Zone) may temporarily store data for processing as described in BR-SA-5 , the data should be limited to only what is needed to complete the processing task. Avoid having full copies of data sets (i.e., files, query datasets, etc.) in the Application and Presentation Zones. 

BR-F-22: The CMS TRA Defines a Services Framework Architecture 

The Services Framework Architecture defines services based upon function, edge services, application services and data services. Mediation services are required to be applied to each service to provide compliance and security. 

The Services Framework is a service-fabric that integrates services as required with two key requirements - data should be at least three independent challenges away from the open Internet and a service should front access to the application data. 

To comply with the defense-in-depth principles, the architecture should implement a service as the authorized channel to access the application data. Users from external networks, CMSNet, or the CMS LAN must be authenticated to access the service fronting the data. Additionally, the file storage service must be configured to deny access to all unauthenticated users and all unauthorized users. 

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration, PL-8 - Information Security Architecture, SA-8 - Security Engineering Principles, SC-2 - Application Partitioning, and SC-32 - Information System Partitioning. 

Rationale: 

The evolution of cloud and modern technologies has created more dynamic environments. As a result, the Service Framework Architecture, has been developed to align Defense-in-Depth principles with new cloud architectures. 

 

CMS Technical Review Board 

This topic is based on the TRB Research Spotlight CMS TRB Engagement Guidance, Updated June 7, 2024. 

The Technical Review Board (TRB) is the CMS governing body that oversees and provides guidance on CMS’s IT investments to ensure they are consistent with the Agency’s IT strategy and architecture. The TRB provides intellectual continuity of high-level architectural decisions and direction, and promotes IT reuse, information sharing, and systems integration across the Agency. 

The TRB engages with project teams by request at key decision points throughout the CMS IT project life cycle in accordance with the project’s demonstrated compliance with 

  • The CMS Target Life Cycle (TLC) 
  • Agency enterprise architecture principles 
  • CMS policies, procedures, standards, and guidelines 

The TRB encourages developers, especially developers using the newer, iterative development methodologies (e.g., Agile and Lean), to review and plan their project implementation to minimize the impact of changes resulting from TRB guidance received during the development process. For more information about engaging with the TRB, please visit the CMS TRB SharePoint site. 

The Technical Review Board offers different engagement types reviews of any proposed system . Project teams can engage with the TRB for various purposes, facilitating project progress and adherence to CMS standards. The TRB provides technical consults on CMS projects related to Infrastructure, Software, Interface, Security, Performance, and Technology at various stages of the system lifecycle. These consultations may occur on an ad-hoc basis or follow a regular cadence, allowing for discussions with the TRB to benefit from their expertise without impeding project timelines. 

TRB Services 

The CMS Technical Review Board (TRB) is a technical assistance resource for project teams across the agency at all stages of their system’s life cycle. It offers consultations and reviews on an ongoing or one-off basis, allowing project teams to consult with a cross-functional team of technical advisors. It also guides project teams on adhering to CMS technical standards and leveraging existing technologies. 

Teams can consult regarding: 

  • ask for help with a technical problem 
  • review potential solutions or ideas with the TRB and other Subject Matter Experts (SMEs) 
  • schedule an ongoing cadence of technical consultations 
  • consult with SMEs from across the agency 
  • consult with the TRB about CMS guidelines and standards 
  • request research or information about a particular technical topic 

Requesting TRB Services 

The project team can request TRB services at any point in the life cycle of the project. The TRB CaaS and TRB Consults can all be requested by submitting a Technical Assistance request in EASi system at easi.cms.gov/trb or by reaching out via email to the TRB mailbox at CMS-TRB@cms.hhs.gov. Such requests must be initiated by a CMS Employee, preferably the business owner, project manager, or technical lead. 

Preparing for TRB Sessions and TRB CaaS Engagement 

To help facilitate the TRB Consults, TRB requires the project teams to fill in the EASi intake form with a clear business context, any relevant technical diagrams, and the areas where help is sought and upload any relevant artifacts before the scheduled meeting. 

The TRB Consults and TRB CaaS engagements are informal and do not have any prescribed templates for the presentation. The project teams are recommended to provide a document explaining the system and the changes/questions that the team is seeking guidance on, along with a background and brief description of the system and the system diagrams. Most of the technical discussion centers around the system architecture/design diagrams, and the project team should prepare system architecture/design diagrams that depict the current and proposed changes with all the system components. Clear and complete diagrams will lead to a more meaningful engagement. 

The TRB-published review templates are located on the CMS Enterprise website at: 
TRB Architectural Diagrams Sample and Diagram Requirements - Instructions 

What Project Materials are required for the TRB Engagements? 

To help facilitate the TRB Design and Operational Readiness Session engagements, TRB requires the project teams to fill in the TRB Templates noted above and send them to the TRB mailbox in advance of the scheduled meeting. 

TRB Consults and TRB CaaS engagements being informal, do not have any prescribed templates for the presentation. The project teams are recommended to provide a document that explains the system and the changes/questions that the team is seeking guidance on, along with a background and brief description of the system and the system diagrams. A majority of the technical discussion centers around the system architecture/design diagrams and the project team should prepare system architecture/design diagrams that depict the current and proposed changes with all the system components. Clear and complete diagrams will lead to a more meaningful engagement. 

Who Should Attend the TRB Engagements? 

The following members should attend: 

  1. A CMS employee with oversight of the project, such as CMS’ Contracting Officer Representative (COR) and/or Government Task Lead (GTL)/ CMS Project Manager, must be present at the sessions. (Note: The TRB cannot meet without one of these individuals). 
  2. CMS Business Owner Representative. 
  3. Information Systems Security Officer (ISSO). 
  4. Contractor Project Manager (PM) responsible for the overall effort 
  5. Lead Business Analyst who can address the business purpose, context, and process. 
  6. Lead Technical or Solution Architect who can address the architecture and design. 
  7. Development lead who can address the development process, DevSecOps, and overall testing. 

What do I need to do on the day of my Scheduled TRB Session? 

TRB Session meetings are conducted at CMS’ Central Office, and zoom sessions are always available. In the case when the CMS Central Office is closed, TRB Session meetings will be held entirely virtually. To attend the meeting virtually, please click on the TRB Meeting link that is sent in the meeting invite. 

We recommend that you login a few minutes early to ensure that you can share your screen and display your presentation properly. 

Please make sure to input your full name when logging into Zoom for the TRB Meeting. Due to the sensitive nature of the presentations, we require that all participants identify themselves on Zoom. Unidentified participants will be removed from the call. 

What to expect after the TRB Session? 

The project teams will receive a TRB Advice Letter within five business days for any Design/Operational Readiness/Consults. TRB Consult as a Service engagement meetings run by the TRB will be provided with a response in the form of meeting notes within a couple of days of the engagement. For continuous ongoing TRB CaaS engagement, the TRB notes might be sent out based on a need-by-need basis as determined by the TRB. 

References 

 

TRA Architecture Change Request (ACR) Process 

The ACR Process establishes the standard CMS process for request, review, approval, and publication of changes to the CMS TRA. The ACR Process consists of the major activities as shown in Architecture Change Request Process and below. The CMS TRA is accessible via the TRA websites. The complete version of the TRA is provided via an internal TRA website tra.cloud.cms.gov that is accessible to CMS employees and contractors with CMS network access. A public version is available to anyone at https://security.cms.gov/policy-guidance/cms-technical-reference-architecture. The public version has CMS sensitive information redacted. Information about the CMS Target Life Cycle (TLC), often referenced in conjunction with the CMS TRA, is accessible on the cms.gov public site.The OIT ISPG site CyberGeek is also available to the public. 

TRA Release Process (page 8) 

Request for Change 

Architecture Change Requests are now integrated with the overall TRA Release Process. Users may submit change requests through the TRA website, using the Topic Feedback feature, or optionally the Review site’s Annotator Feature. In either case, this request will create a Jira ticket to track it. The process below remains available. 

CMS employees submit requests for changes to the CMS TRA using an Architecture Change Request (ACR) Form that is delivered to the Division of IT Investment Management and Policy (DIIMP) via an email to the CMS IT_Governance@cms.hhs.gov mailbox. CMS Contractor partners and/or other employees who identify the need for an Architectural Change should request their Government Task Lead or other CMS employee submit a change request on the contractor’s behalf. 

Review and Prioritization 

The CEA and TRB reviews submitted ACRs to determine how they will be addressed and either closes the request or schedules the change request for processing. 

Reasons for closing a change request at this point include, but are not limited to: 

  • Duplicate requests 
  • Out-of-scope activities (e.g., production implementation steps) 
  • Lack of information to support any actions taken on behalf of the request 

Proposed Updates to the TRA 

Based on input from the CEA, the TRA Release Coordinator works with subject matter experts (SME) to identify changes to the CMS TRA to incorporate into an upcoming release. The TRA Release Coordinator ensures that all selected ACRs are addressed with the stakeholders and SMEs and coordinates the appropriate changes across all impacted parts of the CMS TRA. 

TRA Release Review – Initial TRB Consult 

When the proposed updates are completed, the TRA Release Coordinator submits the revised TRA to the TRB for initial review and schedules a TRB consult to discuss the proposed changes. The TRA Release Coordinator then updates the release based on TRB Consult comments and feedback received during the review period. Some comments may be addressed by creating a new ACR to be assigned to a future release. 

TRA Release Review – Group Review 

The TRA Release Coordinator submits the release for review by CMS Group Directors, CMS Business Units, and CMS data center operators for review and comment. The TRA Release Coordinator then updates the release based on comments and feedback received during the review period. Some comments may be addressed by creating a new ACR to be assigned to a future release. 

CIO / CEA Review and Signature 

After the TRB and Group reviews are completed and all comments are addressed, the TRA Release Coordinator submits the release to the CIO and CEA for review, approval, and signature. If the CIO or CEA have any comments, the TRA Release Coordinator addresses the comments and resubmits the release for final review and approval. After all CIO / CEA comments have been addressed, the CIO and CEA sign the release. 

Release Publication 

After the CIO and CEA sign the release, the TRA Release Coordinator publishes the updated CMS TRA. As prescribed in BR-F-2, The CMS TRA Applies to All CMS Processing Environments, the latest published CMS TRA applies immediately for new systems; for existing systems, compliance with CMS TRA updates is required within twenty-four (24) months of publication. 

 

CYBERSECURITY

Security Services

Security Services Introduction

This Security Services Chapter was developed to be compliant with the CMS Information System Security and Privacy Policy; to complement the CMS TRA by providing detailed engineering guidance for security services (e.g., firewall, intrusion detection system/intrusion detection and prevention [IDS/IDP]) to CMS Processing Environments hosting CMS applications; and to communicate the engineering decisions determined by CMS/Contractor partnerships to date. The CCIC Integration topics provide additional security services requirements for CMS Processing Environments in the areas of security tools, system and device logs integration, incident response, continuous monitoring, and penetration testing.

For the purposes of this chapter, CMS defines a security service as a virtual or physical appliance (device or application) used for security management purposes such as firewalls, switches, routers, and network-based intrusion detection systems (NIDS). In some cases, the security services are supplemented by agents installed on CMS systems.

Relevant Documents

This chapter is not all-inclusive; it complements and incorporates CMS’s existing policies, standards, and procedures, thereby offering an architectural view of the standards. Should any conflict occur, the following standards and any successor documents shall take precedence:

Except as noted, the foregoing documents may be downloaded from the CMS Information Security and Privacy Libraryand the Cybergeek CMS Policies and Guidance website.

 

General Security and Privacy Requirements

The CMS TRA uses the term “CMS Sensitive Information” as defined in the NIST Computer Security Resource Center (CSRC) Glossary, and subject to Executive Order 13556 -- Controlled Unclassified Information. See also in CyberGeek, What is considered “sensitive information”. This definition includes all data that require protection due to the risk and magnitude of loss or harm, such as Personally Identifiable Information (PII), Protected Health Information (PHI), and Federal Tax Information (FTI).

The cybersecurity and privacy requirements in this chapter support the engineering details necessary to implement, manage, and audit all infrastructure components that provide information security for the CMS Processing Environments. Deviations to this technical architecture require explicit approval from the CMS Technical Review Board (TRB) and must be documented as part of the System Security Plan (SSP) and Information Security Risk Assessment (ISRA).

Furthermore, if an application developer/maintainer cannot comply directly with this technical architecture, the project must implement compensating information security and privacy controls and clearly document reductions in risk due to the compensating controls in a proposed risk-based decision for acceptance/rejection by the Authorizing Official (AO) for the application. Information System Security Officers (ISSO) must also conduct a Privacy Impact Assessment (PIA) and a Security Impact Analysis (SIA) for any changes, as required by CMS ARS Security Control CM-4.

New systems and systems with significant changes must obtain an ATO before commencing production operations. Business owners must arrange for testing and submit a complete ATO package to the Information Security and Privacy Group (ISPG) in accordance with current procedures.

CMS TRA Multi-Zone Architecture Security Requirements

The multi-zone architecture supporting the CMS Processing Environments separates each zone by firewalls (virtual or physical) to promote application security. The Foundation section of the TRA (see CMS Services Framework) details how the Multi-Zone Architecture, built upon a services framework, supports CMS security requirements. The services framework defines edge, application, data, management and security services. These services are mapped to an analogous zone (for example application services are mapped to the application zone), except for edge services which are mapped to the presentation zone. The zones generally have clearly defined boundaries when implemented at a CMS data center. The zones may not be so clearly defined when implemented within the cloud. Note that in the topics that follow, the word “server” is used in its most generic context. Servers in this context can refer to physical servers, virtual servers, or even services provided by a cloud service provider. The outermost zone — the Presentation Zone — supports web servers. The middle zone — the Application Zone — supports business logic for the applications. The innermost zone — the Data Zone—contains the data storage, database servers, and data services used by the applications. The Management Zone supports typical systems management functions and network services that include management functions such as Public Key Infrastructure (PKI), Domain Name System (DNS), system backup, patch management, and security monitoring. To maintain separation of management and security activities from those activities inherent to an application, the Management and Security Zones are configured using separate network segments connected to secondary network interfaces of the zone devices and the corresponding devices in the Management and Security Zones. CMS Multi-Zone Architecture illustrates the CMS TRA Multi-Zone Architecture.

In CMS data centers, all inbound and outbound traffic to each zone routes through highly available switches and firewalls to reduce the possibility of compromising the entire network and is protected by highly available IDSs. These network-centric devices must be implemented as dedicated appliances or virtual instances implemented on non-general purpose operating systems. Physical network segmentation (for on premise) or virtual network segmentation (for cloud) must separate all TRA zones.

CMS TRA Multi-Zone Architecture

Presentation Zone

The Presentation Zone contains the servers supporting the static web page content for CMS applications. No application or database logic processing will be executed on servers in the Presentation Zone. The Presentation Zone will proxy communication requests to the Application Zone (i.e., no application-related connection will be allowed into the Application Zone that has not first originated in the Presentation Zone). The Presentation Zone is expected to perform application firewall (e.g., layer 7) analysis and content inspection processing. The purpose of this is to verify the integrity of content. Services in the Presentation Zone must not write to disk any data submitted by a client or any sensitive information passed to the client.

In CMS data centers, all inbound and outbound traffic to the Presentation Zone must be protected by highly available, CMS-dedicated firewalls, and network- and host-based IDS/IDPs. The firewalls and IDS/IDPs provide a concentration of security services for the Presentation Zone that include, at a minimum, packet and protocol filtering, packet inspection, information hiding, and audit logging consistent with CCIC requirements, and are required in all CMS processing environments. All host names and network IP addresses behind the Internet-facing border firewall must be masked to the external network.

Application Zone

Application servers (including web-based applications) are located in the Application Zone. As required, this zone can host services to support any non-web-based application need. For example, many applications need to exchange files with other CMS business partners. In those cases, the server providing file transfer functions must reside in the Application Zone and must be proxied by file transfer services in the Presentation Zone.

The Application Zone can also host proxy services to the Data Zone. This occurs when LDAP is used for authentication services. The LDAP store must reside in the Data Zone and an LDAP proxy will be hosted in the Application Zone, thus providing the authentication service to devices that reside in either the Presentation or Application Zones. When resources located in the Data Zone require LDAP services, an LDAP proxy will be required in the Data Zone (i.e., all access to the LDAP data store must be via an LDAP proxy). System requirements vary widely based on business and technology needs; therefore, the foregoing example shows how proxy services may supplement system design to maintain the zone architectures.

All application-related connections inbound to the Application Zone must originate from the Presentation Zone. Inbound connections from the Data Zone (to the Application Zone) will be allowed to support approved application messaging protocols. In CMS data centers, all inbound and outbound traffic to the Application Zone must be protected by highly available, CMS-dedicated firewalls and network- and host-based IDS/IDPs. The dedicated firewalls and IDS/IDPs provide a concentration of security services for the Application Zone that include, at a minimum, packet and protocol filtering, packet inspection, information hiding, and audit logging consistent with CCIC requirements, and are required in all CMS processing environments.

When implementing virtual firewalls, stateful inspection firewalls must be on a dedicated Virtual Machine (VM) to ensure (a) restriction of access controls to the security administrators and (b) proper segregation and maintenance of duties. A single, firewall-dedicated VM may be used to host more than one firewall instance.

Data Zone

The Data Zone contains all data sources, including data stores supporting directory services / authentication functions, data warehouses, and data marts as well as the applications’ operational databases. All interactions to and from the Application Zone and the Data Zone must pass via a commonly-used access method. Examples can be found in Data Services Access Methods. CMS permits the use of database-stored procedures that may reside on database servers in the Data Zone.

In CMS data centers, all inbound and outbound traffic to the Data Zone must be protected by highly available, CMS-dedicated firewalls and network- and host-based IDS/IDPs. The dedicated firewalls and IDS/IDPs provide a concentration of security services for the Data Zone that include, at a minimum, packet and protocol filtering, packet inspection, information hiding, and audit logging consistent with CCIC requirements, and are required in all CMS processing environments. All non-Management Zone connections inbound to the Data Zone must originate from the Application Zone and use a TRB approved protocol.

Programs in the Data Zone may not originate any communications to endpoints outside the Data Zone. Data Zone-to-Data Zone communications are permitted.

WAN Connections

WAN connections between like zones in different data centers must maintain the same level of Defense-in-Depth protections. Various types of zone-to-zone connections are documented in the Wide Area Network Services chapter.

WAN connections must:

  • Utilize mutual authentication and encrypted tunnels to secure zone-to-zone connections
  • Include centrally managed, stateful inspection firewall, IDS, and IDP at both ends of all tunnels

Please refer to Wide Area Network Services for more information and requirements concerning Wide Area Networking.

Management and Security Zones

The CMS TRA requires isolation of Management and Security Services on dedicated network zones. Remote access into Management and Security Zones is restricted to Internet Protocol Security (IPSec) VPNs utilizing multi-factor authentication.

Each host in a zone must have separate network interfaces for the Security and Management Zones. Network traffic in the Management and Security Zones must use separate Virtual LANs (VLAN) on the respective Management or Security network interfaces. The CCIC Integration chapter identifies mandatory security tools and security functions applicable to CMS data centers.

Workstations in the Management and Security Zones must be configured to assure the integrity and protection of the various zones. To ensure rigorous anti-virus protection mechanisms, CMS requires sandboxing technologies and/or proxy services between the Management and Security Zones and the Presentation, Application, and Data Zones.

Workstations in the Management and Security Zones must be isolated from all networks other than the zone(s) that they manage.

In CMS data centers, all inbound and outbound traffic to the Management and Security Zones must be protected by highly available, CMS-dedicated firewalls and network- and host-based IDS/IDPs. The dedicated firewalls and IDS/IDPs provide a concentration of security services that include, at a minimum, packet and protocol filtering, packet inspection, information hiding, and audit logging consistent with CCIC requirements, and must be implemented in all CMS processing environments.

Audit Log Generation and Collection Requirements

The CMS ARS requires the use of security technologies to monitor potential system intrusions. Security documentation must define all events that require investigation and the associated response times. Associated security tools must identify how each security event has been investigated and what associated response actions have been taken. Therefore, all devices must be configured to create and retain log information to support security monitoring functions and/or forensic investigations. The CCIC Integration chapter contains a complete list of system and device logs that must be collected.

The CMS ARS Security Controls AU-8, SC-45 and SC-45(1) define the requirements for time synchronization.

To prevent unauthorized changes to logged information and limit interactive access to CMS Processing Environments as well as allow cross-component log correlation, CMS requires support for centralized log analysis. All devices capable of generating log information — including, but not limited to, host-based IDSs (HIDS), network-based IDS/IDPs, firewalls, operating systems, middleware, databases, and applications — must be centrally managed and engineered to either send log information to the Security Zone directly or to allow the logs to be pulled into the Security Zone. Both real-time and batch delivery are acceptable. (The CMS ARS provides detailed requirements for logging data and collection frequency.) Logs from all components in all zones must be aggregated in the Security Zone and analyzed by the local system developer and maintainer (SDM) Security Operations Center (SOC), and in turn integrated with the CCIC (as described below in Security Architecture and Engineering).

Trusted Internet Connections

Office of Management and Budget (OMB) Memorandum M-08-05, Implementation of Trusted Internet Connections, November 20, 2007, states that all federal agencies must optimize and standardize the security of individual external connections. In addition, security controls must be implemented within all federal network operating environments.

Use of a TIC is also required by CMS ARS Security Control SC-07(3) - Access Points (High, Moderate) (CSP: Mod) P1, which provides, “Limit the number of external network connections to the HVA;” and further provides as Implementation Standards for High & Moderate: Std.1 – “Implementation must route external connections via a Trusted Internet Connection[JD7]  (TIC) portal.”

The mission of the Trusted Internet Connections program is to provide a means for monitoring, isolating, and securing federal external network connections in the event of cyberattacks. The program requires the Agency to route all traffic to and from external networks through TIC Zones managed by TICAPs operated by the Agency or third parties, such as a Managed Trusted Internet Provider Service (MTIPS). The TIC Zone has the following features:

  • External connection termination point
  • Monitored by EINSTEIN
  • Intrusion Detection System for protection of federal information technology (IT)
  • Capable of alerting the US-CERT (United States Computer Emergency Readiness Team) to the presence of malicious or potentially harmful computer network activity
  • Network connections and data filtering
  • Full packet capture and storage

CMS follows a TIC deployment strategy aligned with the Department. HHS provides TIC services for all HHS Operating Divisions, including CMS data centers and applications. The National Institutes of Health (NIH) operates the HHS TIC services.

CMS application owners and data center operators are responsible for extranet connections to external networks, and for working with the HHS Office of Information Security (OIS) and NIH TIC teams to ensure the extranet connections are properly implemented and extranet traffic is monitored through HHS TIC services. This includes extranet connections to cloud providers and to other government networks.

Malicious Code Protection

Malicious code can be introduced into a system by many different means, including:

  • Web accesses, electronic mail and attachments, and portable storage devices
  • Exploitation of information system vulnerabilities
  • Installation or execution of custom-built software or Commercial Off-the-Shelf (COTS) software

Malicious code protection mechanisms must be employed at information system entry and exit points to detect and eradicate malicious code. Information system entry and exit points include, for example, firewalls, electronic mail servers, web servers, proxy servers, remote access servers, workstations, notebook computers, and mobile devices.

Malicious code protection mechanisms may employ a variety of technologies and methods to limit or eliminate the effects of malicious code, including for example, anti-virus signature definitions and reputation-based technologies. Pervasive configuration management and comprehensive software integrity controls may be effective in preventing execution of unauthorized code. Other safeguards include secure coding practices, configuration management and control, procurement from trusted sources using secure delivery methods, and monitoring practices to help ensure that software does not perform functions other than the functions intended.

Malicious code protection mechanisms must be updated whenever new releases are available in accordance with CMS Risk Management Handbook: Configuration Management (CM)procedures. The minimum frequencies for updating are defined by these CMS ARS controls:

  • Malware, e.g., Anti-Virus (AV): SI-3 - Malicious Code Protection
  • Vulnerability scan: RA-5 - Vulnerability Monitoring and Scanning
  • Software Integrity check: SI-7(1) - Integrity Checks
  • Boundary/IDS: SI-4(2) - Automated Tools and Mechanisms for Real-Time Analysis
  • Vulnerability scanning: RA-5(2) - Update Vulnerabilities to be Scanned

Please refer to CMS ARS Security Control SI-03 for additional guidance on malicious code protection. Please also refer to BR-CCIC-23, Network Security Endpoint Protection Capability,in the CCIC Integration chapter.

Scanning

Malicious code protection mechanisms that perform scanning of files, systems, or networks must be configured to:

  1. Perform periodic scans of the information system in accordance with the specific type and frequency prescribed in the CMS ARS.
  2. Perform real-time scans of files or content from external sources at endpoint, and/or network entry/exit points, as the files are downloaded, opened, or executed.
  3. Block and quarantine malicious code and send alerts to the administrator in response to malicious code detection.

Scanning Submitted Files

Before applying any operational function to the contents of a submitted file, it must be scanned for malware. Contents of newly submitted files must be scanned in the Presentation or Application Zone before transfer to the Data Zone.

Scanning Information Systems

On servers, malicious code scanning services must be configured to perform periodic scans of critical system files and periodic full system scans as specified in the CMS ARS or by FedRAMP (Federal Risk and Authorization Management Program) requirements.

For cloud service providers, the organization must configure malicious code protection mechanisms to:

  • Perform periodic scans of the information system and real-time scans of files from external sources as the files are downloaded, opened, or executed in accordance with organizational security policy; and
  • Block or quarantine malicious code, send alerts to the administrator, and send alerts to FedRAMP in response to malicious code detection.

Malicious code scanning results are reported to the CMS Cybersecurity Integration Center Security Information and Event Management system[JD8]  in compliance with CMS ARS Security Control AU-06.

Detection / Incident Response

Organizations may determine that, in response to the detection of malicious code, different actions may be warranted. For example, organizations can define actions in response to malicious code detection during periodic scans, actions in response to detection of malicious downloads, and/or actions in response to detection of maliciousness when attempting to open or execute files.

Malicious code must be reported to the local SOC and to the CCIC. As part of an incident management plan, organizations must be prepared to take immediate or automated action to mitigate damage from malicious code when it is detected, as well as comply with requirements and incident management procedures described in the CCIC Integration chapter.

The organization must also address the receipt of false positives during malicious code detection and eradication and the resulting potential impact on the availability of the information system.

CMS Continuous Monitoring Requirements

CMS Continuous Monitoring Requirements are detailed in the CCIC Integration chapter.

 

Business Rules

The following business rule defines general security policies within the CMS environments.

BR-SEC-Gen-1: Traffic between and within Zones Must Be Available in Unencrypted Form for Security Analysis Purposes

When an application’s communications between and within zones is encrypted as required by the CMS ARS, the environment must use mechanisms capable of either decrypting content for security analysis or analyzing security of the content before transmission or after receipt.

There are specific instances where encrypted traffic must remain encrypted and not be made available in an unencrypted form for security analysis purposes. Examples include, but are not limited to, LDAP authentication and administrative sessions, such as Secure Shell (SSH) or Remote Desktop.

If a system developer and maintainer cannot comply directly with this business rule, the SDM must implement compensating information security and privacy controls and clearly document reductions in risk due to the compensating controls in a proposed risk-based decision for acceptance/rejection by the Authorizing Official for the environment.

Also note FedRAMP Guidance for M-21-31 and M-22-09, on Alternatives to network inspection. These bullets are extracted from OMB M-22-09

  • Current OMB policies neither require nor prohibit inline decryption of enterprise network traffic. Agencies are expected to balance the depth of visibility they need with the risks presented by broadly trusted network inspection devices.
  • Network traffic that is not decrypted can and should still be analyzed using visible or logged metadata, machine learning techniques, and other heuristics for detecting anomalous activity. This is consistent with the Trusted Internet Connection (TIC) initiative, as updated in OMB Memorandum M-19-26, which gives agencies the flexibility to maintain appropriate visibility without needing to perform inline traffic decryption.
  • OMB Memorandum M-21-31, “Improving the Federal Government’s Investigative and Remediation Capabilities Related to Cybersecurity Incidents”, describes required fields that agencies must log consistently throughout their enterprise, including packet capture logs. M-21-31 does not require full traffic inspection, but specifies fields that should be captured when such inspection is in place. M-21-31 describes this conditional

Related: CMS ARS Security Controls include: SI-4 - Security Monitoring and SC-8 - Transmission Confidentiality and Integrity.

Rationale:

All application traffic between and within zones must be available in an unencrypted form to permit security services and authorized security personnel to perform auditing and monitoring of communications.

Secure Baseline Configuration

BR-SEC-Gen-2: Software and Hardware Components Must Adhere to a Secure Baseline Configuration

All configurable software and hardware components must adhere to a secure baseline configuration authorized to operate by the CMS Chief Information Officer (CIO). All operating systems, network devices, databases, middleware, or other equipment and software must be configured in accordance with CMS Configuration Management (CM) and the CMS ARS, CM-6 - Configuration Settings and CM-7 - Least Functionality. These configurations must be defined in the applicable host environment System Security Plan and must be authorized to operate by the CMS CIO. Note: CMS ARS requirement CM-6 addresses both standard configuration definition and deviations from a standard configuration.

The business owner and SDM must actively manage configurations. Configurations must be validated periodically using the tools hosted in the Security Zone , such as Security Content Automation Protocol (SCAP)-based testing tools and Defense Information Systems Agency (DISA) scripts (STIGs).

Related: Configuration Management; CMS ARS Security Controls include: CM-6 - Configuration Settings and CM-7 - Least Functionality.

Rationale:

All applications and devices must adhere to a secure baseline configuration to provide protection against unauthorized access and malicious attacks, and to permit security services and authorized security personnel to perform auditing and monitoring of communications.

BR-SEC-Gen-3: Disable All Unnecessary Features and Capabilities

All configurable software and hardware components must disable all unnecessary features and capabilities. In addition to implementing the CMS Secure Baselines, any configurable component in the CMS Processing Environments, such as web servers, application servers, other servers, firewalls, or routers, must disable all unnecessary features and capabilities, Internet services, protocols, and applications not expressly required by a CMS application. All compilers, editors, database templates, development tools, unnecessary interpreters, etc. must be removed from the CMS Processing Environments.

Related: Configuration Management; CMS ARS Security Controls include: CM-6 - Configuration Settings and CM-7 - Least Functionality. CMS ARS Security Controls include: CM-6 - Configuration Settings and CM-7 - Least Functionality.

Rationale:

Eliminating unnecessary features, access points, protocols, and developer tools reduces the number of potential attack vectors, as well as possible points of failure.

Internet Connectivity

BR-SEC-Int-3: HTTPS on CMS Public-Facing Websites and Services on the Internet

OMB Memorandum M-15-13, Policy to Require Secure Connections across Federal Websites and Web Services, June 8, 2015, requires that all publicly accessible federal websites and web services only provide service through a secure connection. OMB Memorandum M-23-22, Delivering a Digital-First Public Experience, September 23, 2023, expands on this.

Publicly accessible websites and services are defined in M-15-13 as online resources and services available over HTTP or HTTPS over the public internet that are maintained in whole or in part by the federal government and operated by an agency, contractor, or other organization on behalf of the agency. They present government information or provide services to the public or a specific user group and support the performance of an agency’s mission. This definition includes all web interactions, whether a visitor is logged in or anonymous.

HTTPS is a combination of HTTP and Transport Layer Security (TLS). TLS is a network protocol that establishes an encrypted connection to an authenticated peer over an untrusted network. Only current versions of TLS are to be allowed for connections to CMS public facing websites. See NIST SP 800-52 Rev 2 Guidelines for the Selection, Configuration, and Use of Transport Layer Security (TLS) Implementations for additional information on TLS versions and requirements.

For websites:

To meet the M-15-13 requirement of enforcing HTTPS, the affected CMS websites:

  • Employ server-side redirects. Using only client-side redirects, such as a <meta refresh> tag or JavaScript, is not considered adequate.
  • Employ HTTP Strict Transport Security (HSTS) and HSTS preloading.
  • Allowing HTTP connections for the sole purpose of redirecting clients to HTTPS connections is acceptable and encouraged.
  • HSTS headers must specify a max-age of at least 1 year.
  • The parent domain and each of the publicly reachable subdomains must set a HSTS policy with a max-age of at least 1 year.

Preloading marks entire domains as HTTPS-only and allows browsers to enforce this rigorously and automatically for every subdomain. HSTS preloading a parent domain allows agencies to avoid inventorying and configuring an HSTS policy for every individual subdomain; however, this approach also automatically includes all subdomains present on this domain, including intranet subdomains. All subdomains must support HTTPS to remain reachable for use in browsers supporting HSTS. All .gov domains registered after September 1, 2020 are preloaded (Source: An intent to preload).

For API-based services:

CMS public Internet-facing application programming interface (API)-based services must support only HTTPS. Exceptions must be documented in the system design documents (SDD), CMS FISMA Controls Tracking System (CFACTS), and risk assessment for both the system and the environment. If HTTP is in use for such an existing service, a migration plan should be developed. Helpful information on migration planning can be found at www.councils.gov/cioc.

References: OMB M-15-13 and M-23-22. Please refer also to The HTTPS-Only Standard

Related CMS ARS Security Controls include: SC-8 - Transmission Confidentiality and Integrity, AC-14 - Permitted Actions Without Identification or Authentication, AC-22 - Publicly Accessible Content, and PL-4 - Rules of Behavior.

Rationale:

M-15-13 states that all browsing activity should be considered private and sensitive. The HTTPS‑Only standard will eliminate inconsistent, subjective determinations across agencies regarding which content or browsing activity is sensitive in nature and will create a stronger privacy standard government wide. Federal websites that do not convert to HTTPS will not keep pace with privacy and security practices used by commercial organizations and with current and upcoming Internet standards. This failure leaves Americans vulnerable to known threats and may reduce their confidence in their government.

Administrative Access

BR-SEC-Gen-4: Administrative Access to CMS Services and Devices

CMS applications and devices must be configured to prevent the operation of all system administrative functions except those that originate from the Management and Security Zones. When necessary, application administrative functions can be accessed via the Application Zone by defining application administrative roles and documenting the associated risks and compensating controls in the SSP and ISRA.

Where possible, sessions from the Management Zone must not be interactive except where production emergencies dictate such activities or where there are technology limitations. For example, system changes must be tested in a validation environment and required changes pushed into production via automated tools. Where interactive sessions are required, all activity must be securely logged and security management must review the activity to ensure proper separation of duties. For additional information on application administrative access, please refer to the Access Control and Identity Management, Privilege Administration topic.

RelatedConfiguration Management (CM); CMS ARS Security Controls include:
SC-2 - Separation of System and User Functionality, SC-08 - Security Function Isolation, and PE-02 - Physical Access Authorizations by Role.

Rationale:

Restricting network to administrative functions by network port and source reduces the number of potential attack vectors and helps ensure only authorized users and administrative systems have access to sensitive functions.

Malicious Code Protection

BR-SEC-Gen-17: Software Assurance Measures

Software assurance measures, including scanning for malware (malicious software), must be applied to all information systems in the CMS Processing Environment for new or changing code prior to gaining production status, prior to change management activities, and/or as part of routine workflows that support an application.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection and SI-10 - Information Input Validation.

Rationale:

Malware may be introduced through user input, file transfers, eXtensible Markup Language (XML) and other transactions, deployments of new or modified software, or direct access to a system. Passive and active malware scanning of files or commands introduced through all these vectors is a critical part of Defense-in-Depth.

BR-SEC-Gen-18: Malicious Code Protection in CMS Processing Environments

Malicious code protection mechanisms must be employed at information system entry and exit points to detect and eradicate malicious code. Information system entry and exit points include, for example, firewalls, electronic mail servers, web servers, proxy servers, remote access servers, workstations, notebook computers, and mobile devices. CMS TRA – Cybersecurity, Security Services, Malicious Code Protection, discusses this in more detail. CMS Hybrid Cloud describes scanning services in CMS approved enterprise security tools and services.

The CMS ARS dictates the required specific type and frequency of virus and malware scanning services by system classification (High, Moderate, and Low).

Related BR-SEC-Gen-18a; CMS ARS Security Controls include: SI-3 - Malicious Code Protection and SI-10 - Information Input Validation.

Rationale:

Malware may be introduced through user input, file transfers, XML and other transactions, deployments of new or modified software, or direct access to a system. Passive and active malware scanning of files or commands introduced through all these vectors is a critical part of Defense-in-Depth.

BR-SEC-Gen-18a: Submitted Files Must Be Scanned in the Presentation or Application Zone

Before applying any operational function to the contents of a submitted file, it must be scanned for malware. Contents of newly submitted files must be scanned in the Presentation or Application Zone before transfer to the Data Zone. Files with detected malware must be quarantined.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection and SI-10 - Information Input Validation.

Rationale:

Malware may be introduced through user input, file transfers, XML and other transactions, deployments of new or modified software, or direct access to a system. Passive and active malware scanning of files or commands introduced through all these vectors is a critical part of Defense-in-Depth. CMS Hybrid Cloud describes scanning services in CMS approved enterprise security tools and services.

BR-SEC-Gen-19: User or External Inputs Must Be Validated

All user or external inputs submitted into the CMS Processing Environments must be validated.

CMS information systems must check the validity of defined information inputs (defined in the applicable System Security Plan) for accuracy, completeness, validity, and authenticity as close to the point of origin as possible.

Checking the valid syntax and semantics of information system inputs (e.g., character set, length, numerical range, and acceptable values) verifies that inputs match specified definitions for format and content.

Related CMS ARS Security Controls include: SI-10 - Information Input Validation.

Rationale:

Software applications typically follow well-defined protocols that use structured messages (i.e., commands or queries) to communicate between software modules or system components. Structured messages can contain raw or unstructured data interspersed with metadata or control information.

If software applications use attacker-supplied inputs to construct structured messages without properly encoding such messages, then the attacker could insert malicious commands or special characters that can cause the data to be interpreted as control information or metadata. Consequently, the module or component that receives the tainted output will perform the wrong operations or otherwise interpret the data incorrectly.

Prescreening inputs prior to passing to interpreters prevents the content from being unintentionally interpreted as commands. Input validation helps to ensure accurate and correct inputs and prevent attacks such as cross-site scripting and a variety of injection attacks.

BR-SEC-Gen-21: Malware and Malicious Code Scanning Results Must Be Sent to the Security Zone

Malware and malicious code scanning results must be sent to the Security Zone for further analysis/investigation. Note that cloud implementations may utilize shared storage (e.g., AWS S3) as the security zone for result storage and security scanning. Files with detected malware must be quarantined.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection, SC-03 - Security Function Isolation, and enhancements under SC-7 (i.e., SC-7(13) - Supplemental: Isolation of Security Tools, Mechanisms, and Support Components) for CSP.

Rationale:

Monitoring and analyzing malicious code scanning results and malware is a function of the Security Zone .

BR-SEC-Gen-20: XML Firewalls Must Be Used to Authenticate XML Exchanges

When a system supports XML exchanges, corresponding protections against malicious attacks must be implemented such as XML firewalls to authenticate the XML and validate the security of the content. XML security devices are typically deployed in the Presentation Zone; validated code is passed back to an application residing in the Application Zone.

With the convergence of network-based security devices, it would be unrealistic to propose a single solution for the prevention of XML-based malicious attacks; XML protections can be found and leveraged in existing firewalls and XML gateways or through XML firewalls.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection and SI-10 - Information Input Validation.

Rationale:

System-to-system data exchanges using XML can also be the subject of malicious attacks. Employing separate physical or virtual protection devices reduces the possibility that the protection mechanism becomes a target of malicious attacks.

BR-SEC-Gen-22: All Information Systems Must Have a System Risk Assessment in CFACTS

The system risk assessment must document any malware and malicious code vulnerabilities and suitable plans of action and milestones established in CFACTS. When these vulnerabilities extend to any CMS Processing Environments or other systems, the related ISSOs and system maintainers must also be notified to include the risk in their risk assessments.

Related: Configuration Management (CM); CMS ARS Security Controls include: SI-3 - Malicious Code Protection and PM-4 - Plan of Action and Milestones Process. CMS ARS Security Controls include: SI-3 - Malicious Code Protection and PM-4 - Plan of Action and Milestones Process.

Rationale:

CFACTS is CMS’s primary tool for managing system risk assessments for all CMS information systems.

Vulnerability Management

BR-SEC-Gen-7: Vulnerability Management for CMS Networks, Services, and Devices

Security processes and tools must be defined within the CMS Enterprise Vulnerability Program. Vulnerability identification tools must reside in the Security Zone, or as close to the scanned assets as possible, to reduce security risks and reduce scanning through firewalls. An inventory of all system components must be documented and all associated processes for monitoring vulnerability alerts must be followed. Each vulnerability alert must be documented and the associated action taken and recorded.

To achieve a CMS enterprise-wide vulnerability management and monitoring capability, the infrastructure will allow continuous collection of all vulnerability data at a single location within CMS for centralized analysis and reporting.

The goal of the CMS Enterprise Vulnerability Program is to maintain the capability to collect such data on all CMS IT assets.

Related: Configuration Management (CM); CMS ARS Security Controls include: RA-5 - Vulnerability Monitoring and Scanning and CM-8 - Information System Component Inventory.

Other assessments that fall under other CMS ARS controls can look for vulnerability as well, including for example, pen testing (CA-8); special assessments, such as Risk and Vulnerability Assessments (RVA) under CA-2 (Control Assessments) and RA-3 (Risk Assessment); privacy (RA-8 - Privacy Impact Assessments); and continuous monitoring (CM-7).

Rationale:

Having a vulnerability monitoring plan with inventory and prepared alert actions is prerequisite for CMS’s ability to continuously monitor vulnerabilities and respond quickly to alerts. Documenting alerts and actions taken when an incident occurs is also necessary for determining impacts, determining if additional steps are needed, and supporting further analysis.

Vulnerability Scanning

BR-SEC-Gen-10: Periodic and Continuous Network Vulnerability Scanning Is Required

Periodic and continuous network scanning for security-related vulnerabilities must be performed on all CMS networks.

In addition to periodic network scanning to identify security-related vulnerabilities at every CMS site, CMS requires a continuous CMS enterprise-wide vulnerability scanning capability that must include:

  • Establishing a baseline for CMS assets of known and accepted vulnerabilities at each site as well as Agency-wide
  • Scanning and producing periodic reports as defined within the CCIC Integration chapter and in accordance with CMS ARS requirements
  • Defining a remediation approach, including a prioritization based on score improvement impact at the site and across CMS

Related: Configuration Management (CM); CMS ARS Security Controls include:
RA-5 - Vulnerability Monitoring and Scanning.

Rationale:

Network scanning can uncover configuration vulnerabilities and unauthorized hosts or services. This may include vulnerabilities that may not have been known earlier when configurations were first developed. Analysis and remediation of vulnerabilities identified through network scanning aids in continuous improvement of CMS’s security standing.

Configuration Compliance Management

BR-SEC-Gen-11: Periodic and Continuous Configuration Compliance Scanning Is Required

Periodic and continuous scanning for configuration compliance must be performed on all CMS networks. Security processes and tools must be defined within a configuration compliance management program. The supporting tools must reside in the Security Zone to reduce errors in device configurations. CMS requires a documented inventory of all system components and adherence to all associated processes for monitoring configuration compliance. Each instance of non-compliance must be documented and the corrective action taken and recorded.

CMS requires a CMS enterprise-wide configuration scanning capability (requirements available through the Department of Homeland Security (DHS) Continuous Diagnostics and Mitigation (CDM) Program) that must include:

  • Establishing a baseline for CMS assets and configuration baselines for these assets at each site as well as Agency-wide. Baseline configurations must be compliant with CMS ARS CM-6 and documented in the SSP.
  • Scanning and producing periodic reports as defined within the CCIC Integration chapter
  • Defining a remediation approach, including a prioritization based on score improvement impact at the site and across CMS
  • Establishing and periodically testing baselines of approved protocols/services to guard against the introduction of unauthorized configurations within the environment
  • Using inventory management tools (requirements available through the DHS CDM Program) to periodically scan the environment to ensure it contains only approved assets

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-06 - Configuration Settings and CM-08 - Information System Component Inventory.

Rationale:

Configuration scanning can uncover configuration vulnerabilities and unauthorized hosts or services. Analysis and remediation of vulnerabilities identified through network scanning aids in continuous improvement of CMS’s security standing.

OMB M-19-03 brings enhanced configuration (compliance) management and monitoring controls to HVA (High Value Assets). The change affects “Planning, Budgeting, Contracting, Procurement, Lifecycle, Strategies,” including:

  • Formalized configuration management (CM-2, CM-3, CM-5, CM-6)
  • Enhance assessment, monitoring, and remediation (AU-6, CA-2, CA-7, CA-7(4), PM-14, SI-2, SI-4)
  • Continuous monitoring (CA-7)
  • Additional risk and vulnerability assessments (e.g., CA-8, RA-3, RA-5, SA-3, and SA-17)
  • Plan of Action and Milestones creation and tracking (e.g., CA-5 and PM-14)
  • Expedited remediation (SI-2)

Host Intrusion Detection

BR-SEC-Gen-12: Host Intrusion Detection Capabilities on All IT Components

A centrally managed IDS / IPS capability is required to monitor network communications on all networks and subnets of any environment requiring a CMS Authority to Operate. In addition, centrally managed IDS / IPS sensor agents are required in all systems, appliances, devices, services, and applications for which such agents are available. It is understood that agents may not be available in all cloud implementations, and implementors should work with their ISSO to address this requirement.

As defined in the CMS ARS, Host Intrusion Detection and monitoring capabilities must include capabilities to detect changes on the hosts (sometimes referred to as file integrity monitoring). System configuration file changes on those hosts monitored by a HIDS must be validated against the corresponding change management requests; files without appropriate change tickets must be investigated as potential security concerns. HIDS agents must be centrally managed and configured to send results to a server installed in the Security Zone for further analysis / investigation. Shareable storage (e.g., AWS S3) may serve as the security zone for the storage and sharing of logs. Events monitored must be documented in CFACTS under the appropriate information security and privacy control. CMS may install additional IDSs to enhance network security monitoring.

Related CMS ARS Security Controls include: SI-4 - System Monitoring.

Rationale:

Intrusion detection can uncover unauthorized changes or activities on CMS hosts.

Log Aggregation and Correlation

BR-SEC-Gen-15: Logs Must Be Securely Collected, Aggregated, and Analyzed

Logs must be securely collected, aggregated, and analyzed in the Security Zone and local SDM SOC and integrated with CCIC. Security documentation must define the events, categorization, and required responses. Each security event must be documented, including all corresponding response actions taken. The CCIC Integration chapter contains a complete list of system and device logs that must be collected.

Related CMS ARS Security Controls include: AU-2 - Event Logging, AU-4 - Audit Log Storage Capacity, AU-6 - Audit Record Review, Analysis, and Reporting, AU-7 - Audit Record Reduction and Report Generation, AU-9 - Protection of Audit Information, and AU-12 - Audit Record Generation.

Rationale:

Monitoring and analyzing logs can detect unauthorized activities, system or resource issues, or malicious attacks on CMS hosts. This ability is further enhanced when logged events from multiple related systems may be aggregated and correlated through a central log monitoring service.

Network Time Protocols

BR-SEC-Gen-23: CMS Network Time Protocol Services

The appropriate CMS Network Time Protocol services must be used across all CMS environments as outlined below.

  • The primary domain controllers for awscloud.cms.local domain synchronize time with external NTP sources (NIST time servers).
    • All Windows instances that are joined to this domain synchronize their time with the AD Domain Controller.
  • RHEL instances are configured to reach out directly to the NIST time servers.
  • AL2 instances use Amazon Time Sync Service for time synchronization.

NTP services must be used to synchronize system clocks across the CMS environment to Coordinated Universal Time (UTC). NTP services must reside in the Management Zone and be configured to peer with CMS-approved NTP time servers (see RFC 5905, Network Time Protocol Version 4: Protocol and Algorithms Specification, June 2010).

Related CMS ARS Security controls include: SC-45 - Supplemental: System Time Synchronization, SC-45(1) - Supplemental: Synchronization with Authoritative Time Source

Rationale:

Synchronization of system clocks across the environment is necessary to ensure proper analysis of audit / logging information, and to promote application data consistency and interoperability. Event correlation depends on the assumption that the timestamps of multiple events, logged in separate places, reflect the same relative timeframes.

Because the CMS enterprise environment spans multiple cloud service providers, regions, and data centers, it is essential that all systems and services synchronize to the same CMS-provided time services hierarchy.

Firewalls

The CMS TRA defines two classes of firewalls:

  • Host-based firewalls – firewalls that run locally on and protect only the hosting computer or device
  • Network firewalls – hardware or software firewall devices dedicated to providing firewall protection at the borders of one or more networks

All references in the CMS TRA to firewalls apply to both classes unless explicitly stated otherwise.

The CMS TRA also discusses “border” or “boundary” firewalls protecting CMS security boundaries or networks. Such firewalls are network firewalls, unless explicitly stated otherwise.

Although network firewalls are required at security boundaries, host-based firewalls providing endpoint protection improve the standing of CMS’s Defense-in-Depth and may be deployed as a mechanism to address certain CMS ARS controls.

Firewall Configuration

The following CMS TRA Security Firewall (SEC-FW) configuration business rules support the information security and privacy controls required by the CMS ARS (e.g., information flow, least privilege, and audit).

BR-SEC-FW-1: Separate Network Interfaces for Each Network Segment and Zone

Each network firewall in the CMS Processing Environments must be provisioned with separate interfaces dedicated to each network segment and zone with which it connects. This separate interface may be physical or logical and achieved through the implementation of routing and network rules. Some clouds (e.g. AWS) have other mechanisms, such as security groups, to separate interfaces as well. Include separate interfaces for:

  • Each zone the firewall connects
  • Each gateway or network (e.g., WAN, Internet, and other) the firewall connects
  • Automatic failover and load balancing
  • Connectivity to the Management and Security Zones

Rationale:

Network segmentation using firewalls helps preclude attacks traversing between segments.

BR-SEC-FW-2: Adhere to CMS Security Hardening Guidance

Network firewalls must be configured in accordance with the CMS policy defined under CMS ARS Security Control CM-6, Configuration Settings, which defines a prioritized list of the required standards. With the CM-6 guidance taking precedence, both network and host-based firewalls must be configured to:

  • Deny all requests, protocols, services, destination ports, and destination Internet Protocol addresses, except as expressly permitted in support of a CMS application or service. This includes disabling or preventing Proxy Address Resolution Protocol (Proxy ARP), IP Spoofing, Source Routing, Internet Control Message Protocol redirect (IPv4 and IPv6), and ICMP echoes if not needed.
  • Reject all traffic on its network interfaces that appears to come from a broadcast network, reserved network, or loop-back network, unless there is a business justification and is explicitly approved by CMS. Host-based firewalls may permit loop-back network traffic.
  • Reject all traffic that enters a given network interface of a firewall that appears to come from a network address not within the valid address range for that interface.
  • For border firewalls that connect to an external non-CMS network (i.e., Internet), the firewall must also be configured to reject all traffic at its external interface that appears to come from an internal network address.
  • Respond to denied requests by dropping the packet and sending resets.
  • Filter and allow packet inspection of all traffic through the (host-based or network) firewall.
  • Log all activities in accordance with current CMS ARS requirements.
  • Upon startup, prevent routing within the environment and prevent exposure of vulnerabilities while booting.
  • During an outage of any sort, network firewalls must revert to a configuration that denies all traffic pending the re-enablement of services by the firewall administrator.
  • Network firewalls must ensure that residual information from a previous information flow or internal firewall data is not revealed or transmitted in any way. Resources must be overwritten or cleared before they may be available for reuse.

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-6 - Configuration Settings, AC-6 - Least Privilege, and CM-7 - Least Functionality.

Rationale:

These are network firewall configuration practices in alignment with the CMS ARS and the DISA Security Technical Implementation Guide (STIG) and follow industry best practices.

BR-SEC-FW-3: External Connections to CMS Zones

No connections, services, or requests of any kind may be allowed from external hosts (i.e., external to specific Certification and Accreditation [C&A] boundaries established by CMS) into the Presentation Zone or any other CMS zone , unless the appropriate business justification is made along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, RA-3 - Risk Assessment, and PL-2 - System Security and Privacy Plans.

Rationale:

CMS does not permit unauthorized access to its systems or data. Interconnection must be formally approved and documented.

BR-SEC-FW-4: Access to Services of the Firewall

No connections, services, or requests of any kind may be allowed to the firewall itself except from the Management Zone unless the appropriate business justification is made along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

CMS permits connections from the firewall itself to other relevant CMS services, such as logging to the Security Zone , when there is appropriate business justification along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-6 - Configuration Settings, AC-6 - Least Privilege, and CM-7 - Least Functionality.

Rationale:

In the framework of firewall rules, the firewall itself may be a destination host if it supports features such as remote management or VPN gateway services. This BR is intended to block all traffic to the firewall except from the Management Zone. Therefore, unneeded firewall features—which should also be disabled—such as VPN gateways, are effectively blocked unless there is a justified business need and permitted by CMS.

BR-SEC-FW-5: Filtering Traffic between Zones

All network traffic to or from a zone (Presentation, Application, or Data) must pass through a firewall to enter or leave the zone. These firewalls may be physical or virtual, or implemented by other means (e.g., Security Groups in AWS, Network Security Groups in Microsoft Azure, etc.). Firewalls protecting a given zone must:

  • Prevent all traffic between the Presentation, Application, and Data Zones except as explicitly permitted by CMS
  • Filter and allow packet inspection of traffic between zones, and entering or exiting from external networks

No connections, services, or requests of any kind may be allowed to enter a CMS zone unless there is appropriate business justification along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

Rationale:

Filtering traffic entering each zone provides protection against unauthorized traffic.

BR-SEC-FW-6: Firewalls Transmit Logs and Notifications to the Security Zone

Firewalls will transmit audit logs, event logs, administrator notifications, and security alarms to the Security Zone. At a minimum, all firewalls and perimeter network devices must produce audit records required by the current CMS ARS; Security Controls AU-2 and AU-3 describe the minimum events to be logged.

Rationale:

Monitoring and analyzing logs and notifications is a function of the Security Zone .

BR-SEC-FW-9: Network Traffic Entering a Zone Must Terminate in That Zone

All network traffic entering a zone through a network firewall must terminate in that zone. The destination within the zone must be a service that receives and processes or transforms the application data carried within the network traffic. The firewall must block any traffic entering the zone that does not terminate at a service within the zone.

Examples of acceptable destinations within a zone include an application server, message bus, or proxy server, as follows:

  • An application server that validates and processes transaction requests
  • A file transfer server that receives files
  • A proxy server that examines, and potentially rewrites, the HTML headers
  • A communication server that receives an X.12 request, validates the format and sender, and forwards the request to the next service
  • An XML gateway that validates the object
  • A management interface used to configure or manage a switch, router, HIDS agent, or other networking device or agent

Examples of unacceptable destinations for network traffic entering a zone include:

  • A switch or router that forwards the traffic
  • A load balancer that merely inspects HTML headers and cookies, or rewrites the source / destination address or port
  • A firewall that permits the network traffic to exit the zone
  • A destination outside of the zone
  • A destination in any zone , such as a backup server, HIDS agent, or IDS server, unless the source is in the same zone

Please note that compliance with this BR supports the following best practice for network firewall configuration:

A network firewall need only be aware of the address ranges, subnets, and components of the networks directly connected to the firewall’s network interfaces. The firewall would have no firewall rules or routing tables permitting traffic to destinations in address ranges of other networks such as those beyond other adjacent network firewalls.

Although not required, CMS strongly encourages adoption of this best practice for firewall configuration.

Any exceptions to this BR or the recommended best practice require there be an appropriate business justification along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

Please refer to BR-SEC-FW-5: Filtering Traffic between Zones.

Rationale:

This BR supplements the business rule guidance in CMS TRA – Foundation, Business Rules with criteria for network traffic between components of an application in the CMS TRA Multi-Zone Architecture. The CMS TRA Multi-Zone Architecture is designed to separate the components of an application, thus creating multiple points in which network-based transactions may be inspected, challenged, and monitored. Among these points are some components of the CMS application and the CMS network infrastructure, the network zones in which CMS components and data reside, and the network firewalls protecting those zones. By having multiple points between the requester and CMS systems and data, CMS multiplies the complexity of its defenses and deters attackers.

Besides contributing to CMS defenses, compliance with this BR is an important contribution to the architecture’s overall value to CMS, namely, promoting consistency and efficiency through an enterprise view of service sharing, application development and reuse, operations and maintenance, and security services, while encouraging application design to emphasize performance, scalability, high availability, and security in depth.

BR-SEC-FW-10: Impede Attempts to Traverse Network Zones

The network infrastructure of CMS Processing Environments must impede or challenge any network traffic attempting to traverse (i.e., pass through) a zone. Unless explicitly approved by CMS, no authorized network traffic ever enters and or immediately exits a zone without being challenged or transformed by something within the zone as described by BR-SEC-FW-9. Therefore, any network traffic attempting to traverse a zone may be unauthorized. Although all network traffic must be monitored, unauthorized traffic may be discouraged or impeded using one or more of the following mitigation techniques:

  • Monitor the network for unusual or unexpected traffic between sources and destinations
  • Limit the paths and destinations in the routing tables of network infrastructure devices such as firewalls and routers. Monitor the routing tables for unauthorized changes.
  • Avoid using the same Transport Protocol Port Numbers on network firewalls to enter a given zone as used to exit the zone. For example, if a given port number is used to enter the Application Zone from the Presentation Zone, that same port number may not be used to exit the Application Zone and enter the Data Zone. This practice discourages unauthorized traffic from traversing a zone and would force a determined attacker to scan for a path to other zones. Such scanning increases the probability of detecting the attack.
  • Avoid configuring a given application to use the same Transport Protocol Port Numbers for traversing between consecutive zones. For example, an application may use one port number between the Presentation and Application Zones but should use a different port number between the Application Zone and Data Zone.
  • Use non-standard Transport Protocol Port Numbers whenever possible.

Any mitigation techniques used should be documented in the SSP. Any exceptions to this BR require the appropriate business justification along with documentation of the corresponding risk analysis for the system and environment in CFACTS and complete implementation of compensating information security and privacy controls from the current CMS ARS.

Rationale:

Impeding attempts to traverse network zones helps to discourage or defeat simple network attacks, and to detect misconfigured network devices or unauthorized network scanning. It also helps to defeat or detect malicious code outbound exfiltration attempts.

BR-SEC-FW-11: Utilize Firewalls from Two or More Different Vendors

In CMS data centers, CMS requires using network firewalls from two (2) or more different vendors at the various levels within the network to reduce the possibility of compromising the entire network.

The CMS TRA Multi-Zone Architecture has a natural flow of transactions from external networks to Presentation Zone to Application Zone to Data Zone (and in that sequence). Network firewalls are the gateways at the junctures between each zone. A network firewall at one juncture must be from a different vendor than the firewall at the next juncture in the sequence. This includes the juncture with external networks.

Please refer to BR-SEC-FW-10 - Impede Attempts to Traverse Network Zones .

Related CMS ARS Security Controls include: ARS SC-7 - Boundary Protection.

Rationale:

Employing firewalls from two (2) or more different vendors at the various levels within the network reduces the possibility of compromising the entire network. It reduces the likelihood that both firewalls share the same vulnerability or misconfiguration.

Firewall Administration

The following CMS TRA Security Firewall Administration (SEC-FWA) business rules support the information security and privacy controls cited in the CMS ARS. 

BR-SEC-FWA-1: Administrative Access to Firewalls

All system administration for firewalls must take place either from a local console directly connected to or part of the firewall hardware, or via secure connections from the Security Zone. In addition, CMS ARS Security Control IA-2 requires MFA for privileged user accounts.

Rationale:

It is essential to the security of CMS Processing Environments that only authorized individuals can administer firewalls, and that firewall administration only takes place over protected interfaces intended for firewall administration. Firewall administration is a function of the Security Zone .

BR-SEC-FWA-2: Firewall Implementation

Traffic-filtering network firewalls must not reside on a general-purpose computer or server or make use of a general-purpose operating system. Network firewalls must not use any general-purpose storage capabilities (write-once storage capabilities for recording audit logs are permissible) and must not be capable of executing arbitrary code or applications.

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-7 - Least Functionality.

Rationale:

Firewalls are critical to network operation and security. Implementing firewalls using dedicated and purpose-built hardware, software, and operating systems helps ensure reliable and correct operation of the firewall through the principle of least functionality.

BR-SEC-FWA-3: Disable Non-Firewall Functions

Network firewalls must not implement any services or capabilities other than firewall functions, e.g., DNS services, email services, File Transfer Protocol (FTP) services.

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-6 - Configuration Settings, AC-6 - Least Privilege, and CM-7 - Least Functionality. CMS ARS Security Controls include: CM-6 - Configuration Settings, AC-6 - Least Privilege, and CM-7 - Least Functionality.

Rationale:

Firewalls are critical to network operation and security. Disabling functions unrelated to required firewall functions helps ensure reliable and correct operation of the firewall through the principle of least functionality.

BR-SEC-FWA-4: Firewall Functional Requirements

All firewalls must provide the following functions at a minimum and must restrict the ability to perform these functions to an authorized administrator:

  • Performing startup and shutdown
  • Creating, deleting, modifying, and viewing information-flow security policy rules that permit or deny information flows
  • Creating, deleting, modifying, and viewing administrator profiles
  • Modifying and setting threshold for the number of permitted authentication attempt failures
  • Modifying and configuring logging parameters, including the definition of auditable events, audit log size limitations, and audit log retention policy
  • Restoring authentication capabilities for users who have met or exceeded the threshold for permitted authentication attempt failures
  • Enabling and disabling external hosts from communicating with the firewall
  • Modifying and setting the time and date
  • Archiving, creating, deleting, and emptying the audit trail
  • Backing up administrator profiles, information-flow security policy rules, and audit log data where the backup capability is supported by automated tools
  • Installing, modifying, or removing software
  • Starting and stopping services
  • Recovering to the state following the last backup

Related: Configuration Management (CM); CMS ARS Security Controls include: CM-6 - Configuration Settings, AC-6 - Least Privilege, and CM-7 - Least Functionality.

Rationale:

The listed functions are necessary to the proper operation, monitoring, and maintenance of a firewall.

Application Layer Filtering Firewalls and Gateways

Application layer gateway firewalls, also known as proxy-based firewalls, can monitor and filter on the application layer (i.e., layer 7 of the Open Systems Interconnection [OSI] model). They can look deep within the network packets’ content for inconsistencies, invalid or malicious commands, and executable programs. For example, a proxy for the SMTP protocol can be used to detect invalid commands and parameters used for potential attacks. HTTPS application proxies can check for tampering of cookies and invalid character encoding and more.

The guidance in this topic is intended to address all types of Application Layer Filtering (ALF) services. There are many services and products available that perform ALF and support many different protocols (e.g., SMTP, HTTP, and SOAP) and message/document formats and description languages (e.g., XML and JSON). Common examples of ALF services include, but are not limited to, proxy-based firewalls, XML firewalls, and Web Application Firewalls (WAF).

These configuration requirements fall into three categories: General (ALFG), Authentication and Authorization (ALFA), and XML-based Protocol Protection (XMLP), described in the subsequent subtopics of this chapter.

Configuration requirements listed in these subtopics are derived from NIST SP 800-95, Guide to Secure Web Services, and National Security Agency (NSA) Net-Centric Enterprise Services (NCES) Profile of Web Service Security: Simple Object Access Protocol (SOAP) Message Security (WSSE), NSA Profile 20080522.

ALF General Configuration Requirements

Configuration requirements in this subtopic apply to all ALF services processing digitally signed data elements such as, but not limited to, web service messages and XML or JSON documents. They are derived from NSA NCES WSSE, section 4.5 and NIST SP 800-95, section 3.6.

BR-SEC-ALFG-1: Accept Only Enveloped or Detached Signatures

ALF services must accept only digitally signed data elements (e.g., web service messages and XML or JSON documents). ALF services must accept only enveloped or detached signatures.

Rationale:

Digital signatures provide provenance—such as authenticity or origin—and can be used as a measure of integrity (trust).

BR-SEC-ALFG-2: Apply Exclusive Canonicalization

ALF services must be configured to apply only Exclusive Canonicalization with or without comments (Exclusive C14N Canonicalization transform).

BR-SEC-ALFG-3: Expand All Non-Character Entries

ALF services must be configured to expand all non-character entries when parsing a message or document.

BR-SEC-ALFG-4: Protect Against Web Service Attacks

All ALF services must be configured to protect against attacks, including but not limited to, the following types:

  • Recursive/oversized payload attacks – Attempts to perform a denial of service against the application service by sending messages designed to overload the parser
  • External reference attacks – Attempts to bypass protections by including external references that will be downloaded after the request message / document has been validated but before it is processed by the application
  • Routing Detours – Attempts to misdirect a message, causing it to be routed to an unauthorized location or to a non-existent location, creating a denial of service
  • Schema poisoning – Supplying a schema with the message / packet / document such that the message validator will use the supplied schema and allow a malicious message / document to be validated without error

All denied requests, protocols, services, destination Internet Protocol addresses, and ports must be logged in a manner compliant with the implementation standards within CMS ARS Security Control SC-7 (e.g., logged information must be available to the CCIC for analysis and alerting).

BR-SEC-ALFG-5: Reject Certain Types of Digital Signatures

All ALF services must be configured to reject digital signatures containing the following:

  • Uniform Resource Identifier attributes containing fragment identifiers (these follow a hash mark, as opposed to a query part)
  • For XML digital signatures, the MgmtData child element of the KeyInfo element (instead, they should use <xenc:EncryptedKey> and <xenc:DerivedKey> elements)

BR-SEC-XMLG-1: (Retired after TRA 2019R1): Accept Only Enveloped or Detached XML Signatures

BR-SEC-XMLG-2: (Retired after TRA 2019R1): Apply Exclusive Canonicalization

BR-SEC-XMLG-3: (Retired after TRA 2019R1): Expand All Non-Character Entries

BR-SEC-XMLG-4: (Retired after TRA 2019R1): Protect Against Web Service Attacks

BR-SEC-XMLG-5: (Retired after TRA 2019R1): Reject Certain Types of Digital Signatures

ALF Authentication and Authorization Configuration Requirements

Configuration requirements in this subtopic apply to all ALF services using digital signatures to identify servers, applications, or users for the purposes of authentication and/or authorization, or message integrity checking. Logging of authentication results (pass and fail) must comply with the CMS ARS requirements.

BR-SEC-ALFA-1: Authenticate Request Originators

The originators of requests to an ALF service acting as a web proxy must be authenticated using the x.509 standard certificate authority-based authentication. Derived from NSA NCES WSSE, section 4.1.

BR-SEC-ALFA-2: Chain of Identity, Authentication, and Authorization

When serving as intermediary web servers, ALF services must be configured to forward the identity in an X509SubjectName element containing the request originator’s distinguished name and those of all intermediate web servers in the request chain, including their own. This provides a chain of identity, authentication, and authorization from the request originator to the destination web server. Derived from NSA NCES WSSE, section 4.1.

BR-SEC-ALFA-3: Prevent Transformations on Signed Data Elements

ALF services must be configured to prevent the application of XML eXtensible Stylesheet Language Transformation (XSLT) or application-specific transforms on signed data elements. Derived from NSA NCES WSSE, section 4.5.

BR-SEC-XMLAA-1: (Retired after TRA 2019R1): Authenticate Request Originators

BR-SEC-XMLAA-2: (Retired after TRA 2019R1): Chain of Identity, Authentication, and Authorization

BR-SEC-XMLAA-3: (Retired after TRA 2019R1): Prevent Transformations on Signed Data Elements

ALF Protocol Protection Requirements

Configuration requirements in this subtopic apply to all ALF services providing protection mechanisms to servers, applications, or users accessing information using XML-based protocols.

BR-SEC-ALFP-1: Protect Against Attacks on Protocols

All ALF services must be configured to protect against attacks, including but not limit to, the following types against XML-based protocols such as Simple Object Access Protocol (SOAP) intrusion detection, or eXtensible HyperText Markup Language (XHTML) schema validation: Derived from NIST SP 800-95, section 3.6.

  • Web Services Description Language (WSDL) scanning – Attempts to retrieve the WSDL of Web services to gain information that may be useful for an attack
  • Parameter tampering – Modification of the parameters a Web service expects to receive in an attempt to bypass input validation and gain unauthorized access to some functionality
  • Replay attacks – Attempts to resend requests to repeat sensitive transactions
  • Structured Query Language injection – Providing specially crafted parameters that will be combined within the Web service to generate a SQL query defined by the attacker
  • Buffer overflows – Providing specially crafted protocol parameters that will overload the input buffers of the application and will crash the application service—or potentially allow execution of arbitrary code

All denied requests, protocols, services, destination IP addresses, and ports must be logged.

BR-SEC-ALFP-2: Protocol Headers

All ALF services must be configured to ignore the value of the lower-layer protocol header. Lower-layer protocols used to transport XML-based protocols must be configured to indicate the XML-based protocol type in their header (e.g., SOAPAction in an HTTP request’s header field). Derived from NSA NCES WSSE, section 4.2.

BR-SEC-XMLP-1: (Retired after TRA 2019R1): Protect Against Attacks on XML Protocols

BR-SEC-XMLP-2: (Retired after TRA 2019R1): XML Headers

ALF Service / Management Administration Requirements

The CMS TRA Security ALF Administration / Management (ALFM) configuration requirements in this subtopic apply to all ALF services and support the controls outlined in the CMS ARS.

BR-SEC-ALFM-1: Limit Administrative Access to ALF Services

All system administration for ALF services must take place from a local console either directly connected to or part of the ALF service hardware or via secure connections from the Management Zone.

BR-SEC-ALFM-2: ALF Services Must Not Support General-Purpose Computing

ALF services may be physical or virtual, non-reprogrammable devices. In addition, these services must not contain any general-purpose storage capabilities (write-once storage capabilities for recording audit logs are permissible) and must not have the ability to execute arbitrary code or applications.

BR-SEC-ALFM-3: ALF Service Functions

All ALF services providing inter-zone firewall or IDS/IDP services must provide the following functions at a minimum and must restrict the ability to perform these functions to an authorized administrator:

  1. Performing startup and shutdown
  2. Creating, deleting, modifying, and viewing information-flow security policy rules that permit, deny, constrain, and manage information flows
  3. Creating, deleting, modifying, and viewing administrator profiles and roles
  4. Modifying and setting threshold for the number of permitted authentication attempt failures
  5. Modifying and configuring logging parameters, including the definition of auditable events, audit log size limitations, and audit log rotation and retention policy
  6. Restoring authentication capabilities for users who have been disabled because they met or exceeded the limit for permitted authentication attempt failures
  7. Enabling and disabling hosts from communicating with the firewall
  8. Archiving, creating, deleting, and emptying the audit trail
  9. Backing up administrator profiles, information-flow security policy rules, and audit log data in a secure manner (e.g., encrypted, via management network) where the backup capability is supported by automated tools
  10. Installing, patching, modifying, or removing software
  11. Starting, stopping, and managing services
  12. Recovering to the state following the last backup

BR-SEC-ALFM-4: ALF Service Configuration Must Be Restricted

Changing the ALF service configuration is a privileged activity that must be restricted to authorized administrators.

BR-SEC-ALFM-5: ALF Service Configuration Changes Must Be Logged and Audited

Configuration changes must be logged and audited in a manner compliant with the CMS ARS requirements.

Related CMS ARS controls include: AC-3, AC-4, and SC-7.

BR-SEC-XMLA-1: (Retired after TRA 2019R1): Limit Administrative Access to XML Gateways

BR-SEC-XMLA-2: (Retired after TRA 2019R1): XML Gateways Must Not Support General-Purpose Computing

BR-SEC-XMLA-3: (Retired after TRA 2019R1): XML Gateway Functions

BR-SEC-XMLA-4: (Retired after TRA 2019R1): XML Gateway Configuration Must Be Restricted to Authorized Administrators

BR-SEC-XMLA-5: (Retired after TRA 2019R1): XML Gateway Configuration Changes Must Be Logged and Audited

 

CMS Cybersecurity Integration Center (CCIC) Integration

Introduction

Background

The Federal Information Security Management Act of 2002 (Public Law [P.L.] 107-347), as amended by the Federal Information Security Modernization Act of 2014 (P.L. 113-283) (FISMA), requires each agency to develop, document, and implement an agency-wide information security program to safeguard information and information systems that support the operations and assets of the agency, including those provided or managed by another agency, contractor (including subcontractor), or other source on behalf of an agency. Agency information security programs apply to all organizations (sources) that have physical or electronic access to a federal agency’s computer systems, networks, or IT infrastructure or use information systems to generate, store, process, or exchange data with a federal agency or on behalf of a federal agency, regardless of whether the data resides on a federal agency or contractor information system. CMS, the Federal Risk and Authorization Management Program (FedRAMP), and NIST have published guidance establishing the minimum security controls for safeguards needed to protect the confidentiality, integrity, and availability of a CMS system and its information.

CMS requires each system developer and maintainer to (1) process, (2) store, (3) facilitate transport of, and (4) host / maintain federal information pursuant to federal, HHS, and CMS Information Security Program policies by following the requirements outlined in this chapter. The term “CCIC” encompasses all CMS Information Security and Privacy Group’s operational programs, including traditional and emerging Security Operations Center roles, tools, and processes associated with these capabilities.

Purpose and Scope

This chapter provides requirements for integrating a FISMA system’s security operations and services capabilities with the CCIC security operations and services capabilities. The requirements include both the high-level requirements and information on locating additional details on relevant requirements.

This chapter supersedes all previous versions of the former CMS Enterprise Security Operations Center (ESOC) Integration Requirements but does not replace or supersede federal information security and privacy requirements, policies, regulations, or standards, including those from FISMA, Department of Homeland Security, HHS, and CMS. This chapter provides additional requirements to strengthen information security and privacy for CMS systems and the Agency. It also includes requirements for implementing the CMS-required security capabilities, including security architecture samples and specific tool requirements.

Most of the business rules in this chapter apply to all systems, regardless of FIPS 199 categorization. This chapter includes additional rules for CMS FISMA systems and data centers supporting CMS when any of the following apply:

  1. The FISMA system is categorized as HIGH or MODERATE under FIPS 199. Please refer to NIST SP 800-60, Guide for Mapping Types of Information and Information Systems to Security Categories, and CMS Risk Management Handbook, Risk Assessment (RA), see NIST SP 800-53A, Assessing Security and Privacy Controls in Information Systems and Organizations, RA-2 Security Categorization.or (PDF).
  2. Any FISMA system asset is designated by CMS as a High Value Asset (HVA).

A high value asset is an asset used as a mission-critical information resource supporting infrastructure providers / suppliers or partnering organizations. The unauthorized disclosure of, modification / destruction of, or disruption of access to information could be expected to have a severe or catastrophic adverse effect on organizational operations, organizational assets, or individuals.

  1. The SDM data center houses a FISMA system that is categorized as HIGH or MODERATE under FIPS 199 or is designated by CMS as HVA.

Assumptions and Constraints

In addition to FISMA, requirements set forth in this chapter support the following federal legislation, requirements, policies, regulations, or standards:

As required by federal law, each SDM must:

  • Follow FISMA, NIST, and other federal guidance in the CMS Policies and Guidance and Federal Policies and Guidance pages (including the CMS Acceptable Risk Safeguards and Risk Management Handbook), the CMS TRA guidance, and the CMS Target Life Cycle (TLC) processes and procedures. Any technologies described within this chapter or the associated onboarding documents must be implemented in accordance with the CMS ARS, the CMS TRA, and the TLC.
  • Obtain approval by the CMS TRB, which includes representation from ISPG, for any variance from this guidance. Failure to follow the requirements outlined in this chapter may result in a Plan of Action and Milestones finding and/or impact ATO eligibility.
  • Procure at least one CMSNet connection from the SDM to the CCIC and deploy the necessary hardware and software for the required information security and privacy capabilities. This includes infrastructure such as space, power, and cooling. CMS recommends the CMSNet connection capacity be 10Mbps or larger.
  • Accept responsibility for internal security monitoring, response, and reporting for information security and privacy incidents (and breaches) within SDM data center(s), including operational changes necessary to support the introduction and operation of the required information security capabilities and services. Alternatively, accept that CCIC performs security monitoring for the system after onboarding.
  • Provide technically qualified staff for operations and maintenance of the required information security capabilities and services compliant with roles and responsibilities defined in the CMS Information System Security and Privacy Policy.
  • Follow appropriate CCIC standard operating procedures, capability Onboarding Guides and other CMS documentation to support alignment with CCIC requirements. This includes supporting federal Information Security Continuous Monitoring (ISCM) required capabilities, such as hardware asset management (HWAM), software asset management (SWAM), vulnerability management (VUL), and configuration settings management (CSM).
  • For information security capabilities and services not addressed within this chapter, SDMs must provide the associated (real-time) continuous monitoring feeds as specified in NIST SP 800-137, Information Security Continuous Monitoring [ISCM] for Federal Information Systems and Organizations, by the DHS Continuous Diagnostics and Mitigation Program or other emerging federal requirements. CMS ISPG will assist and provide guidance to the SDM in deploying CCIC tools, collecting feeds, and integrating data sources.

Within CMS, compliance with FedRAMP alone for cloud deployments is the first step (i.e., an initial baseline) in meeting CMS’s confidentiality, integrity, and availability requirements. Cloud service providers are required to meet the technical and functional requirements stated in this chapter, either through log and event information provided by cloud components (e.g., infrastructure, platforms) in a manner that meets these requirements explicitly or though compensating capabilities (e.g., applications capable of providing the information needed) deployed on the hosted platform(s) or within the hosted application(s). In situations where requirements cannot be met, the FISMA system business owner is responsible for documenting the limitations and defining the resulting residual risk in the Information System Risk Assessment. All CMS FISMA systems and applications deployed on a CSP service must have a CMS-issued ATO. For more on CMS’s security requirements for cloud implementations, please refer to Risk Management Handbook Volume III Standard 3.2, CMS Cloud Computing Standard and IS2P2 policies on cloud computing.

Intended Audience

The intended audience of this chapter consists of the CMS SDMs, including all functional teams, FISMA system staff, or other working groups, that operate and/or maintain the information security and privacy data gathering capabilities. This includes technical staff responsible for maintaining the CMS networks and systems that reside within the networks. Information System Security Officers, Cyber Risk Advisors (CRA), and organizational Data Guardians understand the security and privacy drivers for the business rules in this chapter and can help as a liaison between the CCIC and the CMS SDMs. For the purpose of CDM, the SDM role may include the System Administrators, Application Administrators and Developers, and/or Database Administrators.

Integration with CCIC is a shared responsibility for everyone involved with the design, operation, or maintenance of CMS IT systems and data. Although this section is primarily intended for the audience associated with this chapter, it is relevant to audiences of other parts of the CMS TRA. All readers of any other CMS TRA chapter should refer to this section for additional responsibilities associated with CCIC integration.

Key Roles Supporting CCIC Integration Table defines key responsibilities for major IT roles to support integration with CCIC.

Key Roles Supporting CCIC Integration
IT RoleKey Responsibilities to Support Integration with CCIC
Business Owners
  • Coordinate with security, network, system, database, and application administrators and application developers to manage information security and privacy risk
  • Support the analysis of incidents involving Personally Identifiable Information and the determination of the appropriate action to be taken regarding external notification of privacy breaches as well as the reporting, monitoring, tracking, and closure of PII incidents
Information System Owners
  • Develop, implement, maintain, and oversee system-specific, role-based training applicable to system(s) under the information system owners’ purview
  • Ensure employees and contractors receive the appropriate training and education regarding relevant information security and privacy laws, regulations, and policies governing the information assets the information system owners are responsible for protecting
Local Security Operations Center Administrators
  • Perform real-time network and system security monitoring and triage
  • Perform analysis, coordination, and response to information security and privacy incidents and breaches
  • Perform security sensor tuning and management and infrastructure Operations and Maintenance (O&M)
  • Ensure SOC / Incident Response Team (IRT)-specific tools are implemented and deployed per the CCIC and vendor technical guidance
  • Serve as the FISMA system’s information security and privacy lead for CCIC and HHS Computer Security Incident Response Center (CSIRC)
Network Administrators
  • Ensure that the security posture of the network is maintained during all network maintenance, monitoring activities, installations, or upgrades and throughout day-to-day operations
  • Ensure that appropriate security requirements are implemented and enforced for all networks
System Administrators
  • Ensure that the security posture of systems is maintained during all system maintenance, monitoring activities, installations, and upgrades and throughout day-to-day operations
  • Ensure that appropriate security requirements are implemented and enforced for all systems
Application Administrators and Developers
  • Ensure automated information security and privacy capabilities are integrated and deployed as required
  • Coordinate with the ISSO to identify the information security and privacy controls provided by the applicable infrastructure that are common controls for information systems
  • Understand the relationships among planned and implemented information security and privacy safeguards and the features installed on the system
  • Ensure all development practices comply with the CMS ARS
Database Administrators
  • Ensure that the security posture of databases is maintained during all database maintenance, monitoring activities, installations, and upgrades and throughout day-to-day operations
  • Ensure that appropriate security requirements are implemented and enforced for all databases

The table Primary and Supporting Roles for the Cyber Security Operations/ Risk Management Section of the CCIC, lists the primary and supporting roles for each topic of this CCIC Integration chapter.

Primary and Supporting Roles for the Cyber Security Operations/ Risk Management Section of the CCIC Integration Chapter
Audience / SectionBusiness and System OwnersLocal SOC and Network AdministratorsSystem AdministratorsApplication Administrators and DevelopersDatabase Administrators
Risk ManagementPrimaryN/AN/AN/AN/A
Incident ManagementSupportingPrimarySupporting-ReportingSupporting-ReportingSupporting-Reporting
SOC ManagementSupportingPrimaryN/AN/AN/A
Data Feed and Log IntegrationSupportingSupportingSupportingSupporting 
Advanced InvestigationSupportingPrimarySupporting-ReportingSupporting-ReportingSupporting-Reporting
Information Sharing and CTISupportingPrimaryN/AN/AN/A
Penetration TestingSupportingPrimarySupporting – Awareness and Problem ResolutionSupporting – Awareness and Problem ResolutionSupporting – Awareness and Problem Resolution
Security Architecture and EngineeringPrimaryPrimaryPrimaryPrimaryPrimary
High Value AssetsPrimaryPrimaryPrimaryPrimaryPrimary

 

Table - Primary and Supporting Roles for Information Security Continuous Monitoring (SCM)/ Privacy Continuous Monitoring (PCM)
Audience / SectionBusiness and System OwnersLocal SOC and Network AdministratorsSystem AdministratorsApplication Administrators and DevelopersDatabase Administrators
CyberScopeSupportingPrimarySupportingSupportingSupporting
CDM Asset ManagementN/AN/ASupportingSupportingSupporting
ISCM / CDM Capability ManagementSupportingPrimaryN/AN/AN/A
HWAMSupportingPrimarySupportingN/AN/A
SWAMSupportingPrimarySupportingSupportingSupporting
Config Settings ManagementSupportingPrimarySupporting – Operating SystemsSupporting – Web ApplicationsSupporting - Databases
Vulnerability ManagementSupportingPrimarySupporting – Operating SystemsSupporting – Web ApplicationsSupporting - Databases
SOC as a ServiceN/AN/AN/AN/AN/A

 

Table - Primary and Supporting Roles for CCIC Integration, Perimeter Protections
Audience / SectionBusiness and System OwnersLocal SOC and Network AdministratorsSystem AdministratorsApplication Administrators and DevelopersDatabase Administrators
Perimeter MonitoringSupportingPrimarySupportingN/AN/A
Full Packet Capture / InspectionSupportingPrimarySupportingN/AN/A
Intrusion Detection/PreventionSupportingPrimarySupportingN/AN/A
Malware Detection/PreventionSupportingPrimarySupportingN/AN/A
Network FirewallsSupportingPrimaryN/AN/AN/A
Network Data Loss PreventionSupportingPrimaryN/AN/AN/A

 

Table - Primary and Supporting Roles for the Lateral and Endpoint Protections topic of the CCIC Integration Chapter
Audience / SectionBusiness and System OwnersLocal SOC and Network AdministratorsSystem AdministratorsApplication Administrators and DevelopersDatabase Administrators
Network Security Endpoint ProtectionSupportingPrimaryPrimaryN/AN/A
FIPS 140-2 or FIPS 140-3 Validated EncryptionSupportingPrimaryPrimaryN/AN/A
Insider Threat ProtectionSupportingPrimarySupportingSupportingSupporting
Endpoint Data Loss PreventionSupportingPrimaryPrimaryN/AN/A

 

Table - Primary and Supporting Roles for the CDM Considerations Topic of the CCIC Integration Chapter
Audience / SectionBusiness and System OwnersLocal SOC and Network AdministratorsSystem AdministratorsApplication Administrators and DevelopersDatabase Administrators
CDM Identity and Access ManagementSupportingFutureFutureFutureFuture
CDM Network Security ManagementSupportingFutureFutureFutureFuture
CDM Data Protection ManagementSupportingFutureFutureFutureFuture
Ongoing Assessment and AuthorizationSupportingFutureFutureFutureFuture

Overview

CMS ISPG established the CCIC to deliver a number of important Agency-wide security services designed to protect CMS systems and provide adequate, risk-based, cost-effective cybersecurity. These services help ensure the confidentiality, integrity, and availability of information processed, stored, or transmitted by CMS systems.

The services offered include vulnerability assessment and management as provided by the Vulnerability Assessment Team (VAT), security engineering, incident management, forensics and malware analysis, cyber threat intelligence, and penetration testing.

Historically, the VAT has focused primarily on vulnerability scanning and detection, vulnerability assessment, vulnerability management, and vulnerability reporting. In addition to those services, the VAT also supports both the Division of Implementation and Reporting (DIR) and continuous diagnostics and mitigation (CDM) operations for ISPG. VAT supports DIR by providing:

  • Project management support for deployment of tools and capabilities that support VAT services
  • Consulting and architectural guidance
  • Reporting and monitoring
  • Process and procedure optimization
  • Deployment and administration of tools that support VAT services

For CDM operations, the VAT:

  • Provides vulnerability data
  • Reduces CMS’s exploitable attack surface
  • Assists CMS with responding quickly to new and emerging vulnerabilities

For more information please contact the VAT team at VAT@cms.hhs.gov

The CCIC supports and provides a variety of enterprise-level information security capabilities and services for CMS by:

  • Maintaining an independent and enterprise-wide information security operations perspective
  • Centrally coordinating CMS enterprise information security operations
  • Centrally coordinating and disseminating CMS-applicable information security and privacy threat intelligence
  • Managing aggregation, correlation, and reporting to meet enterprise information security and privacy controls and offer common controls
  • Providing integration requirements for implementing CCIC technical requirements to ensure a more fully integrated information security environment for the entire enterprise

The CCIC is not a replacement for, nor intended to remove the requirement for, the SDM’s own security policy, operations personnel, procedures, and responsibilities. The CCIC extends the Defense-in-Depth strategy by providing an additional layer that provides monitoring, response, and notification across the entirety of CMS. This additional layer augments and enhances the security controls and capabilities for the FISMA system and SDM.

The SDM maintains responsibility for appropriately staffing and implementing its information security and privacy program to meet CMS requirements.

Base CCIC Capabilities

The base CCIC capabilities support the following strategic areas:

  • Cybersecurity Operations / Risk Management – Capabilities associated with the management of risk, incidents, and operations. This includes advanced analytics and security engineering.
  • ISCM / PCM – Capabilities associated with ISCM / PCM as defined by NIST and OMB under the DHS Continuous Diagnostics and Mitigation (CDM) Program. This includes ongoing assessment and authorization.
  • Perimeter Protections – Capabilities associated with information security and privacy protections at the external boundary and enclave perimeter.
  • Lateral and Endpoint Protections – Capabilities associated with information security and privacy protections internal to an enclave (e.g., at the endpoint or a device within the enclave) and insider threats. This includes penetration of the endpoint (e.g., phishing).

Reference CCIC Integration Architecture

CMS does require similar implementations if sensitive data, such as Personally Identifiable Information, Protected Health Information, or Federal Tax Information, are used within non-production environments such as test and development.

The following recommended implementation details should be noted:

  • One or more service architectures exist within each FISMA system enclave. (A “secure enclave” is defined as one or more segments of an internal and/or isolated network consisting of a common set of security policies.)
  • One or more FISMA systems exist within each data center enclave.
  • One or more user enclaves may exist within the FISMA system, across multiple FISMA systems, within the data center, or external to the data center.
  • Monitored perimeters exist between enclaves and boundaries. (Boundary is defined as the separation point between one or more network segments/enclaves.)
  • The Trusted Internet Connection 3.0 reference architecture is now required for deployments.
  • Secured enclaves host the SDM’s Supporting SOC and the SOC tool suite.
  • CMSNet is the secure connection for sharing information security and privacy information with CCIC. (CMS recommends that the CMSNet connection capacity be 10Mbps or larger.)
  • A CMS TRA- and TLC-compliant, tiered architecture is used.
  • Strategically placed sensor components for scanning of all assets via network-based sensors, capability relays, information collectors, or local servers include:
    • Vulnerability scanning tool
    • Configuration setting scanning tool
    • Software asset scanning tool
    • Hardware asset scanning tool
  • Log collection and correlation permits the CCIC-based deployment of the enterprise logging tool to inter-operate with the FISMA system’s (or SDM’s) Security Information and Event Management capability.
  • Network taps provide full packet capture, malware detection and capture, for Intrusion Detection Systems / Intrusion Prevention Systems (IDS / IPS).
  • Layer-3 configuration files are analyzed for vulnerabilities and misconfigurations via a contextual risk management and threat identification application. (Layer 3 (Network Layer) handles packet forwarding and routing.)
  • Other ISCM / PCM information is collected and shared as needed with CCIC for CyberScope and other program requirements.
  • Endpoint protection is deployed on all applicable assets for malware / virus detection to support forensic acquisition and to share indicators of compromise (IOC).

Advanced CCIC Capabilities and Services

The CCIC provides a variety of security capabilities and services to fulfill its mission of monitoring for and responding to information security and privacy incidents and breaches. This subtopic describes the advanced CCIC capabilities and services available to the FISMA system and SDM. These capabilities and services help to facilitate the successful deployment of the required capabilities. Security capabilities and services also referenced in HHS OCIO Rules of Engagement for Security Monitoring and Collaborative Systems.

The following CCIC services are available to the SDM:

  • Incident Management Team (IMT). Incident management, through an IMT, serves as the focus for all incident handling, across all information response teams (IRT) within CMS. See Incident Management below for additional information on the IRT.
  • Forensic & Malware Analysis Team (FMAT). The Advanced Investigation service oversees the handling of any incidents that require digital forensics and malware analysis on affected system(s). See Advanced Investigation for additional information on FMAT.
  • Information Sharing and Cyber Threat Intel (IS&CTI). The IS&CTI service improves the effectiveness of CMS information security and privacy. See Information Sharing and Cyber Threat Intelligence for additional information on IS&CTI.
  • Penetration Testing. CMS established a penetration testing capability within ISPG to proactively assess the susceptibility of CMS systems, personnel, and facilities to attack. See Penetration Testing below for additional information on penetration testing.
  • Security Architecture and Engineering (SAE). The SAE service provides architecture and design of information security and privacy solutions across the CMS enterprise, including government and contractor facilities. See Security Architecture and Engineering below for additional information on.

 

Cyber Security Operations / Risk Management

Risk Management

This subtopic provides the risk management requirements associated with cyber security operations. These requirements specify the following capabilities:

  • Obtaining an ATO for the FISMA system
  • Assessment of information security and privacy risks

The business rules within this topic apply to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-01: Security Authorization of Systems

The business owner / information system owner / SDM must:

  • Categorize the FISMA system in accordance with FIPS 199, and document the system attributes used to identify PII, PHI, and HVAs
  • Develop, periodically review, and maintain information security, privacy, and continuous monitoring plans (NIST, SP 800-18 Rev. 2, Developing Security , Privacy, and Cybersecurity Supply Chain Risk Management Plans for Systems) in accordance with CMS ARS-defined transitional conditions (including major system changes and data calls following an increased threat level).
  • Implement and maintain information security and privacy controls as required by the CMS ARS and the FISMA system information security and privacy plans
  • Ensure all personnel meet information security and privacy requirements within the HHS Information System Security and Privacy Policy and CMS IS2P2 policies for:
  • Support the authorization of the FISMA system by implementing the controls, completing required activities, supplying information, and documenting the details in cybersecurity and privacy artifacts required under the CMS Security Assessment and Authorization Process
  • Manage system interconnections in compliance with CMS-ISA policies governing information sharing and exchange as defined by HHS, CMS, and other applicable information-sharing agreements

For transaction-based FISMA systems, the business owner / information system owner / SDM must:

  • Implement and maintain a transaction recovery system(s) appropriate for the CMS mission or business function

For FISMA systems categorized under FIPS 199 as MODERATE or HIGH, the business owner / information system owner / SDM must:

  • Establish and maintain alternate processing and storage site agreements appropriate for the CMS mission or business function

For FISMA systems categorized under FIPS 199 as HIGH, the business owner / information system owner / SDM must:

  • Develop and maintain IT contingency plans appropriate for the CMS mission or business function:
    • Ensure coordination with organizational elements responsible for related plans
    • Ensure processes and procedures prevent the unauthorized removal of maintenance equipment / media

For FISMA systems categorized under High Value Assets, the business owner / information system owner / SDM must:

  • Verify system attributes
  • Participate in RVAs
  • Complete additional monitoring and reporting requirements as defined by HHS and within OMB M-19-03

Rationale:

Proper understanding of a system’s value to CMS as an asset helps CMS, the SDM, and the CCIC properly prioritize response and remediation activities.

Related CMS ARS Security Controls include: CA-1 - Policy and Procedures, CA-2 - Control Assessments, CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, RA-3 - Risk Assessment, RA-8 - Privacy Impact Assessments, SA-1 - Policy and Procedures, SA-4 - Acquisition Process, SA-9 - External System Services, PT-3 - Personally Identifiable Information Processing Purposes, RA-2 - Security Categorization, AT-3 - Role-Based Training, AC-3(9) - Supplemental: Controlled Release, AC-21 - Information Sharing, CP-2 - Contingency Plan, MA-2 - Controlled Maintenance, and MA-3 - Maintenance Tools.

BR-CCIC-02: Assessment of Information Security and Privacy Risks

The business owner / information system owner / SDM must:

  • Ensure security impact is identified:
  • Ensure protection of PII:
    • Conduct a Privacy Threshold Analysis (PTA)
    • Conduct an initial Privacy Impact Assessment (PIA) and recurring PIA reviews
    • Address and mitigate the risks identified
    • Update the PIA throughout the TLC process
    • Identify PHI and FTI and apply applicable controls
  • Integrate with CCIC information security capabilities and services:
    • Ensure coordination with the CMS Chief Information Security Officer (CISO), Senior Official on Privacy (SOP), business owner(s), ISSO(s), and other stakeholders
    • Provide support for automated control assessment as defined by CMS and the FISMA system’s security and privacy plan(s
    • Configuration setting management
    • Vulnerability management
    • Asset inventory management
  • Conduct independent risk assessments on the FISMA system documenting the results (Assessment processes and procedures are defined under NIST SP 800-30 Rev. 1, Guide for Conducting Risk Assessments, and NIST SP 800-53A, Assessing Security and Privacy Controls in Information Systems and Organizations.)
  • Evaluate level of risk periodically and in accordance with CMS ARS-defined transitional conditions: (including major system changes and data calls following an increased threat level)
    • Notify CMS stakeholders when the risk is found to exceed organizationally defined thresholds
    • Update information security and privacy plans to address changes in risk, mitigations, and compensating controls

For FISMA systems categorized under FIPS 199 as MODERATE or HIGH, the business owner / information system owner / SDM must:

Perform coordinated contingency testing and/or exercises with organizational elements responsible for related plans

For FISMA systems categorized under FIPS 199 as HIGH, the business owner / information system owner / SDM must:

Employ automated mechanisms to schedule, conduct, and document required maintenance and repairs

Maintain up-to-date, accurate, and complete records of all maintenance and repair actions needed, in progress, and/or completed (including both repair / preventive maintenance on hardware and patching/reconfiguring on software)

Rationale:

Proper assessment of a system’s risk helps the CMS CISO, SOP, AO, SDM, and CCIC properly prioritize response and remediation activities. Automated mechanisms can be leveraged to improve overall FISMA system operational efficiency.

Related CMS ARS Security Controls include: CA-1 - Policy and Procedures, CA-2 - Control Assessments, CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, CA-7 - Continuous Monitoring, CA-7(4) - Risk Monitoring, RA-3 - Risk Assessment, RA-8 - Privacy Impact Assessments, SA-1 - Policy and Procedures, SA-4 - Acquisition Process, SA-9 - External System Services, SI-2 - Flaw Remediation, SI-4 - System Monitoring, MA-2(2) - Automated Maintenance Activities, , CM-6 - Configuration Settings, CM-7 - Least Functionality, CM-8 - Information System Component Inventory, PT-3 - Personally Identifiable Information Processing Purposes, PT-6(1) - Routine Uses, RA-8 - Privacy Impact Assessments, and RA-5 - Vulnerability Monitoring and Scanning.

Incident Management

The CMS Incident Management Team coordinates all functions of incident handling and response across CMS’s incident response (IR) teams. In accordance with the CMS IS2P2 Roles and Responsibilities requirements, the FISMA system’s IRT supports CMS’s Security Operations team, which is integrated into the CCIC. The CMS IS2P2 defines the security operations and incident response organizational structure for the CMS enterprise. The FISMA system SOC/IRT operates under the direction and authority of the ISSO and Business Owner /ISO.

A primary responsibility of the CCIC’s Security Operations team is coordinating incident management across the CMS enterprise. The CCIC IMT provides an effective incident handling and response capability by working with the Marketplace SOC, the FISMA System SOC(s), and other IR external organizations. The CCIC provides CMS leadership with up-to-date information (i.e., situational awareness) regarding the status of all security and privacy incidents (and breaches).

The business rule within this subtopic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-03: Onsite Incident Response Team

The Onsite Incident Response Team must consist of technically qualified staff, including trained information security and privacy personnel, to fulfill the roles and responsibilities of the IRT. Small SDMs may choose to defer part or all required capabilities to the CCIC via an appropriate Memorandum of Understanding (MOU) / Memorandum of Agreement (MOA), and by leveraging the CCIC Incident Response Plan. Please consult the RMH for guidance and templates in the Information Security Library on CMS.gov. Otherwise, please request assistance by sending an email to cisso@cms.hhs.gov.

The SDM must:

  • Provide an IRT to respond when a security or privacy incident affects its IT system:
    • Ensure responses are coordinated with the CCIC IMT
  • Be responsible for internal security monitoring, response, and reporting for information security and privacy incidents within the SDM data center, including all operational changes
  • Develop, maintain, and test organizational-level IR plans, processes, and procedures as required by the CMS ARS:
    • Ensure processes and procedures are in line with CCIC-provided workflows for properly escalating incidents with predefined criteria, thresholds, and contacts
    • Include relevant point of contact (POC) and Office of Information & Regulatory Affairs membership information
    • Ensure incident response plan(s) are provided electronically to the CMS CISO, SOP, and the CCIC
  • Provide incident response training for the IRT

The SDM IRT must:

  • Serve as the on-site incident response authority for information security and privacy incident reporting and tracking
  • Coordinate IR activities with the CCIC and the SDM’s ISSO:
    • Assist in information gathering, forensics, response, and reporting activities
    • Respond to advisories, requests, or directives issued by the CCIC
    • Participate in CMS-led incident response and Playbook tabletop / training exercises
  • Provide timely responses to informational requests for SDM security status and posture information

(A timely response is a response within a defined response window associated with the threat level of the security violation. While a response window is typically 24 to 72 hours, the duration is at the discretion of the CMS CISO.)

  • Provide technical support and advice for incident handling, impact assessment, and technical system management, including actions taken should standard operating procedures not cover the circumstances
  • Serve as the focal point for reporting, monitoring, and tracking to closure of reported information security and privacy incidents:
  • Small SDMs may choose to leverage one or more of the CCIC-provided IR-related services:
    • Network Security Monitoring service
    • FMAT service
    • IS&CIT service

Rationale:

The SDM onsite IRT, which acts as a subcomponent of the CCIC IMT, serves as the focal point for reporting, monitoring, and tracking to closure of reported information security and privacy incidents.

Related CMS ARS Security Controls include: IR-05 - (Security) Incident Response, IR-8 - Incident Response Plan, IR-8(1) - Breaches, CA-2 - Control Assessments, AT-2 - Literacy Training and Awareness, AT-3 - Role Based Training, PL-4 - Rules of Behavior, CA-7 - Continuous Monitoring, CA-7(4) - Risk Monitoring, and SI-4 - System Monitoring.

SOC Management

The CCIC uses a variety of security capabilities to monitor, detect, and respond to cybersecurity and privacy incidents and breaches. This subtopic provides the business rules associated with capability implementation within each Supporting SOC operational environment to ensure interoperability with the CCIC (a Supporting SOC is the unit that provides SOC capabilities for the FISMA system / SDM in coordination with the CCIC).

The FISMA system and SDM will work with the CCIC to integrate all required security tools and functional requirements identified in this chapter as part of a unified architecture.

SOC management ensures:

  • Interoperability among information security and privacy applications, utilities, and tools implemented within the FISMA system / SDM Supporting SOC and the CCIC
  • Timely response to CMS requests related to IOCs, data calls, incident or breach investigations, and reporting
  • Relevant privacy controls are addressed to account for sensitive data that may leak into the analysis process

The business rules within this subtopic apply to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-04: Local Secure Enclave to Support CCIC Capabilities

The SDM must:

  • Provide oversight, monitoring, and incident response capabilities within the local data center:
    • Audit and report physical and/or logical access to the equipment
  • Provide and maintain a secured, separate enclave (i.e., create and maintain the Security Zone) of networks, systems, and devices to implement the CMS information security and privacy functionality, including access control mechanisms, tools, tool sensors (collectors), data aggregators, and data (exceptions must be documented and authorized as a deviation) that assures:
  • Physical / logical separation from other SDM networks (including production) and all non-CMS data center operations and security subnets
    • Required boundary protection mechanisms (e.g., IDS / IPS, packet inspection firewalls, dynamic execution environments (also known as detonation chambers), including technical capability requirements specified in CCIC Perimeter Protections.
    • Secure (i.e., encrypted) connectivity to the CCIC via the CMSNet connection
    • Secure connectivity into the enclave for agents (including devices) used to collect required logs and data feeds operating outside the enclave
  • Provide a physical / logical environment supporting CCIC-required and managed information security and privacy capabilities:
    • Rack space, cabling, connectivity, and appropriate environmental support for systems, appliances, and devices or host space and connectivity for virtualized systems, appliances, and devices
    • Physical separation secured against unauthorized access, either by a dedicated cage or by use of a separate, dedicated room / rack (for service providers)
    • Physical and/or logical access limited to authorized individuals (e.g., SDM security, SDM operations personnel) with authorization for access managed by the ISSO and CMS
    • Availability on a 24/7/365 basis per CMS and federal requirements
  • Provide a capability environment supporting CCIC-required and managed information security and privacy capabilities:
    • Event management (e.g., SIEM)
    • Incident management and response
    • Forensics and malware analysis
    • IS&CTI

Rationale:

The SDM’s secure enclave provides a protected subnet that houses applications and tools for protecting applications and systems deployed within the SDM’s data center(s). The secure enclave is the focal point for providing information security and privacy capabilities across the CMS portion of the SDM’s enterprise.

Each SDM is also responsible for maintaining appropriate security and access control for all secure enclaves/boundaries and for implementing the appropriate tools and technologies to meet CMS and federal requirements. Information sharing between the CCIC and the SDM secure enclave / boundary must use the CMSNet network and be encrypted via an IPSec VPN tunnel to maintain confidentiality and integrity of the data in transit. The CMS ARS directs that encryption algorithms in support of IPSec tunnels must meet or exceed CMS requirements (i.e., be a FIPS 140-2 or FIPS 140-3 validated module).

Related CMS ARS Security Controls include: PL-8 - Security and Privacy Architectures, SA-8 - Security and Privacy Engineering Principles, SA-17 - Developer Security and Privacy Architecture and Design, SC-3 - Security Function Isolation (with enhancements), SC-39 - Process Isolation, AC-2 - Account Management, AC-6 - Least Privilege (with enhancements), AU-2 - Event Logging, AU-9 - Protection of Audit Information, PE-3 - Physical Access Control, PE-6 - Monitoring Physical Access, and SC-13 - Cryptographic Protection.

BR-CCIC-05: Interoperability with CCIC

The SDM must:

  • Provide bandwidth necessary to ensure security- and privacy-relevant data are not:
    • Truncated or lost during transit
    • Degraded before aggregation and use by the CCIC’s SIEM
  • Ensure information security and operations staff are adequately and appropriately trained in: (See HHS Memorandum, Requirements for Role-Based Training of Personnel with Significant Security Responsibilities[MT15] )[JD16] 
    • Incident management and response
    • Security operations (detection and analysis)
    • Malware analysis (both static and dynamic) and digital forensics
  • Implement security capabilities identified by vendor name / functional requirement within the timeframes defined by CMS or other federal mandate:
    • Procure, implement, and perform O&M on the required information security and privacy tools and platforms supporting the required information security and privacy tools
    • Implement the required information security and privacy tools in a manner and using an architecture as required by information security and privacy policy
  • Provide authorized personnel within ISPG with administrative access to information security and privacy tools on an as-needed basis for tuning and policy changes:
    • SDM tool administrators and analysts retain responsibility for the routine O&M of the system.
  • Provide the CCIC with a service account with appropriate administrative privileges (i.e., membership in the enterprise logging tools’ administration group) on each of the SDM’s Indexers to facilitate the CCIC Search Head’s identifying the SDM Indexer(s) as a search peer. (Creating the search peer relationship between the SDM Indexers and the CCIC Search Heads is a one-way share. Although the analysts at the CCIC can access peered data stored at the SDMs, the analysts at the SDMs cannot access any data stored at the CCIC.)

Rationale:

The SDM’s secure enclave must be interconnected with the CCIC to support its oversight and enterprise-wide responsibilities. This includes providing the CCIC with the ability to perform ad hoc risk evaluations in an evolving threat environment. Although CMS subject matter experts (SME) could access the Supporting SOC configuration files when needed, the FISMA system / SDM tool administrators and analysts retain responsibility for the routine O&M of the system and data sources protecting CMS IT assets.

The SDM is providing CCIC staff with appropriate privileged / administrative access to support interoperability with the CCIC capabilities. This includes the ability to modify rule sets and tool-based policies when needed. (To remain effective, information security and privacy capabilities must be constantly tuned to meet threats and priorities at the CMS Enterprise level. Privileged / administrative access is required to facilitate such tuning. Examples include updating / adding detection rules based on changing IOCs, adjusting defined configuration baselines, and implementing appropriate threat protections.)

In addition to the specific product-based security- and privacy-related capabilities defined under the CMS security architecture and implemented by the CCIC, FISMA system and SDM security analysts may implement other proprietary and open source tools to provide additional capabilities that support both detection and incident investigation. Any findings produced by these tools during an investigation must be made available to and fully usable by the CCIC.

Related CMS ARS Security Controls include: PL-8 - Security and Privacy Architecture, SA-8 - Security and Privacy Engineering Principles, SA-17 - Developer Security and Privacy Architecture and Design, AT-3 - Role-Based Training, PS-3(3) - Supplemental: Information Requiring Special Protective Measures, AT-2 - Literacy Training and Awareness, AT-3 - Role-Based Training, PL-4 - Rules of Behavior, IR-8 - Incident Response Plan, and IR-8(1) - Breaches.

BR-CCIC-06: Timely Response to CCIC Requests

The FISMA system business owner must:

  • Appoint an ISSO as the official POC for CMS security data calls
  • Notify the CCIC within thirty (30) days of a change of ISSO (CA-2)

The ISSO must:

  • Update POC, network architecture (including topology maps), IP address ranges, infrastructure updates, and security documentation for the systems the ISSO and/or CRA operate on behalf of CMS. The ISSO and/or CRA must:
    • Use CFACTS to record the changes
    • Complete updates within the timeframes defined within the CMS ARS
    • Update volatile information, such as network architecture (e.g., topology maps), no less often than once every thirty (30) days
    • Provide a full list of subnets no less often than once every ninety (90) days to support vulnerability scanning

The SDM must:

  • Notify the CCIC of significant changes to architecture, security posture, or other items that could cause degradation or unexpected results in security monitoring, detection, response, and mitigation activities prior to making a change
  • Notify the CCIC of subnet changes (additions or deletions) within 24 hours of making the change
  • Maintain and provide changes to the system accounts required for credentialed scanning by tools used by the CCIC at least two (2) weeks before the passwords expire or when other changes to the accounts are needed:
    • Changes to the accounts must be communicated to the CCIC using out-of-band channels and techniques meeting and exceeding CMS encryption requirements (a FIPS 140-2 or FIPS 140-3 validated module)
  • Provide timely responses to information requests for SDM security status and posture information. A timely response is a response within a defined response window associated with the threat level of the security violation. Although a response window is typically 24 to 72 hours, the duration is at the discretion of the CMS CISO.
  • Adhere to response times associated with information security and privacy incidents and breaches, and must follow CMS-defined timeframes that are dependent on:
    • Assigned severity of the incident or breach
    • Nature and scope of the incident or breach
  • Report FISMA system information security and privacy incidents and breaches to CCIC and HHS CSIRC as required by federal law, regulations, mandates, and directives, following the RMH. For more specifics on incident reporting and response, please refer to:
    • IS2P2, Section 3.4
    • NIST SP 800-61, Computer Security Incident Handling Guide
    • HHS-OCIO, Policy for Information Technology (IT) Security and Privacy Incident Reporting and Response
    • RMH Chapter 8: Incident Response
  • Respond to information requests, which may include, but are not limited to:
    • Hardware / software patch implementation levels
    • Implementation status for specific IT security- and privacy-related updates
    • System / infrastructure inventories and configurations (including components and software)
    • Specific network traffic levels
    • Security and privacy specific metrics

Rationale:

The SDM’s response to CCIC data calls must support a dynamically changing environment. This requires the ability to respond to ad hoc risk evaluations within the evolving threat environment.

A frequent issue reported by the CCIC is out-of-date contact information. This has resulted in delays when initiating further investigation of a potential incident or responding to a confirmed incident. A second issue is alerts generated as the result of a planned / authorized change. By properly maintaining POCs in CFACTS, many of these false positives can be avoided.

Information sharing between the CCIC and the SDM secure enclave / boundary must use the CMSNet network.

Related CMS ARS Security Controls include: IR-05 - (Security) Incident Response, AC - Access Control family, AT-2 - Literacy Training and Awareness, AT-3 - Role-Based Training, PL-4 - Rules of Behavior, CA-2 - Control Assessments, CA-7 - Continuous Monitoring, CA-7(4) - Risk Monitoring, CM - Configuration Management family, MA - System Maintenance family, PE - Physical and Environmental Protection, RA-5 - Vulnerability Monitoring and Scanning, RA-5(11) - Public Disclosure Program, RA-7 - Risk response, IR-8 - Incident Response Plan, IR-8(1) - Breaches, SC-7 - Boundary Protection, SC-13 - Cryptographic Protection, SI-2 - Flaw Remediation, SI-3 - Malicious Code Protection, SI-4 - System Monitoring, and SI-5 - Security Alerts, Advisories, and Directives.

Data Feed and Log Integration

CMS’s ISPG maintains an enterprise SIEM capability to meet HHS and federal security information logging and analysis requirements, support incident detection and response capabilities, and assist in meeting reporting requirements.

While network-based activity has long been a primary source of monitoring information associated with malicious activity, changes to both advanced threats and industry best practices require a more comprehensive monitoring approach. To address these changing threats, continuous monitoring will supplement CMS’s ability to quickly and efficiently detect security and privacy breaches and insider threats. Please refer to NIST SP 800-137, Information Security Continuous Monitoring for Federal Information Systems and Organizations.

CMS has implemented an enterprise-level security information and event management (SIEM) platform as the data backbone for security operations within both CMS and the CISA Continuous Diagnostics and Mitigation (CDM) Program. This platform collects, aggregates, correlates, and analyzes logs and event information within each SDM from sensors used to monitor network and host activity into a unified, operational intelligence platform.

The business rule within this subtopic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-07: Local Security Information and Event Management Capability

All SDMs must deploy a SIEM toolset compatible with the CMS enterprise SIEM platform which allows the SDM to capture file, process, registry, service, and thread activities on a per host, server, and network basis anywhere within the SDM network used to support CMS. This approach allows the CCIC or CDM team to submit queries to the SDM’s SIEM for analysis and reduces the information returned to the query response. The CCIC can search through the data at all SDMs without having to copy this same data to the CCIC. For information on the requirements for integrating the SDM SIEM with the CMS enterprise SIEM platform and CDM toolset, send email to ISPG_Logs_Onboarding@cms.hhs.gov.

The SDM must:

  • Deploy an enterprise logging tool, such as Splunk or equivalent to provide a local SIEM capability:
  • The product may be used (a) if the product provides native capabilities for correlation and contextualization of detected suspect activities from the set of logs and data sources being consumed by the tool, and (b) is fully compatible with the CCIC enterprise SIEM platform deployment:
    • Full data integration – Data generated by this product can be aggregated into the CCIC enterprise SIEM platform for analysis.
    • No loss of data integrity – Data forwarded to the CCIC enterprise SIEM platform is not modified or truncated.
    • Raw format is available when needed.
  • Collaborate with the CCIC to ensure the SDM’s current and planned infrastructure:
    • Integrates with the CMS federated enterprise logging tool deployment
    • Conforms to the CMS enterprise information security architecture
  • Provide CCIC with a notification within 30 days when a new data source becomes active to ensure the CCIC enterprise SIEM platform instance is receiving the most up-to-date events

SIEMs deployed within the SDM must:

  • Support encryption using FIPS 140-2 or FIPS 140-3 validated modules of all communications, when supported, between the implemented SDM SIEM and any applications, systems, servers, and other monitoring and analysis tools, to protect against the spillage of sensitive information / data
  • Run automated, customizable reports in accordance with ISPG and federal reporting requirements
  • Generate alerts based on triggers and/or signatures that are:
    • Customized / modified
    • Generated / implemented automatically from authoritative external sources (e.g., vendors and other federal agencies)
    • Integrated automatically with ticketing or workflow solutions deployed within the SDM data center(s) to facilitate remediation activities
  • Serve as a data source ingestible by the locally deployed enterprise SIEM platform, such as:
    • Processed data
    • Generated alerts, including the captured data artifacts (raw data) from which the alert was generated without alterations of the data and within the retention limits established by federal, departmental, and agency logging requirements and policies.

Current CMS policy (AU-11) requires retention of online records (event data) for 90 days and archiving of older records for no less than one year. Retention periods may be customized by the SDM and defined within the SSP.

  • Support sharing of processed data through a standardized event format or an application programming interface in a manner that facilitates sharing with:
    • External collection and analysis application(s), such as another SIEM
    • External organizations, another federal agency, or even another SDM
  • Be configured to send unparsed data directly from hosts and sensors via if applicable, the enterprise logging tool’s universal forwarders or syslog:
    • Data should not be forwarded from any other analysis or logging tool, except in the case of syslog aggregators.
    • Communications between enterprise logging system endpoints (forwarders, indexers) must be encrypted using the encryption feature built into the enterprise logging system tool.
    • The anonymization of field data capability must be enabled for any sources containing PII.
    • Communications must be implemented using FIPS 140-2 or FIPS 140-3 validated modules to protect against the spillage of sensitive information/data:
      • Between local enterprise logging system components
      • Between SDM enterprise logging system instances and the enterprise logging system instance at the CCIC
      • Between data sources outside of the SDM enterprise logging tools supporting encrypted communications and the local enterprise logging tools
    • Signed certificates used within the enterprise logging system must come from the CMS Certificate Authority (CA). Self-signed certificates, such as those created by the enterprise logging tools, are readily available to the public and easy to decrypt
  • Be configured to ingest, aggregate, monitor, and analyze web application logs
  • Implement the necessary enterprise logging system indexers, forwarders, and search heads

Rationale:

A SIEM capability supports the automated interpretation of disparate data sources and log files that contain security- and privacy-relevant information collected within a security operations architecture. SIEMs use signatures and behavioral, heuristic, and other content analysis techniques to identify indicators of suspicious or known malicious behavior and activity within monitored systems.

Related CMS ARS Security Controls include: SI-4 - System Monitoring, IR-5 - Incident Monitoring, AU - Audit and Accountability family, CA-2 - Control Assessments, CA-7 - Continuous Monitoring, CA-7(4) - Risk Monitoring, CM - Configuration Management family, RA-5 - Vulnerability Monitoring and Scanning, RA-7 - Risk Response, SC-8 - Transmission Confidentiality and Integrity, SC-13 - Cryptographic Protection, and SC-17 - Public Key Infrastructure Certificates.

Advanced Investigation

The CCIC Advanced Investigation service oversees the handling of any incidents requiring digital forensics or malware analysis on affected system(s). The Forensic & Malware Analysis Team provides recommendations, technical expertise, and analysis of the collection of digital evidence to the CCIC, Marketplace SOC, and Supporting SOCs.

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-08: Local Forensic and Malware Analysis Support

The SDM must:

  • Coordinate incident response activities locally and collaborate with the CCIC, including assisting with:
    • Forensics and malware analysis
    • Incident response and threat analysis
    • Information security and privacy engineering and architecture
    • Obtaining ISPG approval for all SDM digital forensics processes and procedures
  • Coordinate between administrators handling malware detection and FMAT services
  • Comply with the CMS Digital Forensics Services (DFS) standard operating procedures to ensure the secure collection, transport, and analysis of digital forensic data, including:
    • Complete the required CMS Chain of Custody and CMS Evidence Collection forms
    • Provide a capability that collects volatile and non-volatile digital evidence and/or malware artifacts from identified targets in a forensically sound manner
  • Cooperate fully throughout any digital forensics investigation by providing access to/information on relevant IT resources as requested by CMS, HHS, and/or the Office of Inspector General (OIG)

The SDM must ensure malware detection and forensic acquisition solutions are:

  • Deployed on all applicable assets, including Network Security Endpoint Protection (NSEP), to support malware detection and forensic acquisition
  • Capable of:
    • Installation within the SDM data center secure enclave
    • Proactively monitoring (e.g., focused sweeps) its IT assets for IOCs and providing alerts
    • Actively participating in the CCIC IOC sharing program to share SDM locally developed IOCs with CCIC
    • Quickly identifying and isolating compromised systems
    • Deploying software agents on remote machines to perform data collections and acquisitions
    • User-level modifications and customizations, including:
      • Ingesting IOCs using the OpenIOC Extensible Markup Language format
    • Supporting modification of IOCs for variations, differences, or other aspects of the data center environment
    • Export digital evidence in CMS-approved standard file formats (i.e., raw, Expert Witness [E01], or Advanced Forensics Format [AFF])
    • Calculate checksums of the acquired image using CMS-approved standard approaches (i.e., Message Digest number 5 [MD5] or Secure Hash Algorithm 1 [SHA1]) for validation against original evidence

The SDM forensic acquisition solution(s) must:

The SDM malware analysis solution(s) must:

  • Perform analysis of malware threats in an isolated location (e.g., Malware Analysis Lab) using specialized software to:
    • Detect (identify suspected malware)
    • Detonate (extract and execute in a secured sandbox environment)
    • Analyze (identify anomalous behavior)
    • Report (provide an industry best practice-style finding report)

Rationale:

The SDM onsite forensic and malware analysis support, which acts as a sub-component of the CCIC FMAT, serves as the first responder for handling suspected incidents that require digital forensics or malware analysis on affected system(s) for confirmation. To this end, the SDM must work with ISPG to determine what tools are appropriate—and how the tools should be deployed. For example, ISPG will help the SDM determine when malware detection capabilities should be placed in-line to allow for active blocking of malicious activity or implemented using a tap / span methodology.

The FMAT requirements apply to NSEP deployments as well. NSEP functionality, which provides host-based security suite protection on endpoints (e.g., servers in the data center), protects endpoints from compromise by:

  • Detecting / blocking malicious behavior (e.g., host-based IDS, host-based firewalls, and application firewalls)
  • Managing system configurations and vulnerabilities
  • Providing malicious code detection, such as antivirus
  • Providing hardware and software asset management (including allowlisting / denylisting)
  • Encryption (i.e., via a FIPS 140-2 or FIPS 140-3 validated module)
  • Supporting advanced investigation

Related CMS ARS Security Controls include: AU-6 - Audit Record Review, Analysis, and Reporting, AU-7 - Audit Record Reduction and Report Generation, CA-2 - Control Assessments, CM-6 - Configuration Settings, IR-4 - Incident Handling, IR-7 - Incident Response Assistance, IR-8 - Incident Response Plan, IR-4(11) - Integrated Incident Response Team, RA-5 - Vulnerability Monitoring and Scanning, SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, SC-13 - Cryptographic Protection, SI-3 - Malicious Code Protection, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives, SI-7 - Software, Firmware, and Information Integrity.

Information Sharing and Cyber Threat Intelligence

The CCIC IS&CTI service improves the effectiveness of CMS information security and privacy by identifying CMS-specific threat indicators, actors, attack vectors, and breach scenarios to enhance situational awareness, incident detection, response, and coordination operations throughout the enterprise.

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-09: Local Information Sharing and Cyber Threat Intelligence Support

The SDM must:

  • Coordinate and collaborate IS&CTI activities locally and with the CCIC, including:
    • Incident response, response plans, and threat analysis
    • Threat feed and vulnerability / baseline configuration data
    • Prioritization of mitigations
    • Regular notification on response plan and mitigation status
    • Required submission of CyberScope data feeds
  • When applicable, coordinate and collaborate incident response activities locally and with the CCIC as part of the IMT
  • Provide required periodic monitoring and review of (frequency of review, including FISMA systems categorized under FIPS 199 as HIGH or MODERATE and FISMA systems identified by CMS as High Value Assets, is defined within the CMS ARS):
    • Vulnerability and configuration scan reports performed on ALL assets
    • Threat feed and vulnerability / baseline configuration information
    • Boundary protection devices (e.g., Firewall, IDS / IPS) logs and reports
    • Endpoint Protection (often listed as EPP) solution (e.g., AV / anti-malware, vulnerability and configuration scanners) logs, reports, and scan results
  • Ensure deployed IS&CTI capabilities provide:
    • Correlation between IS&CTI feeds and collected information (e.g., logs, reports, and scan results)
    • Analysis supporting prioritization of mitigation / remediation
  • Ensure corrective actions are managed through change management or a POA&M, as appropriate

The SDM IS&CTI solution(s) must:

  • Support assessment of business impact, development of mitigation plans, and implementation of the mitigation
  • Ingest and export threat feed data that conforms to CMS-approved standard formats (i.e., Structured Threat Information eXpression [STIX], Trusted Automated eXchange of Indicator Information [TAXII])
  • Conform to NIST SP 800-137, Information Security Continuous Monitoring, standards and procedures
  • Integrate IS&CTI activities with CMS’s Governance, Risk, and Compliance Management (GRC) tools

Rationale:

The SDM onsite information sharing and cyber threat intelligence support, which acts as a sub-component of the CCIC IS&CTI team, helps the CCIC improve the effectiveness of CMS information security and privacy capabilities by identifying CMS-specific threat indicators, actors, attack vectors, and breach scenarios to enhance situational awareness, incident detection, response, and coordination operations throughout the enterprise.

IS&CTI must coordinate and correlate with the SDM’s NSEP solutions, through functionality such as the host-based security suite providing protection to endpoints (e.g., servers in the data center), to reduce the risk of compromise by:

  • Detecting malicious behavior
  • Managing system configurations and vulnerabilities
  • Providing malicious code detection (e.g., AV)
  • Providing hardware and software asset management (including allowlisting / denylisting)
  • Supporting advanced investigation

Related CMS ARS Security Controls include: AU-6 - Audit Record Review, Analysis, and Reporting, AU-7 - Audit Record Reduction and Report Generation, CA-2 Control Assessments, CM-6 - Configuration Settings, IR-4 - Incident Handling, IR-7 - Incident Response Assistance, IR-8 - Incident Response Plan, IR-4(11) - Integrated Incident Response Team, PM-16(1) - Automated Means for Sharing Threat Intelligence, RA-5 - Vulnerability Monitoring and Scanning, SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, SC-13 - Cryptographic Protection, SI-3 - Malicious Code Protection, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives, SI-7 - Software, Firmware, and Information Integrity.

Penetration Testing

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-10: Penetration Testing Support

For FISMA systems categorized under FIPS 199 as HIGH or MODERATE (including FISMA systems identified by CMS as High Value Assets), the SDM must:

  • Participate in the CCIC Penetration Testing Program

For SDMs participating in the CCIC Penetration Testing Program, the SDM must:

  • Coordinate penetration test activities locally and collaborate with the CCIC to:
    • Provide access and security allowlisting for required CMS testing activities
    • Define the degree and nature of testing to be performed
  • Implement the following capability:
    • Web application inventory mechanism that incorporates applications into the CCIC Penetration Testing Program and identifies exclusions (including both the justification and approval for exclusion from testing)

Rationale:

CMS established a penetration testing capability within ISPG to proactively assess the susceptibility of CMS systems, personnel, and facilities to attack. This capability provides a robust, repeatable penetration testing capability with meaningful and usable results that complement (while avoiding duplication of) other forms of CMS information security testing.

Related CMS ARS Security Controls include: CA-8 - Penetration Testing, RA-5 - Vulnerability Monitoring and Scanning, SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, and SI-3 - Malicious Code Protection.

Security Architecture and Engineering

The CCIC Security Architecture and Engineering service provides architecture and design of security and privacy solutions across the CMS enterprise, including government and contractor facilities. The SAE capability applies a broad knowledge of CMS practices, security architectures, requirements, and solutions to strengthen the overall enterprise security posture. The CCIC SAE service provides an end-to-end view of security across the CMS enterprise to integrate all capabilities and functions effectively by designing solutions (cybersecurity capabilities) to solve evolving cybersecurity and privacy needs within the CMS TLC processes and procedures.

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-11: Local Security Architecture and Engineering Support

The SDM must ensure the availability of technically qualified staff, including trained information security and privacy personnel, to fulfill the roles and responsibilities of the SDM SAE. Small SDMs may choose to defer part or all required capabilities to the CCIC SAE via an appropriate Memorandum of Understanding / Memorandum of Agreement.

The SDM must:

  • Provide local security architecture and engineering support, consisting of:
    • Full understanding of the local architecture, network, and hardware / software inventory
    • Work with CCIC SAE to design local SDM capabilities to be integrated with CCIC
    • Oversee deployment of local SDM capabilities to integrate with CCIC
    • Troubleshoot integration problems after deployment
  • Deploy the following capabilities:
    • Network mapping and risk context analyzer including layer-3 device configuration analysis

Rationale:

The SDM onsite security architecture and engineering support, which acts as a sub-component of the CCIC SAE, serves as the key resource for helping CCIC SAE provide an end-to-end view of security across the CMS enterprise.

Related CMS ARS Security Controls include: CA-2 - Control Assessments, AT-2 - Literacy Training and Awareness, AT-3 - Role-Based Training, PL-4 - Rules of Behavior, RA-5 - Vulnerability Monitoring and Scanning, IR-4(11) - Integrated Incident Response Team, SI-4 - System Monitoring, AU - Audit and Accountability family

High Value Assets

A high-value asset is an asset used as a mission-critical information resource supporting infrastructure providers / suppliers or partnering organizations. The unauthorized disclosure of modification / destruction of, or disruption of access to information could be expected to have a severe or catastrophic adverse effect on organizational operations, organizational assets, or individuals.

Among other requirements, BR-CCIC-01, Security Authorization of Systems, says the business owner / information system owner / SDM must categorize the FISMA system in accordance with FIPS 199, and document the system attributes used to identify PII, PHI, and HVAs. The CIO determines whether a system is an HVA. If the CMS CIO identifies a CMS system as a High Value Asset, an HVA Designation Letter must be on file. Please refer to BR-CCIC-01 for more information.

CMS is required to comply with OMB Memorandum M-19-03, Strengthening the Cybersecurity of Federal Agencies by enhancing the High Value Asset Program, December 2018; CISA’s Binding Operational Directive (BOD) 18-02, Securing High Value Assets, May 2018; CISA High Value Asset Control Overlay, February 2018.

The CISA High Value Asset Control Overlay provides additional protections required for all HVA systems to increase the level of assurance that the HVA will meet the protection needs of the organization and help the federal government manage and reduce known risks to its most valuable assets. The enhanced controls assist in reducing the following risks:

  • Reducing available attack surface
    • Limit lateral movement (e.g., from adjacent components) by requiring segmentation and strict flow control
    • Limit internal connections
    • Coordinate auditing across organizations
  • Protecting against unauthorized access
    • Strict identity and account management practices
    • Reduce and limit access to elevated privileges
    • Enhance access controls on privileged accounts
  • Enhancing protection, control, and monitoring of data to be shared outside the HVA authorization boundary
    • Reduce risk of loss of confidentiality
    • Minimize data shared with known recipients (i.e., specific users, specific roles, and/or specific remote hosts)
    • Link monitoring and reporting to OMB Circular A-130 requirements
  • Consolidating and centralizing device audit and logging to improve monitoring, detection, and response
    • Coordinate across organizations
    • Streamline organizational response to incidents
    • Improve capabilities to detect threats
  • Holding contractors accountable and liable for implementation and effectiveness of implemented security and privacy controls
  • Protecting the acquisition supply chain for HVAs

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS that have been identified by the CMS CIO as an HVA. All systems identified as HVAs must not only meet or exceed the CMS minimal security and privacy control baselines for the FIPS 199-based security impact level (HIGH or MODERATE), but also implement the HVA security and privacy controls defined within DHS’s High Value Asset Control Overlay. When the implementation under the CMS ARS and the DHS HVA overlay differ, the most rigorous implementation is required.

BR-CCIC-26: High Value Assets

If a CMS system will be identified by the CMS CIO as a High Value Asset (i.e., HVA Designation Letter will be on file), the SDM must:

  • Meet security and privacy requirements defined by the CMS HVA Program
    • Meet or exceed applicable security and privacy control baselines for the system’s FIPS 199-based security impact level (i.e., HIGH or MODERATE)
    • Meet or exceed applicable security and privacy control requirements defined within DHS’s High Value Asset Control Overlay
  • Ensure all privacy documentation (e.g., PIA, SORN, and sharing agreements) is completed and approved by the CMS SOP

Rationale:

The SDM developing or maintaining a CMS HVA is required to meet emerging HVA requirements. To this end, the SDM must work with ISPG to determine what security and privacy controls and control enhancements are appropriate—and how the controls and control enhancements will be implemented. For example, ISPG will help the SDM determine when multiple authorizations are required and who is authorized to provide those authorizations before a user account will be created.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, AC-2 - Account Management, AC-3 - Access Enforcement, AC-4 - Information Flow Enforcement, AC-6 -Least Privilege, AC-17 - Remote Access, AC-20 - Use of External Systems, CA-5 - Plan of Action and Milestones, CA-6 - Authorization, CA-7 - Continuous Monitoring, CA-7(4) Risk Monitoring, CA-9 - Internal System Connections, AU-2 - Event Logging, AU-6 - Audit Record Review, Analysis, and Reporting, AU-9 - Protection of Audit Information, AU-10 - Non-Repudiation, AU-16 - Cross-Organizational Audit Logging, CM-5 - Access Restrictions for Change, CM-6 - Configuration Settings, CM-8 - System Component Inventory, PL-2 - System Security and Privacy Plans, PL-8 - Security and Privacy Architectures, PL-10 - Baseline Selection, SC-3 - Security Function Isolation, SC-5 - Denial-of-Service Protection, SC-7 - Boundary Protection, SC-8 - Transmission Confidentiality and Integrity, SC-18 - Mobile Code, SC-28 - Protection of Information At Rest, RA-1 - Policy and Procedures, IA-2 - Identification and Authentication (Organizational Users), IA-3 - Device Identification and Authentication, IA-5 - Authenticator Management, CP-8 - Telecommunications Services, CP-9 - System Backup, CP-9(8) Cryptographic Protection, CP-10 - System Recovery and Reconstitution, PM-7 - Enterprise Architecture, PM-9 - Risk Management Strategy, PM-10 - Authorization Process, PM-12 - Insider Threat Program, IR-4(10) - Supply Chain Coordination, IR-5 - Incident Monitoring, RA-5 - Vulnerability Monitoring and Scanning, RA-8 - Privacy Impact Assessments, RA-9 - Criticality Analysis, SA-4 - Acquisition Process, SA-9 - External System Services, SA-11 - Developer Testing and Evaluation, SI-2 - Flaw Remediation, SI-3 - Malicious Code Protection, SI-4 - System Monitoring, and PE-3 - Physical Access Control.

Information Security Continuous Monitoring / Privacy Continuous Monitoring

In compliance with federal and HHS requirements, CMS will perform periodic manual or automated audits, scans, reviews, or other inspections of the contractor’s IT environment that is used to provide or facilitate services for CMS.

CyberScope

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-12: CyberScope Data Feeds

Each SDM must:

  • Provision, secure, monitor, and maintain the hardware, network(s), and software supporting the ISCM capability infrastructure:
    • ISCM capability may be provided by CMS.
  • Provide required security and privacy continuous monitoring feeds that:
    • Comply with the CISA CyberScope reporting requirements
    • Automate CyberScope data feeds to the CCIC, as required

Where ISPG security tools do not cover automated security domains, SDMs must:

  • On a real-time basis, locally maintain tools to provide the necessary continuous monitoring feeds specified in NIST SP 800-137 and other emerging federal requirements

Rationale:

Recent updates to FISMA reporting guidance focus Agency resources on building the infrastructure and capabilities required to fulfill ISCM / PCM requirements. CISA and HHS have defined specific, automated feed requirements for both CyberScope reporting and improved situational awareness. In compliance with these requirements, CMS will perform periodic manual or automated audits, scans, reviews, or other inspections of the contractor’s IT environment that provides or facilitates services for CMS.

CyberScope currently requires automated scanning and feeds that report the status of asset management, vulnerability management, and configuration management across all assets. To facilitate the aggregation of this data across the numerous existing tools used by contractors, NIST defined Security Content Automation Protocol standards. SCAP includes formats for asset management (Common Platform Enumeration [CPE]), vulnerability management (Common Vulnerabilities and Exposures [CVE]), and configuration management (Common Configuration Enumeration [CCE]). These feeds must be collected in the Lightweight Asset Summary Results (LASR) schema and Extensible Configuration Checklist Description Format (XCCDF).

In addition to meeting all CMS information security and privacy requirements that are documented in the Information Security and Privacy Library, contractors must work closely with the CCIC and the CDM Program Team to implement the federally required CDM capabilities. CISA oversees the definition and implementation of the CISA Continuous Diagnostics and Mitigation (CDM) Program, including the program requirements and participation in the acquisition of tools, as outlined on the General Services Administration (GSA) in the GSA Continuous Diagnostics and Mitigation Tools information website. This topic emphasizes and repeats many of the CISA CDM requirements.

The CMS ARS requires an independent, third-party assessment of the SDM’s security controls to determine the extent to which security controls are implemented correctly, operating as intended, and producing the desired outcomes in meeting security requirements. CMS’s security assessment staff are available for consultation during the assessment process and will review the results before issuing an assessment and subsequent authorization recommendation. CMS ISPG reserves the right to verify the infrastructure and security test results before issuing an authorization recommendation. This recommendation will support the CIO in making the final authorization decision.

Related CMS ARS Security Controls include: CA-7 Continuous Monitoring, CA-7(4) - Risk Monitoring, RA-5 - Vulnerability Monitoring and Scanning, SI-4 - System Monitoring, AU - Audit and Accountability family, CM-8 - Information System Component Inventory, CA-2 - Control Assessments, PM-27 - Privacy Reporting, and PM-21 - Accounting of Disclosures.

CDM Asset Management

CDM’s Asset Management (formerly Phase 1) addresses endpoint integrity through management of hardware and software assets, configuration management, and vulnerability management. These capabilities are viewed as foundational capabilities in the protection of systems and data.

The business rules within this topic apply to all FISMA system and System Developer and Maintainer (SDM) data centers supporting CMS.

BR-CCIC-13: Local ISCM / CDM Management Capability

The SDM must:

  • Provide centralized management and configuration of the content captured from individual components of the information system under ISCM / CDM
  • Provide centralized content repositories, made available to the CCIC, whereby:
    • Information, including system and network monitoring data, is available to the CCIC in a format compliant with CMS and federal (e.g., CDM) requirements
    • Centralized content repository record sources include systems, appliances, devices, services, and applications (including databases)
    • As required by the CMS ARS, raw content must be available in an unaltered format to the CCIC
    • Records are maintained for ninety (90) days and archives of old records for one (1) year to provide support for analysis, incident response, management, disposition, and after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements
  • Assess or support the assessment by third parties of all security controls to verify the security controls are adequately designed, effectively implemented, and appropriately documented.

The term ‘tailorable’ as used in this document references CMS ARS Supplemental controls and control enhancements that CMS encourages Business Owners to consider. The CMS ARS states that many of the mandatory and supplemental controls are customizable (i.e., tailorable) by the Business Owner.

Rationale:

Effective ISCM is the implementation of a strategy that addresses ISCM requirements and activities at each organizational tier (organization, mission / business processes, and information systems). Each organizational tier centrally monitors security metrics and assesses security control effectiveness with established monitoring and assessment frequencies and status reports customized to support tier-specific decision making. Policies, procedures, tools, and templates are implemented at each organizational tier and are managed in accordance with guidance from higher-level organizations in a manner that supports shared use of data within and across tiers.

ISCM / CDM Capability Management is the centralized management, including O&M, of security functions that support the federal government’s ISCM program.

Related CMS ARS Security Controls include: CA-7 - Continuous Monitoring, CA-7(4) - Risk Monitoring, RA-5 - Vulnerability Monitoring and Scanning, SI-4 - System Monitoring, SI-2 - Flaw Remediation, AU - Audit and Accountability family, CM-8 - Information System Component Inventory, and CA-2 - Control Assessments.

BR-CCIC-14: Hardware Asset Management Capability

The SDM must:

  • Provide a managed and tailorable hardware asset management (HWAM) capability to protect endpoint configurations
  • Manage hardware assets (such as authorized assets connected to CMS networks and unauthorized devices disconnected from the network) automatically, according to CMS- and HHS-defined guidance and requirements

The SDM HWAM capability must:

  • Prevent / reduce damage from unauthorized hardware gaining access to the network, as defined by the SDM and CMS
  • Detect and relay hardware inventory and provide a means to prevent access to the network by unauthorized hardware
  • Provide management and reporting of changes
  • Provide status feeds to the CCIC in a format compliant with CMS and Federal (e.g., Continuous Diagnostics and Mitigation) requirements
  • Assign risk to each difference between the authorized hardware inventory baseline and the actual hardware inventory, based on relevant factors
    1. Scores for all authorized and unauthorized hardware.
    2. When applicable, standard federal score for deviations from authorized hardware inventory baseline.
    3. CMS, HHS, or SDM scores for deviations from authorized hardware inventory baseline, provided CMS has adopted a policy that permits the difference from applicable corresponding federal policy.
  • Maintain records for ninety (90) days and archives old records for one (1) year to provide support for analysis, incident response, management, disposition, and after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements by:

The HWAM tool must score (assign a numerical value to) deviations for purposes of computing risk. Scoring includes:

  1. Determining[JD20]  deviation between the authorized hardware inventory and the actual hardware inventory
  2. Ensuring deviations are identified and include both missing and newly identified (i.e. previously unknown / unidentified) hardware assets
  • Report anomalies when deviations between the authorized hardware inventory baseline and the actual hardware inventory are detected and:
  1. Remove[JD21]  (disallow access) or authorize the unauthorized hardware device
  2. Bring the unauthorized asset back into operation, record its non-operational status (with rationale), or de-authorize the hardware when the difference indicates authorized hardware that is missing in operation

To support the HWAM capability, the SDM must:

  • Identify, review, update, and address the status of all devices within the SDM’s enterprise / boundary (i.e., maintain the actual hardware [HW] inventory) :
    • Create, operate, and maintain an inventory of authorized hardware, including perimeter and endpoint devices, in near real time, along with applicable information required to accurately assess risks to the asset
    • Update the inventory regularly with automated hardware discovery processes and tool(s) no less than once every 72 hours.
  • Enable and revoke access, manually and automatically, for perimeter and endpoint devices that connect to the SDM network based on knowledge of the asset and authorization (e.g., grant the asset access to the network when authorized or block access to the SDM network when not authorized)

Rationale:

Once unauthorized or unmanaged hardware is discovered, the SDM will either remove or authorize and manage this hardware. Since unauthorized hardware is unmanaged, it must be assumed to be vulnerable and can be exploited as a pivot to other assets if not removed or managed.

The HWAM capability is achieved through the creation, operation, and maintenance of an authorized hardware inventory baseline that assigns unique identifiers for hardware and tracks key contact properties such as the manager or custodian of the hardware. General requirements for HWAM from CISA include:

  • The process for generating the authorized hardware inventory baseline must be established, operated, and maintained in environments already in operation, following CMS and HHS-defined guidance and requirements.
  • To provide this functionality, SDMs should not assume existing CMS or SDM property management capabilities can assist in generating the authorized hardware inventory baseline. Where available, property management systems should be considered as sources of input for the authorized hardware inventory baseline.
  • The authorized hardware inventory baseline generated by this process must be complete, accurate, and timely. For this purpose, “timely” is at least once every 72 hours.

Related CMS ARS Security Controls include: CM-8 - Information System Component Inventory.

BR-CCIC-15: Software Asset Management Capability

The Software Asset Management (SWAM) capability inventories all software configuration items (SWCI) in IT assets on a network and provides oversight.

The following requirements apply to all CMS FISMA system and SDM data centers supporting CMS.

The SDM must:

  • Provide a managed and tailorable SWAM capability to provide a full inventory of software assets on any FISMA system
  • Manage and dispose of software (i.e., authorize software to operate on or be removed from the network) manually or automatically, according to CMS- and HHS-defined guidance and requirements

The SDM SWAM capability must:

  • Prevent / reduce damage from and/or execution of malware and/or unpatched software assets, as defined by the SDM and CMS
  • Detect and report malware (including, as configured, non-patched but necessary software assets) at a rate comparable to existing AV products and provide a means for removing malware in time to prevent it from executing
  • Provide management and reporting of current changes and software installation actions
  • Provide status feeds to the CCIC in a format compliant with CMS and Federal (e.g., Continuous Diagnostics and Mitigation) requirements
  • Assign risk to each difference between the authorized software inventory baseline and the actual software inventory, based on relevant factors
    1. Scores for all authorized and unauthorized software based on patch level, maintenance / support, configuration, and known vulnerabilities.
    2. When applicable, standard federal score for deviations from authorized software inventory baseline.
    3. CMS, HHS, or SDM scores for deviations from authorized software inventory baseline, provided CMS has authorized the deviation from applicable CMS or federal policy.
  • Maintain records for ninety (90) days and archives old records for one (1) year to provide support for analysis, incident response, management, disposition and after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements by:

The SWAM tool must score (assign a numerical value to) deviations for purposes of computing risk. Scoring includes:

  1. Determining[JD22]  differences between authorized software inventory and actual software inventory
  2. Ensuring differences identified include expected (i.e., known, authorized), unexpected (i.e., unknown, unidentified), and missing software assets
  • Report anomalies when installed software assets are not showing on the provided inventory

To support the SWAM capability, the SDM must:

Identify, review, update, and address the status of all software within its enterprise / boundary (i.e., maintain the actual software inventory) as authorized / unauthorized, managed / unmanaged, or reporting / non-reporting:

  • Create, operate, and maintain its inventory of authorized and unmanaged software in near real time, along with applicable information required to accurately assess risks to the asset and physically locate the asset.
  • Generate/update inventory data that is complete, accurate, and timely. For this purpose, “timely” means no less often than once every 72 hours.
  • Assess all software assets on all perimeter and endpoint devices within its network no less often than once every 72 hours / three (3) calendar days using an automated mechanism.
  • Enable and revoke access to any software asset, manually and automatically, for perimeter and endpoint devices that connect to the SDM network, based on knowledge of the asset and authorization (e.g., grant execution when authorized or block execution when not authorized).

CMS does not specify which SWAM tool a SDM may choose. If a SDM chooses a different tool, the SDM is responsible for ensuring compatibility and interoperability with the CCIC.

Rationale:

Once unauthorized or unmanaged SWCIs are discovered by the contractor’s provided tool(s), the contractor will act to remove these SWCI. Because unauthorized software is unmanaged, it is vulnerable to exploitation as a pivot to other IT assets if not removed or managed. In addition, a complete, accurate, and timely software inventory is essential to support awareness and effective control of software vulnerabilities, security configuration settings, and licensing efficiencies.

The SWAM capability is achieved through the creation, operation, and maintenance of an authorized software inventory baseline that assigns unique identifiers for software and tracks key contact properties such as the manager or custodian of the software. General requirements for SWAM from CISA include:

  • The process for generating the authorized software inventory baseline must be established, operated, and maintained in environments that are already in operation, following CMS- and HHS-defined guidance (e.g., business rules) and requirements.
  • It should not be assumed that CMS and SDM property management functions will assist the process for generating the authorized software inventory baseline; however, where property management systems are available, the property management systems should be considered as sources of input for the authorized software inventory baseline.
  • The authorized software inventory baseline generated by this process must be complete, accurate, and timely. For this purpose, “timely” is at least once every 72 hours.

Related CMS ARS Security Controls include: CM-7 - Least Functionality, CM-7(4) - Unauthorized Software - Deny, and CM-7(5) - Authorized Software - Allow.

BR-CCIC-16: Configuration Settings Management Capability

The CSM capability reduces misconfiguration of IT assets, including misconfigurations of hardware devices (i.e., physical, virtual, and operating system) and misconfigurations of software. Deviations and changes in configuration may make the device or software either more secure or less secure. The CSM tool is not responsible for distinguishing the difference.

The following requirements apply to all CMS FISMA system and SDM data centers supporting CMS.

The SDM must:

  • Provide a managed and tailorable CSM capability to protect endpoint configurations
  • Manage a change control process to document CMS and SDM extensions or exceptions to the authorized core federal benchmarks
  • Implement a CMS-authorized security configuration benchmark as defined within the ARS: The baseline consists of the acceptable value(s) for each relevant configurable setting for each IT asset type.
    • Establish a core benchmark, based on CMS / HHS / federal guidance (e.g., the CMS ARS) and requirements, for hardware devices and software with pre-programmed tests for actual status and an authoritative system, based on federally approved scoring methodologies and standards, to provide core scores for risk assessment
    • Adopt the federal benchmark and/or:
      • Add, modify, or delete non-core configuration settings with review and approval from ISPG
      • Assign an alternate method to score risk for common and CMS / SDM-specific settings for internal SDM use
    • Develop an internal benchmark tool using pre-programmed tests to evaluate actual status and an authoritative system, using federally approved scoring methodologies and standards, to provide core scores for risk assessment

Each configuration baseline setting can be traced to an authoritative reference, such as US Government Configuration Baseline (USGCB), DISA STIG, etc.

The SDM CSM capability must:

  • Prevent / reduce risk of damage from misconfiguration
  • Detect and report changes in configuration (i.e., from the defined baseline) at a rate comparable to existing AV products and provide a means for remediation of unauthorized changes
  • Provide management and reporting of configuration settings and changes
  • Provide status feeds to the CCIC in a format compliant with CMS and Federal (e.g., Continuous Diagnostics and Mitigation) requirements
  • Maintain records for ninety (90) days and archives old records for one (1) year to provide support for analysis, incident response, management, disposition, and after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements
  • Assign risk to each difference between the authorized configuration setting (i.e., the baseline) and the actual configuration setting, based on relevant factors: The CMS tool must score (assign a numerical value to) deviations for purposes of computing risk. Scoring includes:
    1. Scores for settings based on patch level, maintenance / support, configuration, and known vulnerabilities.
      1. When applicable, standard federal score for deviations (e.g., Common Configuration Scoring System [CCSS]) from the authorized software inventory baseline.
      2. CMS, HHS, or SDM scores for deviations from authorized baseline values, provided CMS has authorized the difference from applicable CMS or federal policy.
    2. Maintain and ensure collected data is used for purposes of analysis, incident response, investigation, management, and disposition by determining differences between authorized configuration and actual configuration
  • Report anomalies when differences between the authorized configuration baseline and the actual configuration are detected

To support the CSM capability, the SDM must:

  • Identify, review, and update the status of configuration settings for all monitored devices within its enterprise:
    • Create, operate, and maintain the SDM’s information on deviations from baseline configurations in a timely manner [at least once every 72 hours / three (3) calendar days] along with information needed to assess the risk associated with the deviations
    • Update the baseline deviations regularly with automated software discovery processes and tool(s)
    • Generate inventory data that is complete, accurate, and timely. For this purpose, “timely” means no less often than once every 72 hours.
  • Assess all configuration settings on all perimeter and endpoint devices within its network no less often than once every 72 hours / three (3) calendar days, using an automated mechanism
  • Address configuration setting deviations from a defined baseline, manually and automatically, for perimeter and endpoint devices that connect to the SDM network, based on knowledge of the asset and authorization (e.g., deviation is granted)

Although CMS does not specify which CSM tool a SDM may choose, the solution should meet the technical, operational, and reporting requirements set forth in this document. The SDM is responsible for ensuring compatibility and interoperability with the CCIC.

Validation of CSM Compliance

To support CMS’s tool-agnostic approach to CSM for SDMs, CMS has chosen to follow an approach which decouples the process of an SDM implementing and operating an automated CSM tool versus CMS’s process of validating an SDM’s CSM compliance to an identified standard using a common tool.

  • The SDM will conduct timely CSM compliance scans using CMS’s supplied solution
  • The SDM scans will use a CMS identified DISA STIG baseline and CMS supplied audit files which support FISMA tagging
  • The SDM will create a CSM repository that can be peered by CMS’s instance

Rationale:

Once a misconfiguration of hardware or software is discovered by the provided tools, the SDM is responsible for taking any needed action to resolve the problem or request that CMS accept the risk upon appropriate risk analysis and documentation in conjunction with the ISSO.

CISA estimates that over 80 percent of known vulnerabilities are attributed to misconfiguration and missing patches. Cyber adversaries often use automated computer attack programs to search for and exploit IT assets with misconfigurations, especially for assets supporting federal agencies, and then pivot to attack other assets.

Related CMS ARS Security Controls include: CM-6 - Configuration Settings and CM-7 - Least Functionality.

BR-CCIC-17: Vulnerability Management Capability

The vulnerability management (VUL) capability, including exposure-based vulnerability management, discovers and supports remediation of vulnerabilities in IT assets on a network. This includes, but is not limited to, AV, host, and application scanning (credentialed and non-credentialed); host-based IDS / IPS; network vulnerability and risk analysis; and web application scanning.

CMS primarily uses a standard vulnerability assessment tool to perform vulnerability management and continuous risk monitoring at CMS. The tool leverages a distributed architecture that allows CMS to conduct efficient and nonintrusive vulnerability assessments with very low impact to network or host resources (an executable exists on the target of the scan for the duration of the scan). In addition, by continuously updating its host / vulnerability information, the CDM Program team can provide system administrators, ISSOs, and system owners with current, actionable information without the need for ad hoc scanning.

The following requirements apply to all CMS FISMA system and SDM data centers supporting CMS.

The SDM must:

  • Provide managed and tailorable vulnerability management capabilities to protect endpoints, including:
    • AV
    • HIDS / Host-based Intrusion Prevention System (HIPS)
  • Manage a change control process to document CMS and SDM extensions or exceptions to the authorized core federal benchmarks

The SDM VUL capability must:

  • Deploy a vulnerability assessment tool to assess the presence of known vulnerabilities and prevent / reduce risk of damage from known vulnerabilities
  • Detect and report the status of known vulnerabilities at a rate comparable to existing AV products and provide a means to prioritize required remediation in compliance with the CMS ARS
  • Provide management and reporting the existence of known vulnerabilities
  • Where the CMS standard vulnerability assessment tool is deployed: create and maintain a repository for VUL data that can be peered by CMS’s instance ; or when an alternate tool is used, provide status feeds to the CCIC in a format compliant with CMS and Federal (e.g., Continuous Diagnostics and Mitigation) requirements
  • Assign risk to each vulnerability based on relevant factors
    1. Scores for vulnerabilities based on patch existence, maintenance / support, and known threats exploiting the vulnerabilities.
    2. When applicable, standard federal score for vulnerabilities (i.e., Common Vulnerability Scoring System).
    3. CMS, HHS, or SDM scores for vulnerabilities, provided CMS has authorized the deviation from applicable federal values.
  • Maintain records for ninety (90) days and archives old records for one (1) year to provide support for analysis, incident response, management, disposition, and after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements, by:

The VUL tool must score (assign a numerical value to) vulnerabilities detected for purposes of computing risk. Scoring includes:

  1. Determining[JD23]  the presence of known vulnerabilities within a system or application
  2. Updating vulnerability definitions within one (1) week of the update release
  • Report anomalies when vulnerabilities are detected

To support the VUL capability, the SDM must:

  • Identify, review, and update the status of the presence of vulnerabilities for all monitored devices within its enterprise and do the following:
    • Create, and maintain the SDM’s information on the presence of vulnerabilities in a timely manner (at least once every 72 hours / three [3] calendar days), along with information needed to assess the risk associated with the vulnerabilities
    • Update the presence of vulnerabilities regularly with automated software discovery processes and tool(s)
    • Generate inventory data that is complete, accurate, and timely. For this purpose, “timely” means no less often than once every 72 hours.
  • Assess the presence of known vulnerabilities on perimeter devices within its network no less often than once every 72 hours / three (3) calendar days, using an automated mechanism
  • Address identified vulnerabilities, manually and automatically, for perimeter devices that connect to the SDM network, based on knowledge of the asset and authorization (e.g., deviation granted)

To support CCIC-based vulnerability scanning, the SDM must:

  • Facilitate the deployment and maintenance of the CMS standard suite of tools and meet the following requirements:
    • Follow the procedures and documentation provided by CMS in onboarding the CMS standard solution
    • If required, create a repository for VUL data that can be peered by the CMS instance of the tool
    • Provide the necessary infrastructure to host virtual device profilers (DP) provided by ISPG or work with ISPG to determine if a physical DP is required for scanning
    • Make all necessary firewall and routing changes to all machines that contain or transfer CMS data to allow for authenticated scanning
  • Manage a change control process to document CMS and SDM extensions or exceptions to the authorized core federal benchmarks

SDMs with public-facing websites must:

  • Reduce the risks associated with web-based applications used within CMS through ongoing security testing that facilitates the actionable identification, remediation, and elimination of vulnerabilities and defects
  • Control costs associated with web-based applications used within CMS through the identification, remediation, and elimination of vulnerabilities and defects in development and quality assurance environments
  • Provide a cost-effective monitoring and compliance capability for the security programs associated with web-based applications used within CMS
  • Manage regulatory demands associated with web-based applications used within CMS for protecting sensitive data processed by web-based applications

Although CMS itself uses a specific module to provide AV capabilities, CMS does not specify use of an AV application within the SDM. The SDM may choose any AV so long as the application receives updates regularly (i.e., in compliance with CMS ARS update requirements) and can provide CMS with standards-based status feeds.

CMS does not specify HIDS / HIPS tools for use by SDMs so long as the application can provide CMS with standards-based status feeds.

Rationale:

Vulnerability management is the management of risks presented by known software weaknesses that are subject to exploitation. The VUL function ensures that configuration errors and deficiencies are identified and remediated quickly. (Information security vulnerabilities are deficiencies in software that a hacker can use directly to gain access to a system or network.) Once the VUL tool(s) identifies these configuration errors and deficiencies, CMS and its SDMs act to remove or remediate these from operational systems to ensure the vulnerabilities can no longer be exploited.

The VUL capability scans CMS assets and SDM assets used to support CMS for known vulnerabilities, such as missing patches and improper access controls, and software weaknesses by:

  • Discovering and locating known security vulnerabilities
  • Discovering and locating software weaknesses in network configuration, software applications, and source code
  • Supporting awareness and understanding of exposure risks associated with software weaknesses

Related CMS ARS Security Controls include: RA-5 - Vulnerability Monitoring and Scanning, RA-7 - Risk Response, CA-7 - Continuous Monitoring, and CA-7(4) - Risk Monitoring.

SOC as a Service

CMS data center operators may use remote SOC as a Service (SOCaaS) services as part of a comprehensive plan that meets all the requirements discussed in this CCIC Integration chapter. The CMS CCIC may also offer to provide SOC services for CMS data centers that do not operate their own in-house SOC.

 

Perimeter Protections

A Perimeter Protections capability examines all cybersecurity and privacy events to identify any impact to CMS assets or information. To achieve this goal, this service provides standardized tools and techniques to monitor network perimeters and internal traffic flows for lateral communication.

Perimeter Monitoring

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-18: Perimeter Monitoring Prerequisites

The SDM must:

  • Comply with CMS ARS Security Control SI-4 - Information Security Monitoring requirements:
    • Perform real-time monitoring of traffic that traverses the CMS secure enclave/boundary, including any ingress or egress traffic that crosses an SDM’s internal FISMA system security boundary
    • Identify the most effective monitoring points on the network to monitor perimeter traffic traversing the CMS secure enclave/boundary
  • Ensure that the monitoring points (e.g., network taps):
    • Provide visibility to traffic going from the internal network to the public Internet and from the public Internet into an internal network
    • Implement a capability to decrypt encrypted traffic to ensure proper inspection of all traffic traversing the SDM boundary
    • Allow for source attribution to the originating host prior to any network address translation (NAT) or port address translation (PAT)
    • Support the core competencies of full packet capture, IDS / IPS, and web- and email-based malware detection
  • Adjust monitoring boundaries as new points of ingress/egress are created and/or identified:
    • For new points of ingress / egress, adjustment for monitoring boundaries before transitioning the system into production
    • For previously unknown points of ingress / egress, the adjustment window for revising monitoring boundaries is defined by the CMS POA&M process.

To support the perimeter monitoring capability, the SDM must:

  • Identify, review, update, or remove monitoring techniques based on analysis of network traffic traversing the environment and do the following:
    • Create, test, and verify newly developed techniques for deployment on network sensors and other perimeter monitoring tools
    • Maintain and support existing techniques for deployment on network sensors and other perimeter monitoring tools
    • Address intrusion attempts, suspicious incidents, and other daily cyber operations activities

Rationale:

The FISMA system (or SDM) must implement and support a technical capability to perform full packet capture and analysis of network traffic traversing the perimeter of the data center as it applies to the CMS secure enclave/boundary. The capture and analysis are in support of incident detection and investigations.

Related CMS ARS Security Controls include: SI-4 - System Monitoring; SI-3 - Malicious Code Protection; SC-7 - Boundary Protection; AC-18 - Wireless Access; AT-2 - AT-2(3) - Social Engineering and Mining, Literacy Training and Awareness, AT-3 - Role-Based Training, PL-4 - Rules of Behavior; CM-7 - Continuous Monitoring; PM-13 - Security and Privacy Workforce; PM-14 - Testing, Training, and Monitoring;

Full Packet Capture / Inspection

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-19: Full Packet Capture Capability

The SDM-deployed full packet capture (PCAP) capability must:

  • Capture and export large amounts of raw network packets in a PCAP-compatible format that must:
    • Be implemented via a network tap or a span port on the infrastructure router (A network tap is the preferred method because it provides the most reliable visibility into network traffic. Span ports are only used in cases where a tap cannot be deployed for technical reasons.)
    • Capture network packets at line speed
    • Ensure the point of monitoring connection provides access to unencrypted network traffic either prior to any Transport Layer Security (highest available level) encryption or as part of a solution to decrypt TLS traffic.
    • Ensure packet data storage meets or exceeds CMS data retention requirements
    • Organize captured packets in a logical manner to facilitate the ability to quickly store, search, and retrieve data
  • Process the packets at the protocol level and must:
    • Reconstruct sessions for common protocols such as web, email, File Transfer Protocol, etc.
    • Navigate packet / session information via graphical user interface (GUI)
    • Support deep analytics of raw packets and automated reporting
  • Provide the CCIC with access to both the packet capture and network analysis tool(s) as required

Rationale:

The FISMA system (or SDM) must implement and support a technical capability to perform full packet capture and analysis of network traffic traversing the perimeter of the data center as it applies to the CMS secure enclave/boundary. The capture and analysis are in support of incident detection and investigations.

Related CMS ARS Security Controls include: SI-4 - System Monitoring, SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, AC-4 - Information Flow Enforcement, AC-18 - Wireless Access, and AU-2 - Event Logging.

Network Intrusion Detection / Prevention

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-20: Network Intrusion Detection / Prevention Capability

The SDM-deployed network intrusion detection / prevention capability must:

  • Support the loading of customized signatures generated by either the SDM security personnel or CCIC, based on CCIC-supplied IOCs
  • Generate alerts that can be consumed as data feeds into the CMS enterprise logging tool
  • Implement a solution that provides the key prerequisites defined under Perimeter Monitoring:
    • Provide visibility into traffic traversing the perimeter
    • Provide source attribution to the originating host prior to any NAT or PAT
    • Decrypt encrypted traffic to ensure proper inspection of all traffic traversing the SDM boundary

Rationale:

Network-based IDS / IPS provides tunable, automated alerting capabilities; however, the network-based IDS / IPS does require a moderate degree of management and maintenance to ensure the alerts triggered by the implemented rules are not largely false positives (“noise”) that result in wasted security investigation efforts.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection, SI-4 - System Monitoring, and SC-7 - Boundary Protection.

Malware Detection / Prevention

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-21: Malware Detection / Prevention Capability

The SDM must:

  • Route inbound and outbound Internet traffic containing CMS data through a federally approved TIC service
  • Deploy web-based and email malware analysis, detection, and prevention capabilities when end user platforms (such as servers, appliances, devices, workstations, etc.) can connect to non-CMS networks without first passing through a TIC. These capabilities must:
    • Support auto-configured test environments to safely execute and inspect suspected malware payloads
    • Deploy at strategic network ingress/egress points to ensure complete coverage of both enterprise and enclave perimeters
    • Support collection of data for analysis through either a network tap or a span port on the infrastructure router(s); network taps are preferred, which provide more reliable visibility into network traffic
    • Span ports may only be used in cases network taps cannot be deployed.
      • Span port use must be documented and approved by CMS.
  • Support detection and blocking of detected threats, including:
    • Advanced zero-day exploits
    • Outbound callbacks

The SDM-deployed web-based malware analysis, detection, and prevention capability must:

  • Dynamically generate actionable malware intelligence, including:
    • Detect and alert on web-based attacks:
      • Leverage both signature and signature-less engine(s) for detection
      • Support automated response (i.e., active blocking that is capable of automatically responding to and stopping an attack) to web-based attacks
    • Automated response (active blocking) must be coordinated with all stakeholders (i.e., ISSO, SDM, ISPG, and business owner) and documented before activation
    • Forward alert, security-related, and log data to the CMS enterprise logging tool
    • Monitor assets without reliance on deployed software agents

The SDM-deployed email malware analysis, detection, and prevention capability must:

  • Dynamically generate actionable malware intelligence, including:
    • Detect and alert on email-based attacks:
      • Leverage both signature and signature-less engine(s) for detection
      • Support automated response (i.e., active blocking that is capable of automatically responding to and stopping an attack) to email-based attacks
    • Automated response (active blocking) must be coordinated with all stakeholders (i.e., ISSO, SDM, ISPG, and business owner) and documented before activation
    • Forward alert, security-related, and log data to the CMS enterprise logging tool
    • Monitor assets without reliance on deployed software agents
    • The point(s) of monitoring for the SDM malware analysis, detection, and prevention solution sensors must provide access to unencrypted network traffic, either prior to any TLS (highest available level) encryption or as part of a solution to decrypt TLS traffic.

Rationale:

The FISMA system (or SDM) must implement and support a technical capability to capture traffic traversing from systems within the FISMA system enclave(s) and detect malware within the content.

Related CMS ARS Security Controls include: SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, SC-07(03) - Access Points, AC-17 - Remote Access, AC-17(3) - Managed Access Control Points, IR-4 - Incident Handling, IR-4(11) - Integrated Incident Response Team, SI-3 - Malicious Code Protection, SI-4 - System Monitoring.

Network Firewalls

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-22: Network Firewall Perimeter Requirements

As described in the CMS TRA Multi-zone architecture , the firewalls may be virtual or physical and may be part of a vendor service when the systems are implemented within the cloud (e.g. AWS Virtual Private Cloud (VPC)). The guidance below applies to both CMS data center and cloud implementations.

The SDM must:

  • As detailed in business rule BR-CCIC-04, implement and maintain secured, separate enclaves or zones of systems and devices to implement the CMS-defined security architecture
  • Incorporate a perimeter protections solution that includes:
    • Stateful inspection / application firewalls (i.e., full packet inspection) (as a minimum)
    • Configure firewalls for least privilege—block, deny, or drop by default (deny all, allow by exception)
    • Logging actions and/or other policy / rule hits
  • Collect and analyze perimeter protections solution logs and security-relevant data, as described in Appendix A
    • Facilitate and maintain any firewall or routing changes necessary for security-relevant data and log collection

Rationale:

The FISMA system and SDM maintain secured, separated enclaves or zones of systems and devices as defined by the CMS TRA to provide the proper oversight, monitoring, and incident response capabilities. This requires the implementation of capabilities to protect the secured enclaves or zones within the CMS processing environment. The boundary separating each zone or enclave is protected by firewalls configured to deny by default and only allow by exception.

Each FISMA system and SDM is responsible for maintaining appropriate security and access control of its secure enclaves or zones and for implementing the appropriate tools and technologies to meet CMS and federal requirements.

Related CMS ARS Security Controls include: SC-7 - Boundary Protection, AC-4 - Information Flow Enforcement, AC-18 - Wireless Access, AC-20 - Use of External Systems, AU-6 - Audit Record Review, Analysis, and Reporting, CM-7 - Least Functionality, and CA-2 - Control Assessments.

Network Data Loss Prevention

Data Loss Prevention (DLP) provides consistent protections to block exfiltration of sensitive (especially privacy) data outside the organizational boundary as well as prevent inappropriate use (i.e., outside a documented routine use). DLP capabilities and functions include the following:

  • Interpretation of System-Readable Policies and Formalized Connection Agreements – the ability for the DLP solution to instantiate security and privacy rules that will ensure compliance with cognizant Privacy Notices and applicable laws and regulations
  • Exfiltration Alerts and Prevention – the ability for the DLP solution to monitor for data movement between systems initiated by a user, or even the system itself, and restrict or limit movement of the data based on rule sets or behavioral patterns. The solution is expected to report and alert on deviation from the defined boundaries or thresholds.
  • Protection Orchestration – The ability for the DLP solution to interoperate within a suite of tools or be integrated within a Security Information and Event Management infrastructure that provides the required access control and monitors functionality across systems.

Some DLP solutions further enhance available protections for sensitive information, such as PII, by providing enhanced monitoring and recognition of sensitive information traversing interconnections to ensure:

  • The exchange is restricted to authorized information
  • The exchange is restricted to authorized purposes
  • The exchange is restricted to authorized entities/users

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-27: Network Data Loss Prevention

The SDM must:

  • Support compliance with regulations and mandates (e.g., Privacy Notice(s), CMS policy, and interconnection agreements) on sharing of CMS sensitive information on authorized endpoints and interconnections
  • Provide audit trail information related to execution of DLP methods and the movement of data in a format compliant with CCIC needs
  • Perform DLP using one or more of the following methodologies:
    • Encryption
    • Quarantine
    • Block
    • Notification
    • Allow with user justification
  • Provide monitoring or DLP capabilities in accordance with one or more of the following:
    • Network monitoring, alerting on, and preventing the movement of data over the network using various network protocols (e.g., email, web, file transfer, and instant messaging) to detect or prevent data exfiltration

Rationale:

The DLP capability provides automated protections for sensitive information that can reduce the risk of unauthorized access. DLP will use fine-grained access control methodologies to restrict access to information to ensure that only authorized users, roles, and/or systems may access the information. DLP solutions can also provide prevention countermeasures for data in transit that can minimize the risk from unauthorized exfiltration (i.e., transmission of sensitive data to an unauthorized recipient).

Related CMS ARS Security Controls include: AC-3 - Access Enforcement, AC-4 - Information Flow Enforcement, CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, and SC-7 - Boundary Protection.

Lateral and Endpoint Protections

This topic describes the requirements for protections at the internal endpoints (IT assets within the data center / enclave) using network security endpoint protections, the proper use of encryption, and protections from insider threats.

Network Security Endpoint Protection

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-23: Network Security Endpoint Protection Capability

The SDM must:

  • Implement and support a technical capability for performing network security endpoint protection of assets within the SDM’s data center (including assets within the data center secure enclave):
    • Establish a process / procedure (including defining required user and system accounts) for performing:
      • Continuous AV and anti-malware scanning
      • Periodic and scheduled sweeps of the environment as defined within the CMS ARS
      • Ad-hoc sweeps in response to notifications, alerts, and incident response direction from CMS
      • Forward collected (e.g., scan and sweep) data / results to the local SIEM system
      • Host-based IDS / host-based IPS
      • Encryption capabilities (e.g., the capability to encrypt endpoint storage and communications)
  • Ensure time is synchronized across SDM assets with the CMS-approved time servers
  • Customize received IOCs to address variations, differences, or other unique aspects of the data center environment
  • Use the NSEP tool as a primary or backup mechanism for forensic data acquisition on the platforms on which the NSEP tool can be installed

The SDM network security endpoint protection capability must:

  • Provide confidentiality and integrity protections that meet or exceed the minimal encryption requirements defined under the HHS Standard for Encryption of Computing Devices and Information.
  • Provide proactive monitoring (e.g., using focused sweeps) for IOCs for assets, which must:
    • Report through alerts and logs
    • Ingest custom IOCs using the OpenIOC XML format
    • Ingest threat information in CMS and federally defined standard formats (i.e., STIX / TAXII)
  • Use software-based agents on client platforms to perform data collection/acquisition
    • Quickly identify and isolate compromised systems, which includes mitigation procedures to remediate the compromise and secure the system from similar incidents
  • Monitor the use of information system accounts and services that:
    • Monitor at the operating system and application level
    • Include COTS and custom applications
    • Provide automated monitoring (i.e., auditing) of account creation, modification, and termination (including disabling)
  • Review and analyze auditable events collected from sources listed in Appendix A

For FISMA systems categorized under FIPS 199 as HIGH, the business owner / information system owner / SDM must:

  • Incorporate an NSEP solution capable of performing automatic firewall blocking

Rationale:

NSEPs provide a HIDS / HIPS security suite to protect endpoints (e.g., servers in the data center) from compromise, detect incidents, manage vulnerabilities, and provide malicious code detection. NSEPs provide asset management (including allowlisting / denylisting ), encryption, firewalls, application firewalls, AV, and advanced investigation.

Related CMS ARS Security Controls include: SC-7 - Boundary Protection, RA-5 - Vulnerability Monitoring and Scanning, SI-4 - Information System Monitoring, AU-6 - Audit Record Review, Analysis, and Reporting, CA-2 - Control Assessments, AC-18 - Wireless Access, CM-8 - Information System Component Inventory, and IR-4(11) - Integrated Incident Response Team.

FIPS 140-2 or FIPS 140-3 Validated Encryption

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-24: FIPS 140-2 or FIPS 140-3 Validated Encryption Use

The SDM must:

  • Ensure that encryption implemented meets or exceeds the minimal requirements defined under the HHS Standard for Encryption of Computing Devices and Information. The implemented level of encryption must be sufficient to mitigate the risk based on the sensitivity of the data. The actual physical data residing on the storage device within the tool or service (e.g., Oracle database) would require FIPS compliant encryption be enabled for the data at rest. For example, a server storing highly sensitive information, such as payroll data, whole disk encryption would not be sufficient to protect information at rest.
  • Ensure encryption algorithms to protect transmitted information and information at rest meets or exceeds CMS encryption requirements (i.e., use of a FIPS 140-2 or FIPS 140-3 validated module).
  • Ensure information system(s) protect the confidentiality and integrity of transmitted information and information at rest.
  • Provide endpoint encryption that applies to information at rest (i.e., full disk, file/folder, and/or database level encryption) and information in transit.
  • Provide support for collecting FIPS 140-2 or FIPS 140-3 security control attributes from hardware devices, software products, and configuration settings in support of the CMS and HHS continuous monitoring strategies.

Encryption is required to protect sensitive information in transit on both external and internal networks (within and across zones). The environment must use mechanisms capable of either decrypting transmitted data for security analysis or analyzing the security of the data before transmission or after receipt.

Use of non-encryption for information in transit within any cloud deployments must be fully documented within required security documentation and pre-approved by the AO.

Rationale:

OMB Memorandum M-17-12, Preparing for and Responding to a Breach of Personally Identifiable Information, January 3, 2017, mandates that agencies protect sensitive information. Examples of sensitive information include personally identifiable information (PII), protected health information (PHI), intellectual property, and other information that have a degree of confidentiality. Information sharing between the CCIC and the SDM secure enclave / boundary is through CMSNet. Traffic between the CCIC and the secure enclave is encrypted through an IPSec VPN tunnel to ensure the confidentiality and integrity of the data in transit.

Related CMS ARS Security Controls include: SC-7(24) Personally Identifiable Information, SC-8 - Transmission Confidentiality and Integrity, SC-28 - Protection of Information at Rest, SC-28(1) - Cryptographic Protection, AC-6 - Least Privilege, AC-20 - Use of External Systems, and AC-18 - Wireless Access.

Insider Threat Detection

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-25: Insider Threat Detection

The SDM must:

  • Provide security awareness training on recognizing and reporting potential indicators of insider threat as part of the security awareness training program
  • Address insider threat under the SDM incident handling capability
  • Ensure insider threat policies, processes, and procedures are compliant with CMS-defined insider threat policies, processes, and procedures

Rationale:

An insider threat is a malicious threat to CMS, the FISMA system, or the SDM that comes from individuals within the organization, such as employees, former employees, contractors, or business associates. These individuals typically possess inside information on the organization’s security practices, data, and computer systems. HHS and CMS have implemented policies and procedures to aid in the detection of these individuals who present as insider threats.

Related CMS ARS Security Controls include: PM-12 - Insider Threat Program, , AT-2(2) - Insider Threat (Training), AT-2(3) - Social Engineering and Mining, IR-4 – Incident Handling, AT-2 - Literacy Training and Awareness, AT-3 - Role-Based Training, and PL-4 - Rules of Behavior.

Endpoint Data Loss Prevention

Data Loss Prevention at the endpoint provides consistent protections to block exfiltration of sensitive (especially privacy) data outside the organizational boundary as well as prevent inappropriate use (i.e., outside a documented routine use). DLP capabilities and functions include the following:

  • Multi-Platform Capability/Multi-Database Capability – The capability of the DLP solution to be available regardless of underlying applications, databases, operating systems, or hardware platforms
  • Interpretation of System-Readable Policies and Formalized Connection Agreements – The capability of the DLP solution to instantiate security and privacy rules that will ensure compliance with cognizant Privacy Notices and applicable laws and regulations
  • Role / Attribute-Based Data Protection – The capability of the DLP solution to allow a system administrator to assign data protection schemes (encryption, application and system access controls, hashing, and substitution) to data elements in another system, associate those schemes to defined roles, and assign users to the roles
  • Exfiltration Alerts and Prevention – The capability of the DLP solution to monitor for data movement between systems initiated by a user, or even the system itself, and restrict or limit movement of the data based on rule sets or behavioral patterns. The solution is expected to report and alert on deviation from the defined boundaries or thresholds.
  • Protection Orchestration – The capability of the DLP solution to interoperate within a suite of tools or be integrated within a SIEM infrastructure in a manner that provides the required access control and monitor functionality across systems

Some DLP solutions further enhance available protections for sensitive information, such as PII, by:

  • Providing enhancing monitoring and recognition of sensitive information traversing interconnections to ensure:
    • The exchange is restricted to authorized information
    • The exchange is restricted to authorized purposes
    • The exchange is restricted to authorized entities / users
  • Restricting the ability for unauthorized users and services to create “archival” copies (e.g., copies and backups) on unauthorized devices and media

The business rule within this topic applies to all FISMA system and SDM data centers supporting CMS.

BR-CCIC-28: Endpoint Data Loss Prevention

The SDM must:

  • Provide support DLP methods to protect data on endpoints (i.e., data at rest and data in use) using one or more of the following:
    • Content monitoring and inspection
    • Contextual monitoring and analysis
    • Metadata / tagging monitoring and inspection
  • Support compliance with regulations and mandates (e.g., Privacy Notice(s), CMS policy, and interconnection agreements) on sharing of CMS sensitive information on authorized endpoints via authorized interconnections
  • Provide audit trail information related to execution of DLP methods and the movement of data in a format compliant with CCIC needs
  • Perform DLP using one or more of the following methodologies:
    • Encryption
    • Quarantine
    • Block
    • Notification
    • Allow with user justification
  • Provide a monitoring or DLP capabilities in accordance with one or more of the following:
    • Endpoint monitoring, alerting on, and preventing use or manipulation of sensitive data by end-user activity (e.g., copy, paste, save, open, print operations, and screen captures) to detect or prevent data exfiltration
    • User and/or system monitoring to alert and prevent unauthorized use, storage, and transmission of privacy data by a user and/or system that has other legitimate access to sensitive data

Rationale:

The DLP capability provides automated protections for sensitive information that can reduce the risk of unauthorized access. DLP will use fine-grained access control methodologies to restrict access to information to ensure that only authorized users, roles, and/or systems may access the information. DLP solutions can also provide prevention countermeasures for data in transit that can minimize the risk from unauthorized exfiltration (i.e., transmission of sensitive data to an unauthorized recipient).

Related CMS ARS Security Controls include: AC-3 - Access Enforcement, AC-4 - Information Flow Enforcement, CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, and SC-7 - Boundary Protection.

Future CDM Considerations

CDM Asset Management is discussed above and represents the initial phase of the CISA Continuous Diagnostics and Mitigation (CDM) Program. DHS is continuing to develop the CDM Program and is expanding it to include other capabilities described in this topic. See CDM Technical Capabilities, Volume Two, Requirements Catalog, Version 2.5 as well as the GSA Continuous Diagnostics and Mitigation Tools website.

CDM for Identity and Access Management

CDM for Identity and Access Management focuses on monitoring user privileges and activities to improve infrastructure integrity through the following requirements:

  • TRUST – Provide access control management to ensure there is trust in the people who are granted access. This is also known as Identity, Credential and Access Management (ICAM).
  • BEHAVE – Provide security-related behavior management to help detect when behaviors change in an unanticipated manner, which is often a sign of malicious activity
  • CRED – Provide credential and authentication management to ensure only authorized users have access to the information systems and data

CDM for Network Security Management

CDM for Network Security Management focuses on boundary protection and event management as key components of information system and information security life-cycle management. Network Security Management will encompass the following requirements:

  • Boundary Protection – Ensure that traffic to and from the FISMA system (i.e., network, physical, and virtual boundaries) does not compromise security
  • Plan for Events – Ensure that resources are in place to deal with both routine and unexpected events that can compromise security. These include cybersecurity incidents and contingencies (Acts of God) such as floods and earthquakes.
  • Respond to Events – Ensure that responses are as planned for both routine and unexpected events that can compromise security and require a response to maintain functionality and security. Events include actual cybersecurity incidents and contingencies (Acts of God).
  • Generic Audit/Monitoring – Identify routine and unexpected events that can compromise security. These include actual cybersecurity incidents and contingencies (Acts of God).
  • Document Requirements, Policy – Ensure that documentation and policy processes fully integrate the concepts of information security
  • Quality Management – Ensure that quality management processes fully integrate the concepts of information security
  • Risk Management – Implement operational security (risk management) initially through risk scores and later supplemented by Level 3 maturity metrics

CDM for Data Protection Management

CDM for Data Protection Management focuses on capabilities that help to protect data. Data Protection Management will encompass the following requirements:

  • Data Discovery – Ensures consistent identification of data assets across the organization for processing, storing, and transmitting information at all sensitivity levels
  • Data Protection – Ensures protection of the data itself. Data protection includes the application of cryptographic methods as well as the ability to hide sensitive data field values using data masking or obfuscation methods. This is in addition to the standard method of controlling access privileges for sensitive information.
  • Data Loss Prevention – Ensures consistent protection to block inappropriate exfiltration of sensitive (especially privacy) data outside the organization and prevent unauthorized use (i.e., outside of documented routine use) both internally and externally
  • Data Breach / Spillage Mitigation – Ensures implementation of and adherence to policies, processes, and procedures that an organization develops in response to an unauthorized loss of organization data
  • Information Rights Management – Ensures access to enterprise information (e.g., documents and files) is managed. IRM solutions provide fine-grained and identity-aware protections that are persistent with the data.

Ongoing Assessment and Authorization

Ongoing assessment and authorization, which may be referred to as continuous monitoring, is part of the overall Risk Management Framework (RMF) for information security. The ongoing assessment and authorization process will determine, using as much automation as possible, whether the set of deployed security controls for the FISMA system remains effective despite planned and unplanned changes that occur in the FISMA system and its environment over time.

The ongoing assessment and authorization program is based on NIST SP 800-137, Information Security Continuous Monitoring for Federal Information Systems and Organizations, NIST SP 800-37R2 (RMF), Risk Management Framework for Information Systems and Organizations and OMB A-130 Office of Management and Budget (OMB) Circular A-130, Managing Information as a Strategic Resource. Ongoing assessment and authorization produce greater transparency of the security posture of the FISMA system, which facilitates timely risk management decisions. Security-related information collected through continuous monitoring processes is used to dynamically update the FISMA system’s SSP, Information Security Risk Analysis, and POA&Ms. These updated documents, as well as the real-time operational feeds, enhance the accuracy of the security authorization package while simultaneously providing information about security control effectiveness. Through these documents, CMS can make timelier informed risk management decisions and better address changing threats and conditions.

System and Devices Logs and Data Sources

To support CMS incident detection and response capabilities, meet various CMS reporting needs, and comply with HHS and federal security information logging and analysis requirements, CMS requires that specific log data be collected by the SDM from systems, infrastructure components, and other data sources.

The SDM must collect this log data in a local SDM SIEM instance that is compatible with both CCIC SIEM platform requirements and CMS CDM tools to enable integration with and secure log transport to these CMS monitoring platforms. The CCIC SIEM platform and CDM toolsets provide a variety of software agents and collection mechanisms to collect and securely transport logs from various enterprise systems and devices. Integration and monitoring of these collected logs files require close collaboration with the CCIC and CDM teams. The SDM also must follow audit requirements (AU and CM controls) within CMS ARS, especially settings for event generation cited in AU-2 and AU-12 as well as applicable configuration settings (e.g., STIG and USGCB) directed by CM-6 controls. SDM contractors must work with the CCIC and CDM teams to integrate the corresponding data feeds / logs to ensure secure data in transit and at rest.

For specific information regarding integration with the CMS CCIC SIEM and CDM tools, and which logs are to be provided to each platform, send an email to CMS - ISPG Logs Onboarding mailbox to receive the latest CMS requirements. (CMS - ISPG-DCTSO may be used as an alternate)

The following topics provide a representative baseline of the types of CMS-specific security data that must be collected and forwarded to a local SDM SIEM and in turn integrated into the CCIC SIEM Platform and/or CDM toolsets. The information in the topics below is intended to provide overall guidance but does not provide the complete list of required log or other monitoring data as requirements change over time. Use the process noted above to obtain the details on current logging requirements

Encryption

The suite of tools implemented by the CCIC and required for implementation at the SDM is designed to monitor unencrypted traffic, log files, and data sources.

The SDM is required to:

  • Ensure that any monitored traffic, logs, or data sources are available (e.g., through submitted queries) to CCIC security tools in unencrypted formats:
    • For applications that use TLS (highest available level) for encryption in transit, the monitoring point must occur after the point at which TLS traffic has been unencrypted.
    • If the TLS (highest available level) traffic occurs in an end-to-end fashion, then the contractor must implement the infrastructure necessary to decrypt that traffic (e.g., Metronome).
  • Protect logs and/or data sources in transit across networks between the point of monitoring and the tools consuming them through either of the following methods:
  • Encrypt logs during transmission
  • Provide logical and physical segregation from normal traffic via an out-of-band solution, dedicated VLAN or subnet, or a combination thereof

SIEM forwarders typically have the required encryption capability built in and should be used where technically possible. Once at the destination, log data should be unencrypted and processed by the SDM SIEM platform. FIPS 140-2 or FIPS 140-3 compliant encryption should be enabled and configured within the SDM SIEM platform (not this may not be enabled by default)

Domain Name Service Servers

The US-CERT, other government agencies, incident response teams, the Department of Defense, and commercial vendors frequently publish lists of domains and IP addresses of websites that host malware, botnets, command and control networks, and other malicious content. Access to these known sites and/or Domain Name System lookups for these sites is a very strong indicator that the internal HHS system is already infected or compromised by malware. Access to DNS information will provide an early indicator of a successful compromise or intrusion, which allows the system to be quickly identified and taken offline before data can be destroyed, compromised, modified, or exfiltrated.

Some organizations have had success in monitoring or data mining DNS and Microsoft Active Directory (AD) logs for consecutive DNS lookups, which indicate internal compromises reaching out to establish footholds on other systems. DNS logs are also used for detection of DNS cache poisoning and DNS redirect attacks. For DNS logs to have maximum value, the DNS logs must be directly attributable and traceable to the originating host system.

The SDM must collect DNS logs according to the following requirements:

  • DNS logs must be directly attributable and traceable to the originating host system for all systems that make DNS queries within the SDM.
  • If a caching architecture is implemented in the environment, CMS requires logging data from the primary DNS resolvers for the environment(s) as well as any upstream caching servers within the data center.
  • Configuration data from the systems performing DNS lookups and/or resolution must be provided to CMS on request.

Email Servers

Email servers accept, forward, deliver, and store electronic messages between email clients. Email messages typically consist of a message header and message body. Email server logs contain information that can be used to trace spear phishing attacks and detect malware attachments sent by email.

US-CERT, the Department of the Treasury, and other government agencies publish daily and weekly advisories on email phishing attacks that can be used to analyze email logs.

The SDM must collect email server logs according to the following requirements:

  • If systems within the SDM’s security boundary use email, the email logs must be made available to CMS as a data source for monitoring.
  • Any email filtering systems must be able to provide security-related logging data to CMS as a data source.
  • Configuration data from the systems sending and/or processing email must be provided to CMS on request.

Proxy / URL Filters

A URL filter is a form of proxy server that acts as a go-between for requests from clients seeking hypertext transfer protocol (HTTP) resources from a remote web server. A URL filter server can block access to unsafe websites.

URL filter servers produce logs that track access to known malware hosting sites and URL block list access. These logs can assist a security analyst in identifying compromised hosts on the local network.

The SDM must collect and forward proxy / URL logs according to the following requirements:

  • Logging data from these systems must be provided in such a fashion that the full session information can be captured with the following attributes (at a minimum):
    • Attribution to the originating host system, including IP address and HTTP header information
    • Indication of whether or not the request was served out of a local cache
    • Indication of whether the request was blocked or allowed and what category (e.g., social networking, news media, and shopping) was associated with the request. Category classes provide granular visibility into the type of web traffic traversing a given network.
    • Attribution information for the involved user (if the request involved user interaction)
  • Read-only access to the reporting, alerting, and administrative configuration features must be provided.
  • Additional logging fields within the tool must be turned on should the logging fields be required for security monitoring purposes as determined by ISPG. Table - Example URL Filter Fields provides an example list of URL filter fields that should be captured for monitoring.
  • Proxy/URL filtering implementation must allow for customized categories and rules to respond to notifications and/or alerts from various sources.
Table - Example URL Filter Fields
NameTypeEnd TimeCustomer ResourceAggregated Event CountCorrelated Event CountCategory Significance
Category BehaviorCategory TechniqueCategory Device GroupCategory OutcomeCategory ObjectModel ConfidenceSeverity
RelevanceAsset CriticalityPriorityAgent SeverityAgent Host NameAgent AddressAgent Zone Resource
Agent Asset Identifier (ID)Agent VersionAgent Time ZoneAgent IDAgent TypeAgent NameDevice Event Category
Device SeverityDevice ActionDevice Event Class IDDevice AddressDevice Zone ResourceDevice Time ZoneDevice Vendor
Device ProductDevice Process NameAttacker AddressAttacker Zone ResourceAttacker User NameAttacker Geo Location InfoAttacker Geo Country Name
Target Host NameTarget AddressTarget Zone ResourcesTarget PartTarget Geo Location InfoTarget Geo Country NameRequest URL
Event Annotation Stage ResourceEvent Annotation Modification TimeDevice Custom Number ThresholdDevice Custom String Alert CountReferrerMIME TypeUser Agent
Request Client ApplicationBytes In / OutN/AN/AN/AN/AN/A

 

Firewalls

Firewall log data is a vital part of network security, because the firewall is the first line of defense against unauthorized external access. Firewall logs generate data that can be useful in identifying security events. Port scans, operating system identification, denial of service (DoS) attacks, network mapping attempts, blocked network address access attempts, and inbound perimeter breach attempts are examples of security incidents that can be detected while analyzing firewall logs. Firewall logs can also detect instant messaging, peer-to-peer (P2P) file sharing traffic through HTTP, and firewall policy changes as well as outbound access attempts to blocked URLs, network addresses, and listed malicious websites.

The SDM must collect firewall logs according to the following requirements:

  • Logging data from firewalls must be provided for any firewalls that protect the perimeter or any segment of the SDM systems.
  • Logging data must:
    • Be attributable to the originating system, except where internal proxies are routing traffic for hosts internal to the network
    • Log blocks, denies, drops, or other policy / rule hits
  • Read-only access to Layer-3 devices, the configuration repository, aggregation server, or configuration management database must be provided on CMS request for analysis.
  • The use of additional logging capabilities within the firewall systems may be required, depending on the nature of the operational environment and the monitoring requirements ISPG determines.

Intrusion Detection / Prevention Systems

IDS / IPS logs are valuable tools for detecting various network attacks such as port scans, remote access attempts, malicious remote commands, and executed commands. IDS / IPS logs contain event data that may be used for forensics or in a correlation engine within a SIEM solution. Information on user login failure attempts, Internet messaging traffic, email communication, and P2P traffic is useful in network security monitoring and trend analysis.

The SDM must collect IDS / IPS logs according to the following requirements:

  • IDS / IPS logging/alerting data must be provided, and include (at a minimum):
    • The rule that triggered the alert (by reference identifier or the full rule itself)
    • Category of the rule and any associated risk rating / scoring data
  • IDS / IPS implementations must (at a minimum):
    • Allow for implementation of customized signatures in response to notifications and/or alerts from various security entities
    • Provide the option to automatically save an associated PCAP for triggered alerts that ISPG requests in support of incident response activities or for limited retention durations upon CMS request
    • Provide read-only access to configuration data upon CMS request or the configuration files themselves via a centralized network management facility

DHCP Servers

Dynamic Host Configuration Protocol is a server service that automatically distributes network addresses to requesting clients on the local access network. DHCP logs record the mapping of network addresses to computer names similar to how a phonebook maps names to telephone numbers.

DHCP server logs provide time-stamped mapping of client network addresses to computer hostnames for desktop and laptop network activity tracking.

The SDM must collect DHCP logs according to the following requirements:

  • Baseline activity logging must be provided, including reservation and host details as requested by CMS.
  • A naming convention for systems on the network must be provided to CCIC and updated with changes no less often than once every 90 days.
  • Configuration data from the DHCP servers must be provided to CMS on request but no less often than once every 90 days.

Network / Port Address Translation

Network Address Translation / Port Address Translation logs are useful for identifying the originating system or node on an internal network when an IDS or another entity detects a security event outside the network.

The SDM must collect NAT / PAT logs according to the following requirements:

  • Any routers or other network infrastructure performing NAT / PAT functions must provide logging data sufficient to attribute traffic to originating hosts for correlation to other network and security data.
  • Configuration data from the devices performing NAT / PAT functions must be provided to CMS on request, but no less often than once every 90 days.

Antivirus

Antivirus server logs contain log entries showing the date on which a virus scan detected a virus, the type of virus (e.g., Trojan, spyware, and other forms of malware), the filename, where the virus was found, and the action taken by the scan (e.g., cleaned or quarantined). The AV server logs also contain AV client status information that can be used to identify vulnerable clients on a network.

The SDM must collect antivirus logs according to the following requirements:

  • Data related to antivirus state on the host and management server must be provided, including:
    • Current executable version
    • Current rules version
  • Logging and data related to activities on the host and management server must be provided, including:
    • Most recent scan
    • Any alert, quarantine, cleaning, or other intervention activity
  • Configuration data from host or management server must be provided to CMS on request but no less often than once every 90 days.

Microsoft Active Directory Security Events

Microsoft AD security events contain information about logins to Microsoft systems—user ID, date, time, and location. This data can be used both proactively and forensically. Proactively, the data can aid in the detection of unauthorized user activity: unauthorized system, application, or data access; failed or brute force login attempts; misconfigured application installations; unauthorized password changes; and user account creation. Misuse of administrative login credentials at unauthorized locations, dates, time, or systems is a strong indication of an intruder compromise.

Forensically, access to AD security logs is helpful in post-incident analysis to determine who did what, when, and how after an intrusion has been detected. One of the first steps an intruder usually takes is to delete or remove security event logs from the compromised system. Storing the logs on a secure remote system removes this ability from the intruder. Some organizations have had success in monitoring or data mining DNS and AD logs for consecutive DNS lookups, which indicates internal compromises reaching out to establish footholds on other systems.

The SDM must collect AD logs according to the following requirements:

  • AD logging requirements are heavily dependent on the role of AD within the architecture. ISPG may require additional logging data beyond the relevant minimum baseline configuration data depending on the operational environment.
  • Configuration data from the AD systems must be provided to CMS on request but no less often than once every 90 days.

Microsoft Windows Security and System Events

Microsoft Windows security events contain information about local logins to Microsoft systems—user ID, date, time, and location. This data can be used both proactively and forensically. Proactively, the data can aid in the detection of unauthorized user activity: unauthorized system, application, or data access; failed or brute force login attempts; misconfigured application installations; unauthorized local password changes; and local user account creation. In addition, these logs can indicate changes to local file, registry, or service permission or ownership changes.

Forensically, access to Windows security and system logs is helpful in post-incident analysis to determine who did what, when, and how after an intrusion has been detected. One of the first steps an intruder usually takes is to delete or remove security event logs from the compromised system. Storing the logs on a secure remote system removes this ability from the intruder.

The SDM must collect Windows system and security logs according to the following requirements:

  • Windows logging requirements are heavily dependent on the role of the Windows system within the architecture. ISPG may require additional logging data beyond the relevant minimum baseline configuration data depending on the operational environment.
  • Configuration data from the Windows systems must be provided to CMS on request but no less often than once every 90 days.

UNIX / Linux Security Logs

UNIX / Linux security-related logging information provides a critical stream of security data for auditing, monitoring, and incident response. A wide variety of logging services is available, depending on the UNIX or Linux variant used, although most use syslog or a near equivalent.

The SDM must collect UNIX / Linux security logs according to the following requirements:

  • UNIX / Linux logging requirements are heavily dependent on the role of UNIX / Linux within the architecture. ISPG may require additional logging data beyond the relevant minimum baseline configuration data depending on the operational environment.
  • Configuration data from the UNIX  /Linux systems must be provided to CMS on request but no less often than once every 90 days.

Mainframe Security Logs

Mainframe systems must use a security service that can output its events through a facility such as syslog. These log files can be used to audit activity and provide system-level correlation for network logs and other security data. With the rise of virtualized environments housed on mainframes, the level of detail and scope of security logging is normally site specific outside of established system baseline requirements.

The SDM must collect mainframe security logs according to the following requirements:

  • In addition to the logging data required in the applicable baseline configuration standard for the platform, ISPG may require additional logging data from the mainframe security facilities and monitoring tools in use.
  • Configuration data from the mainframe must be provided to CMS on request but no less often than once every 90 days.

Netflow or Equivalent

NetFlow is an industry-standard protocol that a firewall can use to export statistics about the IP traffic on its interfaces. The SDM must provide the NetFlow data or its equivalent based on the brand of router in use. Netflow is the statistical data provided by Cisco routers. It describes network connections and communications. Comparable “Netflow” data from competing vendors (e.g., Juniper) must also be provided if deployed within an SDM. It is analogous to telephone call logs, showing the two endpoints, time of call, and duration. Netflow does not contain information about what occurred during a call. Intrusion detection systems can use this data to look for suspicious network traffic that hides within normal network communications. Multiple systems sending small packets to the same endpoint on a regular, repetitive timeline is an indicator of compromise. Examples are botnets, command and control networks, or backdoor Trojans calling “home” for instructions. Netflow data can also be matched to the same lists of malicious sites described in the DNS usage log.

The SDM must collect Netflow data according to the following requirements:

  • Raw or unprocessed Netflow data from network devices must be provided for integration.
  • Internal device Netflow must be part of the architecture, because internal network movement cannot be observed from perimeter device Netflow data.

Routers / Switches

Time-stamped router audit logs provide a trail of user connectivity data, including router configuration changes. Logging of failed logon attempts can indicate Telnet and SSH-based attacks on the network infrastructure.

The SDM must collect router / switch / load balancers audit logs according to the following requirements:

  • Configuration data from the routers, switches, and load balancers must be provided to CMS on request.
  • Configuration data must be provided to the CCIC for contextual risk management and threat identification application analysis.

Authentication Servers

Authentication servers provide centralized authentication, limited authorization, and account management for computers to connect and use a network service. Authentication server logs are used to detect unauthorized logon attempts and unusual logon successes (e.g., user logons when a user is known to be on vacation or outside of scheduled hours).

The SDM must collect authentication server logs according to the following requirements:

  • All events identified in CMS ARS Security Control AU-2, Event Logging, must be logged.
  • All logon and logoff events (failed, successful), including suspicious events, must be logged.
  • All logon attempts must be logged.
  • Account lockouts and unlocks must be logged.
  • Account revokes / terminations must be logged.
  • Account creation / deletions must be logged.
  • Account use of escalated privileges (assigned and inherited), including Windows (Administrator, Domain Admins, Service Accounts, Run As accounts) and Linux (Sudo, etc.), must be logged.

Virtual Private Network Servers

A virtual private network server provides a secure, private communication tunnel between two or more devices across a public network.

VPN server logs track security-related events, such as unauthorized VPN access attempts, VPN client logon failures, and remote file access attempts and successes. VPN server logs provide a time-stamped log of VPN user connectivity, including file names and the amount of data transferred between VPN client and a remotely accessed resource. These logs can provide a security analyst with useful forensics data and help to provide a detailed view of a user’s session activity.

The SDM must collect VPN server logs according to the following requirements:

  • All events identified in CMS ARS Security Control AU-2, Event Logging, must be logged.
  • Activity logs must include:
    • User / system attribution
    • Connect / disconnect times
    • Data transfer statistics
    • Remediation / Network Access Control activities related to the host
    • Host information
    • Link information
  • Lists of subnets assigned to VPN via DHCP or statically must be updated no less often than once every 90 days and provided to the CCIC.

Vulnerability Scanners

A vulnerability scanner is software designed to scan computers and use a vulnerability database to scan applications, operating systems, or networks for vulnerabilities. Vulnerability scanners produce data that allow a security analyst to reduce false positives, enabling the analyst to concentrate on real attacks.

Vulnerability logs contain patch status data and missing software update information that provides additional value when combined with inventory and incident management systems, and form part of the foundation for meeting emerging OMB-mandated continuous monitoring requirements.

The SDM must collect vulnerability scanner data according to the following requirements:

  • Notify CCIC of the IP addresses and expected behavior of vulnerability scanners to avoid false positives.
  • Provide log data from requested or routine usage of the vulnerability scanners in support of incident response, monitoring, or other security requirements.
  • Provide relevant results of any additional vulnerability scans not directly tied to current scans from the CMS enterprise vulnerability scanning platform within the SDM. The CCIC may gain additional information for monitoring purposes from vulnerability scanning of internal SDM environments through tools other than those part of the CMS enterprise vulnerability scanning platform.

Wi-Fi

The SDM and CMS can generate information for monitoring activity on CMS’s wireless networks from wireless access points within the SDM used to access CMS capabilities.

The SDM must collect Wi-Fi data according to the following requirements:

  • Notify CCIC of the IP addresses and expected behavior of hosts accessing the Wi‑Fi network to avoid false positives.
  • Provide log data from requested or routine usage of the access points in support of incident response, monitoring, or other security requirements.

Host-Based Sensing

The following artifacts are considered vital to detecting malicious adversarial behaviors on an enterprise network:

  • Process Data. The following artifacts should be collected for process STARTs and TERMINATES:
    • Computer name, user, command line, executable name, process ID (PID), process image path, MD5 and Secure Hash Algorithm 2 (SHA2) w/256-bit digest (SHA256) hash of process, parent process, parent process ID (PPID), parent process image path, system ID (SID), and the signer of the executable
  • Process Threads (injection). The following artifacts should be collected for process threads:
    • Computer name, user, source PID/target ID (TID), Target PID/TID, stack base, stack limit, start address, start function, start module, start module name, sub process tag, user stack base, and user stack limit
  • File Data. The following artifacts should be collected while monitoring file activity on a system, including CREATE, DELETE, MODIFY, READ, WRITE, and ATTRIBUTE MODIFICATION:
    • Computer name, user, company name, create time, file name, file path, fully qualified domain name (FQDN), image path of the process that created the file, PID of the process, parent process PID and image name, MD5 and SHA256 of file created, and the signer
    • Base Address (for Drivers only)
    • Monitor any file creations in specific Windows vital folders (windows, system32, syswow64, etc.)
    • Monitor file creations of executable file types in any location (exe, zip, zipx, msi, py, ps1, scr, dll, cmd, bat, rar)
  • Registry Data. The following artifacts should be collected for any registry activity, including CREATE/DELETE/MODIFY KEY/VALUE:
    • Computer name, user, key name, value name, data, hostname, and value type
    • Monitoring the registry can cause volume issues, because registry edits occur on a regular basis in Windows. If a reasonable way of tracking registry activity is not possible, then at a minimum watching all Service and Autorun keys is recommended.
  • Service Data. The following artifacts should be collected for changes to any service on the system being modified, deleted, or created:
    • Computer name, user, service name, command line, executable making the change, FQDN, and image path of the executable
  • Modules Installed. The following artifacts should be collected for any module that is installed on a system:
    • Computer name, user, base address, FQDN, MD5 and SHA256 hash of module, module name, module path, PID, signer, and TID
  • User Sessions. The following artifacts should be collected for user sessions logging into a system:
    • Computer name, user, destination/source IP (remote logons), destination / source port (remote logons), logon id, logon, type, and the privileges associated to the logged on user
  • Network Data. The following artifacts should be collected for network connections originating or connecting to a network node:
    • Computer name, user, destination / source IP and port, destination / source hostname, destination / source FQDN, start time, end time, flags, image path of process responsible for the connection, PID and PPID of process, packet count, MAC address, protocol, and proto info
  • Others
    • Monitor for windows hooking
    • Collect Windows Management Instrumentation (WMI), security, and system logs

Monitoring file activity can cause volume issues, because Windows is constantly creating temp files and other artifacts that would be considered false positives. If a reasonable plan cannot be devised to account for this, then at a minimum the following actions should be taken to monitor file activity on a system:

Business Application Logic

Business application logs give insight into any unusual business application activity, as identified by the business application owner in response to auditable event types identified under CMS ARS Security Control AU-2. Often, the application developer custom codes these event logs, and the event logs may rely in part on events found in web server access logs, application service activity logs, and database transaction logs. The application operator uses these logs to understand problems in logic and workflow through the application, but they also help users or administrators to detect inappropriate use of the application. Security staff use these logs to help detect attacks such as SQL injection or cross-site scripting.

The SDM must collect application-coded logs designed to comply with the following types of auditable events identified in CMS ARS Security Control AU-2, Event Logging:

  • Application Modifications (e.g., adding, removing, or modifying application items, successful by application admins, and failed attempts by application users)
  • Application alerts and error messages (e.g., logic / capacity / loss errors and logic high-processing alerts)
  • Configuration changes (e.g., successful changes or failed attempts to change application parameters and preferences)
  • Reading, Modification, or Printing of sensitive information (e.g., successful or failed actions against sensitive data as seen by the application logic performed by application users and administrators)
  • Account creation, modification, or deletion (e.g., adding or removing application users or application administrators)

Other Feeds or Logs

Additional data feeds may also be required (and determined in coordination with CMS) depending on the nature of the architecture and work performed by the SDM. CCIC may place additional monitoring assets and capabilities in place to monitor the raw data at those key aggregation points, as ISPG determines necessary.

The SDM must collect additional data feeds as needed by CCIC including, but not limited to, log events and alerts from databases, application servers and hypervisors.

 

Access Control and Identity Management Introduction

Introduction

This chapter presents a general overview of CMS practices and services for access controls and identity management for CMS enterprise environments. Specific guidance on access controls and identity management may be found in the CMS ARS, the CMS Information Systems Security and Privacy Policy, and in the Business Rules topic of this chapter.

Relevant Documents

This chapter is not all inclusive. It complements and incorporates CMS’s existing policies, standards, and procedures, thereby offering an architectural view of the standards. See CMS Information Security and Privacy Overview. Where there are conflicts, the following standards, and any successor documents, will take precedence:

Other related guidance includes:

Key Concepts and Definitions

To ensure the audience fully understands the conceptual foundation for this guidance, the following terms, definitions, and explanations are presented in a sequence that progressively build the foundation for the guidance. The Glossary defines additional terms for the reader’s benefit.

Identity Management

The primary goal of Identity Management is to establish a trustworthy process for assigning attributes to a digital identity and to associate that identity to an individual. Identity management includes the processes for maintaining and protecting the identity data of an individual over the life cycle of the digital identity.

Privilege Administration

The process of establishing and maintaining the entitlement or privilege attributes that comprise an individual’s access profile. Privilege Administration processes entail reviewing, assigning, and removing privileges over time as an individual’s access needs change.

At CMS, business application owners are responsible for implementing user Privilege Administration processes for their applications.

Assignment and administration of privileges within an application, such as application-specific roles and role-based access controls, are not supported by CMS enterprise infrastructure systems, and are the responsibility of the application owner.

RBAC Profile (also known as Job Code or Role)

Also known as Enterprise User Administration (EUA) Job Codes, Production Control Identifiers, Business Profiles, or RBAC (Role-Based Access Control) Profiles. Profiles are used to control which applications a given user has been authorized to access. Each UserID may be assigned zero or more Profiles, and each Profile may impart to the UserID zero or more Access Control Entitlements (ACE). A Profile assigned to a UserID indicates that user is authorized to authenticate (i.e., log into) to a given application or set of applications. Please refer to Privilege Administration and Enterprise User Administration below for additional discussion of Profiles and Roles.

User Certification

The periodic review of a user’s authorization status. The process of certification verifies the user is still an active user, is still authorized to use CMS systems, has received security awareness training, and has reviewed and accepted the CMS Privacy Statement. User Certification is, at minimum, an annual requirement.

Credential Service Provider (CSP)

Not to be confused with Cloud Service Provider (also CSP) used elsewhere in the CMS TRA, a Credential Service Provider is a trusted entity that issues or registers subscriber authenticators and issues electronic credentials to subscribers. A CSP may be an independent third party or issue credentials for its own use.

E-Authentication Assurance Levels

Federal information systems are required to incorporate information security controls to protect the information systems supporting their operations and missions. CMS is required to ensure the adequate protection of its information assets and must meet a minimum level of information security.

Previously, CMS guidance on e-authentication assurance levels aligned with NIST SP 800-63-2, Electronic Authentication Guideline, August 2013, and OMB Memorandum 04-04, E-Authentication Guidance for Federal Agencies, and defines four (4) levels of assurance (LOA) for electronic transactions in terms of the consequences of the authentication errors and misuse of credentials:

  • Level 1 affords little or no confidence in asserted identity’s validity. Identity proofing relies on the user’s own assertions. The authentication token may be a UserID and password.
  • Level 2 provides some confidence in the asserted identity’s validity. Identity proofing requires verifying the individual’s government-issued ID or a financial account and other information. The authentication token may be a UserID and password.
  • Level 3 provides high confidence in the asserted identity’s validity. Identity proofing requires verifying the individual’s government-issued ID and a financial account and other information. Multi-factor authentication is required.
  • Level 4 provides very high confidence in the asserted identity’s validity. In-person proofing is required. Multi-factor authentication is required.

NIST has superseded this definition of e-authentication assurance levels. NIST SP 800-63-3, Digital Identity Guidelines, June 2017 and updated March 2020, supplements OMB Memorandum 19-17, Enabling Mission Delivery through Improved Identity, Credential, and Access ManagementAuthentication Guidance for Federal Agencies, and defines three (3) components of assurance for electronic transactions in terms of the consequences of the authentication errors and misuse of credentials. These guidelines provide mitigations of an authentication error’s negative impacts by separating the individual elements of identity assurance into discrete, component parts. For non-federated systems, agencies select two components, referred to as Identity Assurance Level (IAL) and Authenticator Assurance Level (AAL). For federated systems, agencies select a third component, Federation Assurance Level (FAL).

These guidelines retire the concept of a level of assurance (LOA) as a single ordinal that drives implementation-specific requirements. The components of identity assurance detailed in these guidelines are as follows:

  • IAL refers to the identity proofing process.
  • AAL refers to the authentication process.
  • FAL refers to the strength of an assertion in a federated environment, used to communicate authentication and attribute information (if applicable) to a relying party (RP). FAL is not used at this time per CMS policy.

The separation of these categories provides agencies flexibility in choosing identity solutions and increases the ability to include privacy-enhancing techniques as fundamental elements of identity systems at any assurance level.

CMS has not yet adopted the new e-authentication assurance components in their guidance documents. Because the new components provide functional separation and clarity when describing e-authentication, this release of the CMS TRA uses both the current four (4) levels of assurance (LOA) and the new e-authentication assurance components in its narrative describing identity management and authentication functions and services. This release does not provide a mapping between the current four (4) levels of assurance and the new e-authentication assurance components.

NIST SP 800-63A Enrollment and Identity Proofing

NIST SP 800-63A addresses how applicants can prove their identities and become enrolled as valid subscribers within an identity system. It provides requirements by which applicants can both identity proof and enroll at one of three different levels of risk mitigation in both remote and physically present scenarios.

NIST SP 800-63A sets requirements to achieve a given IAL. The three IALs reflect the options agencies may select from based on their risk profile and the potential harm caused by an attacker making a successful false claim of an identity. The IALs are as follows:

  • IAL 1. There is no requirement to link the applicant to a specific real-life identity. Any attributes provided in conjunction with the authentication process are self-asserted or should be treated as such.
  • IAL 2. Evidence supports the real-world existence of the claimed identity and verifies that the applicant is appropriately associated with this real-world identity. IAL 2 introduces the need for either remote or physically present identity proofing. Attributes can be asserted by CSPs to RPs in support of pseudonymous identity with verified attributes.
  • IAL 3. Physical presence is required for identity proofing. Identifying attributes must be verified by an authorized and trained representative of the CSP. As with IAL 2, attributes can be asserted by CSPs to RPs in support of pseudonymous identity with verified attributes.

NIST SP 800-63B Authentication and Lifecycle Management

For services in which return visits are applicable, a successful authentication provides reasonable risk-based assurances that the subscriber accessing the service today is the same as who accessed the service previously. The robustness of this confidence is described by an AAL categorization. NIST SP 800-63B addresses how an individual can securely authenticate to a CSP to access a digital service or set of digital services.

The three AALs define the subsets of options agencies can select based on their risk profile and the potential harm caused by an attacker taking control of an authenticator and accessing agencies’ systems. The AALs are defined in SP 800-63B as follows:

  • AAL 1. AAL 1 provides some assurance that the claimant controls an authenticator bound to the subscriber’s account. AAL 1 requires either single-factor or multi-factor authentication using a wide range of available authentication technologies. Successful authentication requires that the claimant prove possession and control of the authenticator through a secure authentication protocol.
  • AAL 2. AAL 2 provides high confidence that the claimant controls authenticator(s) bound to the subscriber’s account. Proof of possession and control of two distinct authentication factors is required through secure authentication protocol(s). Approved cryptographic techniques are required at AAL 2 and above.
  • AAL 3. AAL 3 provides very high confidence that the claimant controls authenticator(s) bound to the subscriber’s account. Authentication at AAL 3 is based on proof of possession of a key through a cryptographic protocol. AAL 3 authentication requires a hardware-based authenticator and an authenticator that provides verifier impersonation resistance; the same device may fulfill both these requirements. In order to authenticate at AAL 3, claimants are required to prove possession and control of two distinct authentication factors through secure authentication protocol(s). Approved cryptographic techniques are required.

NIST SP 800-63C Federation and Assertions

NIST SP 800-63C provides requirements when using federated identity architectures and assertions to convey the results of authentication processes and relevant identity information to an agency application. In addition, this volume offers privacy-enhancing techniques to share information about a valid, authenticated subject and describes methods that allow for strong multi-factor authentication (MFA) while the subject remains pseudonymous to the digital service.

The three FALs reflect the options agencies can select based on their risk profile and the potential harm caused by an attacker taking control of federated transactions. The FALs are as follows:

  • FAL 1. Allows for the subscriber to enable the Relying Party (RP) to receive a bearer assertion. The assertion is signed by the Identity Provider (IdP) using approved cryptography.
  • FAL 2. Adds the requirement that the assertion be encrypted using approved cryptography such that the RP is the only party who can decrypt it.
  • FAL 3. Requires the subscriber to present proof of possession of a cryptographic key referenced in the assertion in addition to the assertion artifact itself. The assertion is signed by the IdP and encrypted to the RP using approved cryptography.

Identification and Authentication

The topics below provide a high-level overview of CMS Identity and Authentication (IA) policies and services. For more specific information, see Business Rules and Recommended Practices.

In addition, please consult the CMS ARS guidance on the Identification and Authentication  policy family for detailed guidance.

The primary goal of Identity Management is to establish a trustworthy process for assigning attributes to a digital identity and to associate that identity to an individual. Identity management includes the processes for maintaining and protecting the identity data of an individual over the life cycle of the digital identity. Directory services provide applications with access to repositories of digital identities and associated attributes, as well as services and data required to authenticate user passwords and control access to application resources.

CMS, or a CMS-authorized third party, acts as a Credential Service Provider that issues or registers subscriber authenticators and issues electronic credentials to subscribers. According to NIST SP 800-63, an individual or applicant opts to be identity proofed by a CSP. If the applicant is successfully proofed, the individual is then termed a subscriber of that CSP. The CSP establishes a mechanism to uniquely identify each subscriber, register the subscriber’s credentials, and track the authenticators issued to that subscriber. The subscriber may be given authenticators at the time of enrollment, the CSP may bind authenticators the subscriber already has, or they may be generated later as needed. Subscribers have a duty to maintain control of their authenticators and comply with CSP policies in order to maintain active authenticators. The CSP maintains enrollment records for each subscriber to allow recovery of authenticators, for example, when they are lost or stolen. See NIST Special Publication 80-63-3, Digital Identity Guidelines, June 2017.

While CMS business owners may implement their own solutions, CMS has available several shared Identity, Credentialing & Access Management (ICAM) services to support CMS applications in any CMS processing environment.

Enrollment and Identity Proofing

For CMS staff, contractors, and other government agencies, the enrollment and identity proofing processes are integrated with human resources and contracting processes. These processes require in-person identity proofing and support the provisioning of a CMS badge or PIV card, and if needed, an Enterprise User Administration account (please refer to Enterprise User Administration for more information about EUA services). Similar processes are required for non-organizational users who require a Level 4 or IAL 3 identity assurance level:

  • Organizational user. Organizational users include employees or individuals that organizations deem to have equivalent status of employees (e.g., contractors and guest researchers). Please refer to control IA-2 in NIST SP 800-53.
  • Non-organizational users. Non-organizational users include information system users other than organizational users explicitly covered by IA-2. Please refer to control IA-8 in NIST SP 800-53.

CMS business owners may implement their own processes for subscriber enrollment and authorization for their access to business-specific CMS applications.

  • A recommended solution for enrolling non-organizational users is the CMS ePortal, which implements subscriber enrollment and authorization for CMS applications that use the CMS Identity Management Service (aka. IDM). Please refer to Shared Services for more information about the CMS Identity Management Service.
  • The Scalable Login System (SLS) is an internal CMS web application that supports the Federally Facilitated Marketplaces (FFM) website, healthcare.gov. SLS provides the identity management of consumers that create accounts on healthcare.gov and medicare.gov.

A CMS-authorized, third-party Remote Identity Proofing service may be used to support, but not replace, an identity proofing process for Level 2, Level 3, or IAL 2 identity assurance levels.

Remote Identity Proofing

Remote Identity Proofing (RIDP) is an automated process (not in-person) that uses government and/or commercially available data to validate an individual’s identity by asking questions of the individual. See Quick Start Remote Identity Proofing (RIDP) User Guide

The RIDP process requires the individual to provide:

  • Personal information, such as name, date of birth, address, exactly as recorded on either his / her driver’s license or any Government ID
  • Answers to questions related to user’s personal and financial history

The individual’s answers are verified using both authoritative sources (e.g., state motor vehicle administration, court records, and licenses) and non-authoritative sources (e.g., credit reports, professional organizations, merchants, warranties, and travel industry). A commercial remote identity proofing service with access to these data sources generates the questions and scores the answers. CMS uses a number of commercial services to provide a layered approach.

CMS understands that Remote Identity Proofing is fallible. RIDP uses questions based on publicly available data, or non-authoritative data sources, simply to raise confidence in an individual’s identity, but does not provide direct evidence. When using RIDP, there are some additional considerations that may require defining policies:

  • As some RIDP questions may be perceived by the public as CMS overstepping privacy norms, it is advisable to set limitations on what types of questions may be used for RIDP.
  • There should be restrictions on what CMS and commercial remote identity proofing providers may do with data collected while providing remote identity proofing services.

CMS understands that Remote Identity Proofing is not feasible for all applicants, such as people who do not have established credit histories, lack US or state government IDs, or are unable to provide answers directly. For this reason, an alternative proofing process must be available for applicants who are unable to gain approval via remote identity proofing.

Credential Provisioning

For CMS staff, contractors, and other government agencies, credential provision processes are integrated with human resources and contracting processes. These processes require in-person identity proofing and support the provisioning of a CMS badge or PIV card, and if needed, an EUA account. Similar processes are required for non-organizational users who require a Level 4 or IAL 3 identity assurance level.

For non-organizational users who require a Level 2, Level 3, or IAL 2 identity assurance level, CMS applications may provide their own CSP services, or use the CMS Enterprise Identity Management Service. The CMS Scalable Login System (SLS) supports users of the Federally-Facilitated Marketplaces (FFM) website, healthcare.gov.

Please refer to Shared Services in this chapter for more information about EUA, IDM and SLS services.

For non-organizational users who require a Level 1 or IAL 1 identity assurance level, CMS applications may provide their own CSP services, provided that the service collects minimal or no PII, and meets CMS requirements for security and protection of PII and the CMS applications.

Multi Factor Authentication

Multi-Factor Authentication is a security mechanism to verify the legitimacy of a person or transaction. MFA requires the user to provide, in addition to user ID and password, more than one form of verification to prove user identity. MFA registration is required only once when a user is requesting a role, and it is verified each time a user logs in. During the MFA registration process, the CSP requires registration of a phone number, computer, hardware or software token, or email address to use as an MFA mechanism to add an additional level of security to a user’s account. See CMS’ Identity Management

Although industry has introduced many alternative mechanisms for implementing MFA, not all are acceptable for a given CMS application or situation under NIST or CMS guidelines. These mechanisms should be reviewed for security vulnerabilities and risk before deciding to implement them. For example, please refer to NIST SP 800-63-3B guidelines, section 5.1.3.1 Out-of-Band Authenticators, which provides that, “Methods that do not prove possession of a specific device, such as voice-over-IP (VOIP) or email, SHALL NOT be used for out-of-band authentication.” which provides that, “Methods that do not prove possession of a specific device, such as voice-over-IP (VOIP) or email, SHALL NOT be used for out-of-band authentication.”

Please refer to Shared Services in this chapter for more information about MFA services and integration with EUA, IDM, and SLS services.

Self-Service and Help Desk Identity Life-Cycle Management

Self-service and Help Desk support must be available to subscribers for the following functions: recovering a forgotten user ID or password, enabling access to a temporarily disabled account, resetting a password, and updating user profile information. Restoration of access to a revoked account must be performed exclusively by a Help Desk agent.

Non-User Accounts and Special-Purpose ID Types

In addition to user accounts, Table - Special-Purpose ID Types presents the accounts for special-purpose ID types and a discussion of each.

Table - Special-Purpose ID Types
ID TypeDiscussion
TrainingTraining accounts are used for training users. Training accounts are most commonly used on training workstations. Passwords for these accounts are changed frequently by EUA scripts. Access Control Entitlements are assigned explicitly as needed.
TestingTesting accounts support testing of applications and other systems. ACEs are assigned explicitly or via ESS Job Codes as needed.
SystemSystem accounts are used by “superuser” administrators, by services or daemons that run on individual machines, or by services to authenticate communications between services. ACEs are assigned explicitly. Procedures ensure only authorized individuals have knowledge of or access to passwords for system accounts.

The CMS ARS requires periodic resetting of passwords for non-user accounts. New passwords are generated on schedule, coordinated with business owners or training facility managers, and put into effect through procedures managed by the EUA team (EUA UIDs), Tier 2 (AD), and Enterprise Database Group (EDG).

Reconciling Identities

When multiple directories or databases hold identity profile data on the same or overlapping sets of individuals, the identity profiles should be periodically reconciled between the directories and databases to detect and eliminate duplicates, outdated profile information, and invalid or unauthorized identities.

Identity reconciliation allows CMS a unified view of user / system identities by synchronizing identity records from multiple databases. It detects duplicate records existing across multiple access control systems or within one access control system itself. It correlates the data through data matching techniques that are used to identify duplicate records, inconsistent formats, and other data problems that can weaken system access controls for sensitive data for CMS.

Privilege Administration

Privilege administration is controlling user access to systems based on roles and responsibility. Privilege Administration consists of the following functions:

  • Identifying user privileges
  • Managing user roles
  • Granting user privileges and roles
  • Revoking user privileges and roles
  • Granting roles for operating system or network
  • Listing privileges and roles

Please refer to Shared Services in this chapter for more information about privilege administration in EUA, IDM, and SLS services, as well as Directory Services.

Periodic User Certification and Account Review

User Account Review, a.k.a. User Certification, is the periodic review of a user’s authorization status. User Certification is, at minimum, an annual requirement. For each user of a CMS system, the process of certification verifies that the user:

  • Is still an active user
  • Is still authorized to use CMS systems
  • Has received the required security awareness training
  • Has reviewed and accepted the CMS Privacy Statement

CMS business owners implementing processes for authorizing user privileges and roles must also implement processes supporting User Certification. The EUA, IDM, SLS, and other CMS shared IDM and directory services may implement features to support User Certification processes, such as tracking User Certification expiration dates and compliance.

Privileged Access Management and the CMS Zero Trust Forge

Access Control (AC) Procedures require privileged accounts with elevated access to sensitive systems or data be limited to only the permissions necessary. This is known as Privileged Access Management (PAM). While CMS Hybrid Cloud provides “stock” roles for Role-Based Access Control (RBAC), these may be overbroad. This is a challenge in improving the Zero Trust maturity of systems. All ADO teams are encouraged to use the CMS Zero Trust Forge, which addresses this by guiding and streamlining the definition of custom, granular roles.

Directory Services

Directories are databases specifically designed to store and retrieve individual user profiles, including login credentials, demographic information, roles, and privileges.

Enterprise Directory Services

The CMS Enterprise Directory Services and the Enterprise LDAP data enable the recognition of a UserID as a user identity in applications throughout CMS. Attributes of user accounts managed by EUA are synchronized with data stores for the Enterprise LDAP, LDAP proxies, RACF and certain databases. Applications may access the Enterprise LDAP data to authenticate users, retrieve email addresses, or retrieve other general information about users.

Enterprise Directory Services make the following synchronized data elements and functions available to applications:

  • UserID
  • Physical token ID numbers (SecureID or other)
  • Password verification
  • LDAP groups
  • Assurance level (not presently implemented)
  • Email address

Other information is also available from the Enterprise Directories. The Enterprise Shared Services Group (ESSG) can provide a complete dictionary for the managed data in each directory domain.

CMS maintains multiple managed directory systems offering directory services to all CMS applications and application platforms.

LDAPS and LDAP Proxies

CMS applications may connect directly to the CMS Enterprise LDAP proxies to perform user authentication or access user attributes. An LDAP bind operation may be used to verify user credentials. Note: CMS does not permit anonymous binds to LDAP. Applications requiring direct LDAP access will be assigned system accounts. Only secure LDAP (LDAPS) is accessible to applications, and a security certificate is required to access the secure server.

Applications using LDAPS to authenticate users are responsible for implementing their own mechanisms for managing access requests, including authentication, session management, policy, client detection, naming, and logging.

By default, the Enterprise LDAP proxies are read-only to most applications. Application-specific LDAP domains may be established with TRB approval.

RACF

Mainframe applications may use RACF. When practicable, mainframe z/Linux and z/OS applications should use the same authentication and directory service mechanisms as mid-tier servers.

Reconciling Identities

EUA supports new user registration processes to detect whether a new applicant already has an existing active, disabled, or archived UserID. The UserIDs in EUA are never reissued; a returning former employee or contractor whose UserID was archived is issued a new UserID.

EUA synchronizes passwords and core attributes of EUA-managed UserIDs in all of the supported domains and the EUA team reconciles non-EUA UserIDs in the domains with the EUA data repository to ensure only CMS-authorized UserIDs exist in the domains. All authorized UserIDs in the RACF and AD domains are also in the EUA data repository.

Major Shared ICAM Services at CMS

Enterprise User Administration

Enterprise User Administration is the CMS Identity Management system for CMS staff, contractors, and other government agencies. EUA user accounts are associated with CMS PIV cards and required for access to most CMS staff workstations and internal applications. EUA identity and access management is not available for all CMS application or data centers.

EUA Profiles and Roles

Each CMS user account may be authorized for and assigned one or more profiles. A profile (also known as EUA Job Codes, EUA Profiles, Production Control Identifiers, Business Profiles, or RBAC Profiles) is a specific set of Access Control Entitlements for a specific set of applications. Most profiles are unique to a given application, but a profile may also include ACE for multiple applications. When a user is assigned a profile, EUA may also initiate additional processes to provision or initialize application accounts or resources required for the user.

EUA manages user profile data and synchronizes profile data with the enterprise LDAP, RACF, and other CMS directories.

Privilege Administration in EUA

Job Codes are assigned to a user by a CAA only after a business owner and the user’s line manager approve a request. CAAs may use EUA Workflow to perform all user identity management functions, including adding new users, resetting user passwords, assigning Job Codes to users, and disabling, archiving, or restoring EUA user accounts.

Reconciling Identities

EUA supports new user registration processes to detect whether a new applicant already has an existing active, disabled, or archived UserID. The UserIDs in EUA are never reissued; a returning former employee or contractor whose UserID was archived is issued a new UserID.

EUA synchronizes passwords and core attributes of EUA-managed UserIDs in all supported domains. The EUA team reconciles non-EUA UserIDs in the domains with the EUA data repository to ensure only CMS-authorized UserIDs exist in the domains. All authorized UserIDs in the RACF and AD domains are also in the EUA data repository.

Identity Management System (IDM)

CMS’ Identity Management (IDM) system is an established, enterprise-wide, identity management solution. IDM is leveraged by CMS business applications across the agency. End users of all business applications that integrate with this solution can use a single set of user credentials to access any integrated application.

The previous Enterprise Identity Management (EIDM) has been decommissioned and replaced with IDM.

Scalable Login Systems

The Scalable Login System is an internal CMS web application that supports the FFM website, healthcare.gov. SLS provides the identity management of consumers who create accounts on healthcare.gov.

SLS is a database application that manages the life cycle of consumers’ User IDs, passwords, and supporting data collected from consumers, from initial creation until the account is archived. It provides the identity verification and account management of online accounts that consumers create on the FFM.

SLS collects and shares personally identifiable information with FFM to establish a consumer’s primary online account. This information includes a broad list of information such as the consumer’s name, address, telephone, employer, SSN, gender, ethnicity, citizenship status, household income, identifying information about any dependents or household members, and preferred language.

SLS services are grouped into two functional areas: the Registration service, which verifies each user’s identity through the New User Registration process using RIDP; and the Experian RIDP web service, which enables SLS to remotely verify the identity of the consumer applying for insurance. The Identity Life-Cycle Management Service provides self-service for the consumer-user, allowing the consumer-user to change a forgotten ID or password, enable a temporarily disabled account, restore access to a revoked account, and update the consumer-user’s profile.

Third-Party CSP and Related Services — A Layered Approach to Identity Verification

CMS uses several commercial services to support identity proofing and credential management. This RIDP solution includes several components, including Experian Precise ID, FraudNet, Boku and Ekata which all work together to provide a comprehensive and effective RIDP solution.

Experian Precise ID

Experian’s Precise ID helps CMS verify the identity of individuals by matching personally identifiable information (PII), such as name, address, and date of birth, against a variety of authoritative data sources.

For non organization users, some CMS applications use Experian’s solution for Remote Identity Proofing (RIDP) to validate that sufficient information exists to uniquely identify an individual without requiring document upload or biometric verification.

Importantly, Experian’s RIDP does not support document-based verification (e.g., uploading ID documents), facial recognition, or liveness checks. It avoids the document intensive approach of standard IAL2 because those methods typically result in lower pass rates. Instead, RIDP relies solely on PII and risk analytics, enabling higher verification success while still achieving strong fraud detection.

FraudNet

FraudNet is focused on detecting and preventing fraudulent activity. It uses advanced analytics and machine learning to analyze patterns of behavior and identify potential fraudsters. FraudNet can flag suspicious activity, such as multiple attempts to use the same identity information or attempts to use stolen or fake identities. It is important to note that your Application Maintainer will need to add Experian-provided Javascript collector software to User Interface (UI) to gather information from device for use by FraudNet.

Ekata

Ekata provides additional data to help verify identities. It draws from a variety of sources, such as address directories, public records, and phone directories, to provide CMS with additional information that can help confirm an individual’s identity. More importantly, by analyzing the metadata associated with the email address, Ekata can confirm that the email belongs to the person claiming to be associated with it.

Twilio

Twilio is focused on mobile identity verification. It leverages mobile phone information, such as phone number and carrier, to help verify an individual’s identity.

Taken together, these components provide CMS with a powerful RIDP solution that can help protect against fraud and ensure that the people they are dealing with are who they claim to be.

Federal Data Services Hub

The Federal Data Services Hub is used by systems like the Health Marketplace and Medicaid systems to verify eligibility. It consolidates access to multiple federal databases (SSA, DHS, IRS, VA, and others) and supports Medicaid and health insurance marketplaces by providing real-time verification of key eligibility criteria, including U.S. citizenship.

Taken together, these components provide CMS with a powerful RIDP solution that can help protect against fraud and ensure that the people they are dealing with are who they claim to be.

Okta Identity and Access Management Services

Some CMS applications use Okta identity and access management services for web-based applications, both in the Cloud and behind the firewall. All access to, and communication with, the Okta service is over an HTTPS connection. The Okta service offers:

  • Single Sign-on. Okta supports SSO to web apps in two ways: secure web authentication (SWA) and federation.
  • Directory Integration. User management may be integrated with AD or LDAP, including user provisioning and de-provisioning. Okta refers to this service as a “delegated authentication” where it uses AD or LDAP server credentials to authenticate users.

Access Control

CMS ARS guidance on the Access Control (AC) policy family is self-explanatory. The following subtopics provide an overview of some of the CMS services and practices available to support AC controls.

Please consult the CMS ARS directly for specific guidance.

Local System Access Controls

Local system security requirements are essential prerequisites for local Access Controls described in the CMS ARS. Please refer to the Security Services section in this chapter for a more comprehensive discussion of CMS security requirements.

Account Management

Account control mechanisms are required to authorize and monitor the use of guest / anonymous accounts, and to remove, disable, or otherwise secure unnecessary accounts. The CMS ARS requires removing or disabling default user accounts, renaming active default accounts, and using unique and separate administrator accounts for administrator and non-administrator activities.

Account managers must be notified when CMS information system users are terminated or transferred and associated accounts are removed, disabled, or otherwise secured. Account managers must also be notified when there are changes in the users’ information system usage or need to know.

Terminated or obsolete user accounts must be disabled and, after a defined period, archived. User accounts are never deleted because this action may impact an application’s data or logs. Applications should be designed to allow for such a possibility by recording a user’s full name and other pertinent information rather than relying on the future availability of that information in the directory.

Access Enforcement

Access control policies (e.g., identity-based policies, role-based policies, and rule-based policies) and associated access enforcement mechanisms (e.g., access control lists, access control matrices, and cryptography) are employed to control access between users (or processes acting on behalf of users) and objects (e.g., devices, files, records, processes, programs, and domains) in the information system.

Operating system controls must be configured to disable public “read” and “write” access to files, objects, and directories that may directly impact system functionality and/or performance, or that contain sensitive information. The information system must restrict access to privileged functions (e.g., system-level software, administrator tools, scripts, and utilities) deployed in hardware, software, and firmware, and security-relevant information shall be restricted to explicitly authorized individuals. Default access will be set to “denied” unless explicitly allowed.

Encryption of Stored Information

If encryption is used as an access control mechanism, it must meet CMS-approved encryption standards (FIPS 140-2 & FIPS 140-3 compliant and use a NIST-validated module). (Please refer to SC-13, Use of Cryptography, in the CMS ARS )

CMS generally does not require that its applications include specific encryption functions (e.g. FIPS 140-2 or FIPS 140-3) as part of the business logic. Encryption and encryption key management are often provided through configuration of the servers and infrastructure services supporting the application but this should not be assumed. CMS application owners must ensure that required encryption and key management controls are in place.

Information Flow Control

Information flow control regulates where information is allowed to travel within an information system and between information systems (as opposed to who is allowed to access the information) and without explicit regard to subsequent accesses to that information. The CMS TRA Multi-Zone Architecture includes many of the safeguards that are necessary for Information Flow Control.

Flow control must be enforced over information between source and destination objects within CMS information systems and between interconnected systems based on the characteristics of the information. (Please refer to Security Control AC-4 in the CMS ARS.) System policies and application interconnection agreements should address the types of permissible and impermissible flow of information between information systems and the required level of authorization to allow information flow.

System Use Notifications

The system displays an approved warning / notification message on successful log on and before gaining system access. The warning message notifies users that the CMS information system is owned by the U.S. Government and describes conditions for access, acceptable use, and access limitations. The system use notification message provides appropriate privacy and security notices and remains on the screen until the user takes explicit actions to log on to the CMS information system.

In high-risk systems, mechanisms must be in place to provide users with information about previous logons, both successful and unsuccessful.

The CMS ARS provides important additional requirements for the content and implementation of System Use Notifications.

Session Control

Information systems must identify and terminate all inactive remote sessions (both user and information system sessions) automatically.

The CMS ARS has the following Implementation Standard(s):

  1. Configure systems to disable local access (i.e., lock the session) automatically after a period of inactivity specified by CMS ARS Security Control AC-11 (Device Lock). Require a password to restore local access.
  2. Configure the information system to automatically terminate all remote sessions (user and information system) after period of inactivity specified by CMS ARS Security Control AC-02(05).
  3. Concurrent User ID network log-on sessions are limited to one (1); however, the number of concurrent application / process sessions is limited to what is expressly required for the performance of job duties and must be documented in the System Security Plan if it is more than one (1) concurrent session.

System and Communications Protection

CMS ARS guidance on the System and Communications Protection (SC) policy family is self-explanatory. The following subtopics provide an overview of some of the CMS services and practices available to support SC controls.

Please consult the CMS ARS directly for specific guidance.

Public Key Infrastructure Certificates

All public key certificates used within the CMS information system must be issued in accordance with a defined certification policy and certification practice statement. Registration to receive a public key certificate must include authorization by a supervisor or a responsible official and must be conducted by a secure process that verifies the identity of the certificate holder and ensures that the certificate is issued to the intended party.

Mobile Code

Mobile code (also known as active content) refers to macros or code embedded in a web page, email message, document, or other communicated medium that executes on a user’s device (workstation, cell phone, other). Mobile code technologies include, for example, Java, JavaScript, ActiveX, PDF, Postscript, Shockwave movies, Flash animations, and VBScript.

CMS establishes usage restrictions and implementation guidance for mobile code technologies based on the potential to cause harm to CMS information systems. The organization must document, monitor, and implement controls for the use of mobile code within the CMS information system. The TRB has the authority to permit or deny the use of mobile code.

Usage restrictions and implementation guidance apply to both the selection and use of mobile code installed on organizational servers and mobile code downloaded and executed on individual workstations. Control procedures prevent the development, acquisition, or introduction of unacceptable mobile code within the information system. NIST SP 800-28 provides guidance on active content and mobile code.

Audit and Accountability

CMS ARS guidance on the Audit and Accountability (AU) policy family is self-explanatory. The following subtopics provide an overview of some of the CMS services and practices available to support AU controls.

Please consult the CMS ARS directly for specific guidance.

As appropriate and defined in the CMS ARS, logs should be collected, aggregated, and analyzed in the Security Zone. Security documentation should define the events, categorization, and required responses. Each security event should document the corresponding actions taken. To ensure proper analysis of audit / logging information, a network time protocol should be used to ensure synchronization of time stamps across the environment.

Logging must be enabled for administrator and user account activities, failed and successful log on, security policy modifications, use of administrator privileges, system shutdowns, reboots, errors, and access authorizations.

Applications must include audit mechanisms that support the following CMS ARS requirements:

  • Disclosures of sensitive information, including protected health and financial information, must be recorded. The log should include information type, date, time, receiving party, and releasing party. Verify every 90 days for each extract that the data is erased or its use is still required.
  • For most auditable events, the audit record content must include (1) date and time of the event; (2) the component of the information system (e.g., software component and hardware component) where the event occurred; (3) type of event; (4) user / subject identity; and (5) the outcome (success or failure) of the event. NIST SP 800-92 provides more detailed guidance on computer security log content and management.

The selection of auditable events must be based on a risk assessment to determine which events require auditing on a continuous basis and which events require auditing in response to specific situations. The CMS ARS implementation standards for AU-2, Event Logging, provide a list of auditable events. NIST SP 800-92 provides guidance on computer security log management.

Audit Monitoring, Analysis, and Reporting

Information system audit records must be reviewed and analyzed regularly to identify and detect unauthorized, inappropriate, unusual, and/or suspicious activity. Such activity must be investigated and reported to appropriate officials, in accordance with current CMS Procedures. (Please refer to Security Control AU-6 in the CMS ARS.)

Automated mechanisms must be established and supporting procedures developed, documented, and implemented effectively to enable human review of audit information and the generation of appropriate audit reports. (Please refer to Security Control AU-7 in the CMS ARS.)

Audit Record Retention

Audit records must be retained to provide support for after-the-fact investigations of security incidents, and to meet regulatory and/or CMS information retention requirements. The National Archives and Records Administration maintains criteria for record retention across many disciplines.

The CMS ARS implementation standards for AU-11 and AU-3(3) include guidelines for PII, PHI, and other record types. The organization retains audit records for ninety (90) days and archives old records for one (1) year to provide support for after-the-fact investigations of security incidents and to meet regulatory (e.g., Federal Rules of Evidence) and CMS information retention requirements.

The data centers routinely back up and retain database data, operating system logs, all data storage, and logs from communications and infrastructure services. If an application creates custom application-specific logs or audit trails, they must be identified in the System Security Plan and the Operations & Maintenance Manual, including any unique backup and retention requirements. For example, if an application uses the operating system logging service, or the WebSphere Application Server logging service to create an application-specific log, that log may not be backed up unless it has been identified in the System Security Plan.

Business Rules for Access Control and Identity Management

CMS developed the following business rules to guide the development of the Agency’s Access Control and Identity Management (ACID) architecture and implementations. These business rules and policies are based on the CMS ARS and in part on the following guiding premises:

  • An individual who accesses CMS systems should have only one identity known to CMS. That individual may or may not have multiple credentials for accessing CMS systems, but all such credentials must be associated with a single vetted identity.
  • Business application owners are responsible for implementing procedures and mechanisms to authorize individuals for application-specific roles.

BR-ACID-1: Valid Purpose Required to Access CMS Information Systems

CMS must have an appropriate degree of confidence that its information systems are used only by those individuals with a valid business purpose.

Related CMS ARS Security Controls include: AC-3 - Access Enforcement and
AC-2 - Account Management.

Rationale:

Individuals accessing CMS information systems without a valid purpose are a potential threat to CMS services and data.

BR-ACID-2: Known Identity Required to Access CMS Information Systems

CMS must know with an appropriate degree of confidence the identity of every individual accessing CMS systems.

Related CMS ARS Security Controls include: AC-2 - Account Management, IA-2 - Identification and Authorization (Organizational Users), and AU-10 - Non-Repudiation (High).

Rationale:

Except for public information and services intended for anonymous users, unknown individuals accessing CMS information systems are a potential threat to CMS services and data.

BR-ACID-3: Single Identity Record for Each Individual Accessing CMS Systems

CMS will maintain for each user a single identity record to manage the user’s credentials for authentication to CMS networks and resources.

Related CMS ARS Security Controls include IA-2 - Identification and Authorization (Organizational Users), AC-2 - Account Management, and AU-10 - Non-Repudiation (High).

Rationale:

By maintaining a single identity record for each user, CMS can simplify the user experience, prevent duplicate identities, simplify identity management, and better understand who has access to CMS information systems.

BR-ACID-4: User Identities Must Be Vetted and Managed Using a Common Framework

The identities of every user must be vetted and managed using a common framework to ensure confidence in the user’s identity at each identity assurance level across CMS business applications.

CMS bases its policy for creating and managing identities for agency employees and support contractors on Homeland Security Presidential Directive 12 (HSPD-12), Federal Information Processing Standards Publication 201 (FIPS PUB 201), and the CMS ARS.

Above and beyond the core framework CMS provides for adjudication, the business owner is responsible for reviewing and approving any application-specific criteria and procedures used to vet the identities of new users of their business applications.

Related: OMB 11-11; HSPD-12; FIPS PUB 201; and CMS ARS Security Controls include:
IA-2 - Identification and Authorization (Organizational Users).

Rationale:

Identity vetting using certain criteria is required for compliance with NIST. Moreover, using a common framework for vetting enables CMS business owners to have confidence in the identity of established CMS users (those already having a CMS UserID) requesting access privileges to their business applications.

BR-ACID-5: Use Personally Identifiable Information Only When Appropriate

CMS will use Personally Identifiable Information only when appropriate.

Related CMS ARS Security Controls include: SA-8(33) - Minimization, PM-5(1) - Inventory of Personally Identifiable Information, PM-25 - Minimization of Personally Identifiable Information Used In Testing, Training, and Research, PT-2 - Authority to Process Personally Identifiable Information, PT-3 - Personally Identifiable Information Processing Purposes, PT-4 - Consent, PT-5 - Privacy Notice, PT-6(1) - Routine Uses, PT -7 - Specific Categories of Personally Identifiable Information, SI-12(1) - Limit Personally Identifiable Information Elements

Rationale:

Protecting PII held by the government is crucial to building and maintaining public trust. Since the purpose of Identity Management is to validate a personal identity, PII is essential to the process. CMS will support the government’s efforts to review and use PII only when appropriate.

BR-ACID-6: Minimize Retention of PII in Identity Life-Cycle Management

CMS will minimize retention of PII in identity life-cycle management, and only retain that PII needed for performance of identity life-cycle management functions after issuance of credentials, such as credential maintenance, account maintenance, authorization, or periodic certification.

Related CMS ARS Security Controls include: MP-6 - Media Sanitation, PT-2 - Authority to Process Personally Identifiable Information, PT-3 - Personally Identifiable Information Processing Purposes, SI-12 - Information Management and Retention, SI-12(3) - Information Disposal

Rationale:

Protecting PII held by the government is crucial to building and maintaining public trust. In some processes, such as in remote identity proofing, the PII collected may include both required proof of identity information (e.g., driver’s license, passport, and valid email address), and supplemental information used only to build confidence in an individual’s identity. Sources of supplemental information may include credit bureaus, public records, other agencies, or commercial providers. Once the individual’s identity has been established with sufficient confidence, the supplemental information is no longer needed, and only the required proof of identity information may be retained. Discarding this supplemental information reduces the potential for unauthorized disclosure and eliminates the expense to maintain and update it.

BR-ACID-7: Collect and Use Social Security Numbers Only When Necessary

CMS will collect, retain, use, or disclose Social Security Numbers (SSN) in identity proofing and management functions only when necessary and where no practical alternative exists.

Related CMS ARS Security Controls include: SA-8(33) - Minimization, PM-5(1) - Inventory of Personally Identifiable Information, PT-7(1) - Social Security Numbers, and SI-12(1) - Limit Personally Identifiable Information Elements

Rationale:

CMS may require an individual to provide a SSN to facilitate proofing at LOA 2 or 3. Without an individual’s SSN, identity proof will be more challenging and time consuming. Collecting SSNs may reduce the likelihood of fraud and abuse.

RP-ACID-8: Use Third-Party Data Sources for Identity Proofing and Credential Management

This recommended practice does not apply to vetting or identity management of organizational users (i.e., federal employees, contractors, administrators, etc.).

For non-organizational users, and where practical or as mandated, CMS will use third-party data sources (other government agencies, credit bureaus, commercial identity service providers, or commercial data brokers) to provide information and managed services that support identity life-cycle management functions such as identity proofing and credential management. For additional guidance, please refer to CMS ARS Security Control IA-8, which is the baseline for non-organizational user authentication requirements (i.e., “e-Authentication”).

Related CMS ARS Security Controls include: IA-8 - Identification and Authentication (Non-Organizational Users).

RP-ACID-9: CMS May Delegate Registration and Identity Proofing to Employers

This recommended practice may not be used for vetting or identity management of organizational users (i.e., federal employees, contractors, administrators, etc.).

For non-organizational users, and when appropriate, CMS may delegate the responsibility to perform registration and identity proofing to external organizations for their employees. For additional guidance, please refer to CMS ARS Security Control IA-8, which is the baseline for non-organizational user authentication requirements (i.e., “e-Authentication”).

Related CMS ARS Security Controls include: IA-8 - Identification and Authentication (Non-Organizational Users).

BR-ACID-10: EUA Manages UserIDs of CMS Employees and Contractors

The EUA system is the primary identity management system used to register organizational users and control issuance of UserIDs, passwords, and job codes for access to internal CMS applications (other than Internet-facing applications) for CMS employees, contractors, and other organizational users.

EUA identity and access management is not available for all CMS application or data centers.

Related CMS ARS Security Controls include: IA-2 - Identification and Authorization (Organizational Users) and AU-10 - Non-Repudiation (High).

Rationale:

Centralized identity management for CMS employees, contractors, and other organizational users accessing internal CMS systems is essential to protecting CMS services and data.

BR-ACID-11: CMS Business Owners Provide Privilege Administration

CMS business owners must develop their own processes for privilege administration, role administration, user provisioning of application resources, and related administration and management tools.

Privilege administration processes must establish that a user has an ongoing legitimate need to access a given application in an application-defined role. Identity proofing or possession of a CMS UserID does not represent vetting of an external user’s business relationship with CMS or other entities.

Related CMS ARS Security Controls include: AC-2 - Account Management and IA-2 - Identification and Authorization (Organizational Users).

Rationale:

Privileges and roles are very specific to each application and business. Only CMS business owners are in a position to define the privilege administration requirements and business processes for their applications.

BR-ACID-12: Local User and System Accounts Must Be Auditable

CMS requires strong technical and business reasons before permitting local user and system accounts on CMS processing systems. Like other accounts, local user and system accounts require management, auditing, and reconciliation.

Local user accounts are those that exist only in a single instance of an application (such as a database management system, business application, service daemon, infrastructure tool, or utility) or operating system on an individual server (either virtual or physical) and are not shared with or recognized by applications on other servers.

System accounts are used by superuser administrators, by services or daemons that run on individual machines, or by services to authenticate communications between services.

For all local user and system accounts, applications should use an Application LDAP domain, the CMS Enterprise LDAP domain, or a directory service synchronized with the Enterprise LDAP domain.

Exemptions to using an LDAP or other directory service for local accounts may be permitted for performance or other technical reasons. Exemptions must be documented in the system design documents, CFACTS, and risk assessment for both the system and the environment. Appropriate compensation information security and privacy controls must be implemented.

Related CMS ARS Security Controls include: AU-10 - Non-Repudiation (High).

Rationale:

Using a LDAP or other directory service helps support user account management, auditing, and reconciliation across CMS. Like other accounts, local user and system accounts require management, auditing, and reconciliation.

BR-ACID-13: OIT Is Responsible for Identity Management of Users with Credentials Provisioned in the CMS Enterprise Directory

For users with credentials provisioned in the CMS enterprise directory, including the CMS LDAP, Active Directory, and RACF directories, the Office of Information Technology and the Office of Support Services and Operations are responsible for identity proofing, identity management, and user credential management for users accessing internal CMS systems. Application owners are responsible for authorizing users to access their applications.

In some cases, the business owner may be responsible for proofing the user’s identity before making a request to a CMS Access Administrator to issue the user a CMS UserID.

Rationale:

Centralized identity management for CMS employees, contractors, and other users accessing internal CMS systems is essential to protecting CMS services and data.

 

NETWORK SERVICES 

CMS NETWORK SERVICES 

Network Services Introduction 

The CMS Processing Environment, which includes data centers, networks, applications, and cloud-hosted services, must comply with CMS Information Systems Security and Privacy Policy (IS2P2) and the CMS Acceptable Risk Safeguards (ARS). This section of the CMS Technical Reference Architecture (TRA) focuses on network services and provides supplemental engineering guidance for complying with the CMS ARS requirements for appropriate boundary protections (SC-7), monitoring (SI-4), malware protections (SI-3), protecting information at rest (SC-28) and in motion (SC-8), integrity (SI-7), remote access (AC-17), and other requirements. The reader should consult the IS2P2, the CMS ARS and the other topics of this chapter for more details and specific guidance on network service requirements. There are certain portions of this section that only apply to CMS data centers and not to CMS cloud implementations, where the network infrastructure would be a cloud-provided service. Clarification will be provided when the TRA content applies only to CMS data centers. 

Basic Description of the CMS Networking Environment 

The CMS networking environment includes the following components: 

1. The internal CMS Local Area Network (LAN) and Wi-Fi – connects personal computers and other devices in CMS office spaces and other CMS facilities. 

2. CMS Data Centers – provide gateways connecting all the CMS networks and provide hosting of CMS business applications. 

3. CMSNet – a secure private Wide Area Network (WAN) that: 

  • Provides connectivity to CMS applications and systems for employees and CMSNet Business Partners (CBP) 
  • Provides multi-zone connectivity for data hosting environments 
  • Procures services via a government-only contract vehicle 
  • Provides transport of data, Voice over Internet Protocol (VoIP), and video services 

4. CMS Extranet – a secure private WAN that provides connectivity to CMS applications for Extranet Business Partners (EBP). Services are procured via private commercial agreement among customers and CMS WAN provider at CMS’s discretion and approval. 

5. VPN – a Virtual Private Network (VPN) gateway providing secure Internet-based connectivity for servers. 

6. Internet – connectivity of CMS networks to and from the public internet is through a Managed Internet Gateway and the Department of Health and Human Services (HHS) Trusted Internet Connection (TIC), which: 

  • Provides access for consumers to various CMS-published resources 
  • Provides CMS employees, guests, and select CBPs access to Internet resources, i.e., cloud services providers (CSP), social media, streaming services 

The CMS networks support both Internet Protocols (IP) versions 4 and 6. Programs may use IPv6 transition mechanisms or native IPv6 transport on IPv6-compliant systems connected to CMS networks as described in the following chapter. 

Within the CMS data centers, the CMS TRA Multi-Zone Architecture, described in TRA Foundation, CMS Multi-Zone Architecture) provides Defense-in-Depth for computing environments. The multi-zone architecture, built upon a services framework, details a flexible architecture that can support both data center and cloud implementations. Within CMS data centers, the boundaries between presentation (edge), application and data zones are more clearly defined than in the cloud where some services may be supplied by the cloud vendor. For more detailed information, please review the Service Framework chapters within the Foundation section (see CMS Services Framework) (Internal Link). 

In CMS data centers, Application infrastructure resides in three zones — Presentation, Application, and Data, which span the CMS data centers. Security and Management Zones provide infrastructure and supervisory services to manage the zones. 

Firewalls and other security and monitoring components are located at appropriate boundaries between networks and zones. Transactions within a zone are permitted without restriction, unless the traffic is firewalled between data centers. Transactions traversing the zones are controlled and protected via firewalls and other security mechanisms. Transport and Management Zones provide infrastructure and supervisory services to manage the zones. 

CMS encourages resource sharing and reuse, and the concept of shared services. Applications can communicate within a zone both within a data center/cloud and between data centers or cloud regions. 

In a standard 3-zone data center implementation, users or resources granted access to a zone must meet the security challenges necessary for access to that zone. In addition, corresponding zones implemented in two (or more) data centers must have equivalent security configurations/posture. Hence, data zone components in one data center can communicate with data zone components within another data center, since their security postures are equivalent. 

In the cloud, where the network boundaries are less rigid, the zones may not be as clearly defined. In the cloud, protected application resources are generally implemented within a private cloud (non-public subnet), without a defined ‘zone’. Applications from CMS data centers or other applications within the cloud can be provided access to this private cloud with the appropriate network configuration. However, the same concepts of meeting security challenges and having equivalent security configurations also apply to inter-zone communication in CMS cloud implementations. The application developer will need to verify that security has been appropriately addressed when accessing cloud resources from a CMS data center or another CMS cloud implementation. In addition, applications granted an Authorization to Operate (ATO) can run in multiple data centers concurrently and share information between data centers as needed. 

Responsibilities for configuration and monitoring of applications and the network environment, as well as incident response, require coordination and cooperation between: 

  • Business application owners and system maintainers 
  • Data center and network operators and supporting contractors 
  • The CMS Cybersecurity Integration Center (CCIC) 
  • Trusted Internet Connection Access Provider (TICAP) 

CMS Stakeholder Requirements 

CMS’s stakeholders can be categorized into five communities of interest: CMS data center operators, CMS Offices, CMS Business Partners that include other government agencies, Extranet Business Partners, and beneficiaries. Each of these stakeholders has varying access requirements for the CMS applications, for inbound and outbound access to the Internet, and for communications among the various CMS business partners. 

CMS Data Center Operators 

The CMS data centers host CMS applications and provide services to CMS stakeholders. The CMS data centers also host the CMSNet, Extranet, and Internet gateways that serve as secure entry and exit points for all transactions and communications between the stakeholders. In providing this service, the data centers bridge and transport traffic between networks and facilitate seamless transaction processes. 

CMS Employees 

CMS employees access the Internet over a Managed Internet service. CMS employees at regional offices and satellite offices access the Managed Internet Gateway service over the CMSNet. 

CMS Business Partners 

The CMS Business Partners consist of the CMSNet Business Partners (CBPs), Extranet Business Partners (EBPs), and their communities. The following subtopics describe the network access services for each business partner type. 

CMSNet Business Partners 

CMS CBPs are the major users of the CMSNet. Some of the primary CMS CBPs include HIGLAS (Healthcare Integrated General Ledger Accounting System), Medicare Administrative Contractors (MAC), Unified Program Integrity Contractors (UPICs), Beneficiary Call Centers, Benefits Coordination & Recovery Center (BCRC) and CMS data centers. The CBPs use the CMSNet to access CMS business applications and to communicate with other business partners. CBPs access CMS applications through the CMSNet gateways hosted at the CMS data centers. The CMSNet Managed Service provides a secure site-to-site connection from CMS CBP locations to the CMSNet gateway at a CMS data center. CMSNet’s cloud-based infrastructure enables CMSNet sites the capability of access to any CMS data center or communication with other CBPs via logical configuration based on business requirements. CMSNet includes failover capability at data centers including multiple active-active connections or capability of routing to another data center. 

Extranet Business Partners 

CMS EBPs are other CMS business partners that require access to CMS’s non-web applications or sites. They include banks, clearinghouses, Managed Care Organizations (MCO), Coordination of Benefits (COB) contractors, and aggregators. 

EBPs are not permitted direct access to the CMSNet. Instead, EBPs are provisioned to a separate VPN cloud called the “CMS-Extranet,” established exclusively for use by EBPs. All traffic traversing the “CMS-Extranet” is encrypted according to Federal Information Processing Standards (FIPS) 140-2-compliant standards. 

CBP and EBP Communities 

A CBP can communicate with another CBP on CMSNet using a secure connection; however, CMS policy prohibits an EBP from communicating with another EBP on the Extranet. It is possible that a CBP on the CMSNet may need to communicate with an EBP on the Extranet and vice versa. When this occurs, the CBP traffic reaches the CMSNet gateway at a CMS data center over a secure connection on the CMSNet, and the EBP traffic reaches the Extranet gateway at a CMS data center over a secure connection on the Extranet. Traffic may be bridged between the CMSNet gateway and the Extranet gateway as long as the access policy permits the two business partners to communicate with each other. 

CMS Beneficiaries 

CMS beneficiaries use the Internet to access CMS business applications. Beneficiaries access CMS applications at a CMS data center via the Internet gateway. The CMS data centers are responsible for securing the Internet gateway and providing reliable service to the beneficiaries. 

CMS Network Services Overview 

This topic describes network services available to network users and application developers. 

CMS Internal Domains 

Primary and secondary DNS zones are distributed among CMS data centers to allow application and name resolution sharing across multiple sites. Please refer to the Domain Name System Services (Internal Link)  topics for detailed discussion of CMS DNS services topics for detailed discussion of CMS DNS services. 

Common Platform Services 

Common Platform services are technology services provided within the data centers in support of all hosted applications. Platform Services vary by data center hosting provider, and system maintainers must review specifics when selecting a provider. Platform services may include, but are not limited to, the following: 

  • Web Content Management 
  • Domain Name System – please refer to Domain Name System Services (Internal Link) 
  • Dynamic Host Configuration Protocol (DHCP) – configured for AWS Virtual Private Network (VPC) and for Azure Virtual Network (VNet) 
  • Network Time Protocol (NTP) 
  • Logging and log analysis – see CMS Hybrid Cloud guidance for Logging configurations and Splunk log ingestion 
  • Lightweight Directory Access Protocol (LDAP)— use CMS Active Directory E-LDAP 
  • Enterprise File Transfer (EFT) / Managed File Transfer 
  • Load Balancing 
  • Database Administration 
  • Storage Area Network (SAN) Administration 
  • Operating System (OS) Administration 
  • Message Queuing (MQ) 

Application developers should use Enterprise Shared Services and common platform services where possible rather than implementing specialized or single-use software. Enterprise Shared Services and common platform services ensure adherence to CMS standards and simplify support and management. Consult Required CMS Hybrid Cloud configurations for AWS Commercial accounts to ensure that platform services are configured correctly. 

CMS Continuous Monitoring 

CMS maintains an enterprise-wide, continuous monitoring program to improve situational awareness and provide near real-time risk management, compliant with NIST Special Publication (SP) 800-137, Information Security Continuous Monitoring (ISCM) for Federal Information Systems and Organizations, September 2011, and aligning with other SPs, including SP 800-37 and SP 800-53. In support of federal requirements to perform security and privacy continuous monitoring, CMS requires manual or automated audits, scans, reviews, or other inspections of CMS processing environments. CMS uses key data and metrics from many sources, including but not limited to, network monitors, asset management, vulnerability management, configuration management, malware detection, patch management, Security Information and Event Management (SIEM) integration, and other sources. 

CMS Continuous Monitoring Requirements are detailed in CMS Cybersecurity Integration Center Integration. 

Trusted Internet Connections 

CMS follows a Trusted Internet Connections deployment strategy aligned with HHS. The mission of the TIC program is to provide a means for monitoring, isolating, and securing federal external network connections in the event of cyberattacks. The program requires the Agency to route all traffic to and from external networks through TIC Zones managed by Trusted Internet Connection Access Providers. 

HHS provides TIC services for all HHS Operating Divisions, including CMS data centers and applications. The National Institutes of Health (NIH) operates the HHS TIC services. 

CMS application owners and data center operators are responsible for extranet connections to external networks, and for working with the HHS Office of Information Security (OIS) and NIH TIC teams to ensure the extranet connections are properly implemented and extranet traffic is monitored through HHS TIC services. This includes extranet connections to cloud providers and to other government networks. 

More information about the TIC program is available from DHS at Trusted Internet Connections. 

Email Services 

Various policies govern CMS email, Mail Transfer Agent (MTA), Simple Mail Transfer Protocol (SMTP), and other email-related services. These policies address concerns, which include: 

  • Insecure SMTP and MTA services in production data centers may simplify exfiltration of data by malicious agents or software. 
  • Official email from CMS must always have a “.gov” sending address. 
  • Official government correspondence, such as invoices and notices, must be retained in accordance with federal and CMS policies. 

For these reasons and more, CMS requires the use of  CMS Enterprise Email as a Service as opposed to Messaging Application Programming Interface (MAPI) and Internet Message Access Protocol (IMAP) protocols by business applications within a CMS data center or cloud environment. To send email, applications must use a CMS secure SMTP relay or a secure web service other than SMTP to request that a security-hardened email proxy server prepare and forward the outbound message to the CMS Enterprise Email service. 

Pursuant to Executive Order 14028: Improving the Nation's Cybersecurity, the use of secure SMTP relays is required as of September 1, 2024. 

For guidance concerning email to or from CMS applications, please refer to CMS TRA – Application Development section, business rule (BR) BR-SA-10, and NIST SP 800-45 and its supplement NIST SP 800-177. Related CMS ARS Security Controls include SI-8, Spam Protection. Security Controls include SI-8, Spam Protection. 

IPv6 

IPv6 is the next-generation Internet Protocol (IP), designed to replace version 4 (IPv4) that has been in use since 1983. IP addresses are the globally unique numeric identifiers necessary to distinguish individual entities that communicate over the Internet. The global demand for IP addresses has grown exponentially with the ever-increasing number of users, devices, and virtual entities connecting to the Internet, resulting in the exhaustion of readily available IPv4 addresses in all regions of the world. 

Over time, numerous technical and economic stop-gap measures have been developed in an attempt to extend the usable lifetime of IPv4, but all of these measures add cost and complexity to network infrastructure and raise significant technical and economic barriers to innovation. It is widely recognized that full transition to IPv6 is the only viable and sustainable option to ensure future growth and innovation in Internet technology and services. It is essential for the Federal government to consistently expand and enhance its strategic commitment to the transition to IPv6 in order to keep pace with and capitalize on industry trends. Building on previous initiatives which date back to 2003, the Federal government remains committed to completing this transition. 

The Federal government's IPv6 initiative has served as a vital catalyst, fostering commercial development and adoption of IPv6 technology. In the last five years, IPv6 momentum in industry has dramatically increased, with large IPv6 commercial deployments in many business sectors driven by reducing cost, decreasing complexity, improving security and eliminating barriers to innovation in networked information systems. Several large network operators, software vendors, service providers, enterprises, state governments, and foreign governments have deployed significant IPv6 infrastructures. In fact, many of these organizations have migrated, or are planning to migrate, to "IPv6-only" infrastructures to reduce operational concerns associated with maintaining two distinct networking regimes. 

Information and guidance on the IPv6 policies can be found at HHS Policy Transition IPv6 

Specific questions and information concerning the CMS IPv6 transition can be directed to the mailbox at IPv6_Transition@cms.hhs.gov 

 

WIDE AREA NETWORK SERVICES 

Introduction to Wide Area Networks 

CMS has established WAN architecture and design requirements to guide implementations and application development. These requirements help ensure that CMS Processing Environment contractors deploy implementations consistently in accordance with the CMS TRA vision. They also communicate system-level requirements to CMS application developers to fully inform them of these dependencies in developing their solutions. This topic presents CMS’s WAN architecture and design requirements. 

CMS WAN Architecture 

The CMS WAN consists of the CMS Private Network (CMSNet), QIES VPN, the Internet, and the CMS Extranet. Together, these WAN elements serve the communication needs of Agency data centers, business partners, employees, and beneficiaries. 

The CMS WAN consists of the following components: 

  • CMSNet – A secure private WAN that: 
  • Provides connectivity to CMS applications and systems for employees and CMSNet Business Partners 
  • Provides multi-zone connectivity for data hosting environments 
  • Procures services via government-only contract vehicle 
  • Provides transport of data, Voice over Internet Protocol (VoIP), and video services 
  • CMS Extranet – A secure private WAN that provides connectivity to CMS applications for Extranet Business Partners. Services are procured via private commercial agreements among customers and a CMS WAN provider at CMS’s discretion and approval. 
  • QIES VPN – A virtual private network providing secure Internet-based connectivity for the Quality Improvement & Evaluation System user community. 
  • Internet – Connectivity of CMS networks to and from the public internet is through a Managed Internet Gateway and the HHS Trusted Internet Connection Access Provider. The Internet: 
  • Provides access for consumers to various CMS-published resources. 
  • Provides CMS employees, guests, and select CBPs access to Internet resources, i.e., cloud services providers, social media, and streaming services. 

General Requirements 

The CMS TRA contains a high-level overview of the CMS WAN services. The CMS WAN requirements support the engineering details necessary to implement the CMS WAN infrastructure. These requirements will be implemented consistent with the CMS Information Security Policy Standards and Guidelines; the CMS ARS; DISA Network Infrastructure Security Technical Implementation Guide; and NIST SP 800-81, Secure Domain Name System (DNS) Deployment Guide. 

CMS Multi-Zone Environment 

The CMS WAN consists of multiple VRF segments that constitute the multi-zone environment and are referred to collectively as CMSNet. Each VRF in the WAN maps to a zone in the infrastructure architecture as defined in CMS TRA – Foundation. 

CMS Processing Environments consist of one or more of the following zones to provide Defense-in-Depth: 

  • Presentation Zones (PZ) – To support the presentation of content. Presentation Zones are accessible to external networks via firewalls through a Trusted Internet Connection. 
  • Application Zones (AZ) – To support business logic for applications and creating dynamic user presentations. 
  • Data Zones (DZ) – To contain data and data services used by applications. 
  • Management Zone – To support specialized services, such as Public Key Infrastructure, Domain Name System services, and system management services. 
  • Security Zone– To support security services. 

WAN connections between like zones in different data centers must maintain the same level of Defense-in-Depth protections. Various types of zone-to-zone interconnections are documented in this chapter. 

WAN connections must: 

  • Utilize mutual authentication and encrypted tunnels to secure zone-to-zone connections 
  • Include a Web Application Firewall and IDS / IPS capable of deep packet inspection at both ends of all tunnels 

WIDE AREA NETWORK SERVICES 

Business Rules and Recommended Practices 

Security 

The Security Services section in this chapter contains more detailed requirements that CMS contractors must follow for hardening the WAN components. 

BR-WAN-S-0: Use Mutual Authentication and Encrypted Tunnels between Data Centers 

Connections between CMS data centers must: 

  • Utilize mutual authentication and encrypted tunnels to secure zone-to-zone connections 
  • Include a WAF or firewall and IDS / IPS capable of deep packet inspection at both ends of all tunnels 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement, SC-28 - Protection of Information at Rest, SC-13 - Cryptographic Protection, and SC-8 - Transmission Confidentiality and Integrity. 

Rationale: 

The CMS ARS requires using appropriate approved cryptographic mechanisms (such as digital signatures and cryptographic hashes) to protect the integrity of data while in transit from source to destination outside of a secured network, to prevent unauthorized disclosure of information and to detect changes to information during transmission. 

BR-WAN-S-1: The WAN Must Implement FIPS 140-2 or FIPS 140-3 Compliant Encryption 

The WAN must implement an encryption capability that is FIPS 140-2 or FIPS 140-3 compliant using encryption products that have been validated under the Cryptographic Module Validation Program. 

When cryptographic mechanisms are needed, the information system uses encryption products that have been validated under the Cryptographic Module Validation Program (please refer to https://csrc.nist.gov/projects/cryptographic-module-validation-program ) to confirm compliance with FIPS 140-2 or FIPS 140-3 in accordance with applicable federal laws, Executive Orders, directives, policies, regulations, and standards. 

Related CMS ARS Security Controls include: SC-13 - Cryptographic Protection, AC-3 - Access Enforcement, AC-17 - Remote Access, AC-18 - Wireless Access, AC-19 - Access Control for Mobile Devices, AU-9 - Protection of Audit Information, CM-2 - Baseline Configuration, 
IA-7 - Cryptographic Module Authentication, SC-8 - Transmission Confidentiality and Integrity, SC-12 - Cryptographic Key Establishment and Management, and SC-28 - DoubProtection of Information at Rest. 

Rationale: 

In accordance with FISMA, the CMS WAN will implement the security controls based on the FIPS 199 categorization of the system and service. This general support system is considered HIGH and must comply with the following controls for encryption standards: SC-13 - CRYPTOGRAPHIC PROTECTION and associated controls: 

FIPS publications 

Cryptographic Module Validation Program 

BR-WAN-S-2: CMS Business Partners Only Access the Presentation Zone 

CMS Business Partners only access the Presentation Zone unless they have a defined business need to access other zones. Exceptions must be documented in the system design documents, CFACTS, and risk assessment for both the system and the environment. Appropriate compensating information security and privacy controls must be implemented. 

Related CMS ARS Security Controls include: AC-3 - Access Enforcement, AC-4 - Information Flow Enforcement, AC-6 - Least Privilege, CM-6 - Configuration Settings, and CM-7 - Least Functionality. 

Rationale: 

This business rule adheres to the zone requirements in CMS TRA Foundation, Processing Environments and Multi-Zone Architecture. 

BR-WAN-S-3: Communication between CMS Data Centers Is Only Permitted between Like Zones 

Communication between CMS data centers is permitted only between like zones (i.e., presentation to presentation, data to data). 

Related CMS ARS Security Controls include: AC-6 - Least Privilege, CM-6 - Configuration Settings, and SC-7 - Boundary Protection. 

Rationale: 

This rule maintains the Defense-in-Depth aspects of the CMS TRA Multi-Zone Architecture while extending the architecture across CMS data centers. It ensures that inter-zone transactions within a data center are only possible by using the required mechanisms and protections. Inter-zone transactions between data centers would add complexity and risk to the required inter-zone mechanisms and protections, create uncertainty over responsibilities between data centers, and may expose additional vulnerabilities. 

BR-WAN-S-4: Business Partner Access Restrictions 

The CMS WAN is designed to facilitate communication between CMS applications and CMS Business Partners. There are different restrictions for CBPs using CMSNet or the CMS Extranet. 

CMSNet CBPs are subject to the following restrictions: 

  • CMSNet CBPs may communicate with one another as determined by business requirements. 
  • CMSNet CBPs may access the Internet via CMSNet based on business requirements and as approved by CMS. 

Extranet Business Partners are subject to the following restrictions: 

  • CMS must approve any connection to specific location(s) or application(s) for which the EBPs have a business need. 
  • EBPs are only allowed to access the Transport Zone and Presentation Zone; access to other zones requires TRB approval. 
  • EBPs are not allowed access to the Internet via the CMS Extranet connection. 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement, 
AC-5 - Separation of Duties, and AC-17 - Remote Access. 

Rationale: 

Some CBPs may require Internet access or access to services from other CBPs to fulfill CMS contractual or other business obligations. 

In contrast to CMSNet CBPs, EBPs reside outside of the defined CMS WAN security boundary. Allowing EBPs to access the public Internet via CMS networks would present undesired extra costs (due to increased demand on bandwidth and other resources) and risks (due to EBP actions) to CMS. 

BR-WAN-S-11: Customer Edge Devices Must Be Configured Securely 

Customer Edge (CE) devices must be configured in accordance with CMS security policies, and secured in response to CMS, CISA Cybersecurity Advisors, and NIST National Vulnerability Database (NVD) advisories. 

Related CMS ARS Security Controls include: CM-6 - Configuration Settings, CM-7 - Least Functionality, and SI-2 - Flaw Remediation. 

Rationale: 

This rule adheres to CMS security policies, DISA STIG, (CIS) guidance, US-CERT and NIST NVD advisories, and best practices. 

BR-WAN-S-12: Management of CMSNet Is Via a Dedicated Logical Network 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement, 
SC-7 - Boundary Protection, SC-8 - Transmission Confidentiality and Integrity, and SC-32 - Non-Mandatory: Information System Partitioning. 

Rationale: 

NIST SP 800-53 rev 5 has a requirement for information system partitioning (SC-32) to assure Defense-in-Depth and minimize exposure of critical processes. Accordingly, management access is on a separate dedicated logical network. 

IP Addressing 

BR-WAN-IP-1: WAN Services and Devices Will Be Internet Protocol Version 6 (IPv6) Capable 

The CMS networks support both IPv4 and IPv6. Programs may use IPv6 transition mechanisms or native IPv6 transport on IPv6-compliant systems connected to CMS networks. 

an annual independent evaluation of the system must be assessed - Information Flow Enforcement. 

Rationale: 

This rule adheres to the OMB mandates for IPv6. 

BR-WAN-IP-2: The WAN Provider Will Use IP Space Provided by CMS 

The WAN provider will use IP space provided and maintained by CMS. 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement. 

Rationale: 

HHS provides CMS with IP space for use on the WAN. 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement. 

RP-WAN-IP-3: CMS or Data Centers Provide IP Space for the Multi-Zone Environment 

CMS or data center operators provide IP space for the Multi-Zone Environment. 

Rationale: 

The HHS-provided IP space is limited. 

BR-WAN-IP-5: Extranet Business Partners Provide Their Own IP Space 

EBP-provided publicly routable IP address space is required for EBPs to use the Extranet. 

Rationale: 

CMS IP space is reserved for entities that are connected directly to the CMS WAN (CMSNet). 

Change Management 

BR-WAN-CM-1: WAN-Related Service Requests Will Be Maintained on the CMS SOR 

All WAN-related service requests and management reporting related to service requests will be initiated and maintained on the CMS System of Record (SOR). 

Related CMS ARS Security Controls include: CM-3 - Configuration Change Control. 

Rationale: 

The CMS WAN Service Provider is responsible for maintaining all WAN components, configurations, and tools in accordance with CMS Change Management and Configuration Management processes. The CMS Network Service Provider is also responsible for tracking all updates to these systems for consistency across the CMS Processing Environments. 

BR-WAN-CM-2: CMSNet Services Must Be Certified Annually 

In accordance with FISMA mandates, an independent assessment of the system must be performed annually to maintain an Authorization to Operate (NIST 800-53 rev 5, CA-2 Security Assessments). 

Related CMS ARS Security Controls include: CA-2 - Control Assessments and CA-6 - Authorization. 

Rationale: 

As a federal agency, CMS and HHS compliance with FISMA is required for general support systems. 

BR-WAN-CM-3: Ensure Timely Version, Patch, and Configuration Management Practices Relative to the Identification and Release of New Security Features 

CMS WAN security objectives focus on protection of the WAN component hardware, software, and data traversing the WAN against threats to confidentiality, integrity, and availability. The security measures implemented must be commensurate with the function of the WAN component—i.e., Internet-facing components will require the highest level of security relative to other types of components. 

Multicast Routing 

BR-WAN-M-1: IP Multicast Routing Support across the CMS WAN 

IP Multicast Routing will be configured across specific CMS WAN elements as necessary, including CMSNet gateways, and CMS regional and satellite offices. 

To implement Multicast Routing across the WAN, CMS requires the following: 

  • Protocol-Independent Multicast-Sparse Mode (PIM-SM) must be enabled between the Service Provider PE and CMS CE router. 
  • PIM-SM must be enabled between the CMS CE router and the LAN switch. 
  • IP Multicast Boundary list must be implemented on the CMS CE routers to restrict certain multicast group addresses from traversing the WAN. 
  • PIM version 2 must be supported to provide for the distribution of Rendezvous Point information via Bootstrap Router (BSR) protocol. 
  • PIM-SSM (Source Specific Multicast) must be supported for a privately addressed multicast group range on the CE router. 
  • CMS sites that utilize Hot Standby Routing Protocol (HSRP) must use a dynamic routing protocol (OSPF) to support PIM. (PIM does not support scenarios in which infrastructure devices are pointing routes to a HSRP address.) This enables the use of PIM for IP multicast support and ensures routing redundancy. 

Related CMS ARS Security Controls include: AC-4 - Information Flow Enforcement, AC-6 - Least Privilege, CM-7 - Least Functionality, and SC-7 - Boundary Protection. 

Rationale: 

A variety of network-based applications such as Voice Over Internet Protocol (VOIP), audio and video streaming, and others require Multicast Routing. The CMS WAN must be able to support such applications for CMS and CMS Business Partners when there is a CMS business need. 

Other 

BR-WAN-O-2: Access to the Public Internet Will Comply With HHS TIC Policy 

CMS follows a TIC deployment strategy aligned with the Department of Health and Human Services, as required by the HHS Policy for the Implementation of Trusted Internet Connections (TIC) (Internal Link), March 21, 2021.  

Related CMS ARS Security Controls include: AC-17(3) - Managed Access Control Points and SC-7(3) - Access Points. 

Rationale: 

This rule adheres to guidance from OMB M-19-26, Update to the Trusted Internet Connections (TIC) Initiative, September 12, 2019, which rescinds previous memoranda providing TIC guidance, (M-08-05, M-08-16, M-09-32, M-16-27) and provides an enhanced approach for implementing the TIC initiative that provides agencies with increased flexibility to use modern security capabilities. This memorandum also establishes a process for ensuring the TIC initiative is agile and responsive to advancements in technology and rapidly evolving threats. 

 PREFERRED - Secure user access to CMS from Windows and Macintosh endpoints requires the Zscaler Client Connector (ZCC) agent on both GFE and contractor-owned systems. Connectivity from these clients uses the Zscaler Zero Trust Exchange. 

DOMAIN NAME SYSTEM SERVICES 

Introduction to Domain Name System Services 

This chapter presents a general overview of CMS practices and services for the Agency’s Domain Name System Services (DNS) architecture, design, and implementations supporting CMS enterprise environments. Specific guidance may be found in the CMS ARS, the CMS Information Systems Security and Privacy Policy (IS2P2), and in DNS Business Rules. .  

The DRaaS-CACHE Domain Name System provides a highly fault-tolerant and distributed DNS resolution environment for CMS. It includes physical and virtual appliances deployed in redundant fashion, accessible by all CMS serving entities on CMSnet. Public-facing DNS is managed by the Office of Communications and provided by Akamai Information about DNS requests is provided at CMS Hybrid Cloud: Introduction to DNS request. 

Additional references: 

DNS Business Rules 

Because internal DNS services are provided by DRaaS-CACHE, many of these rules apply to enterprise-level or datacenter-level DNS. Most are applicable for individual internal application systems that do not implement their own physical DNS infrastructure. 

BR-DNS-1: Keep Domain Names Flat 

CMS requires that all domain names be kept flat (close to root) and simple to all users. 

BR-DNS-2: The Physical Location of Name Servers in CMS Data Centers Must Be Easily Identifiable 

This is applicable in data centers that utilize physical equipment. 

BR-DNS-3: Minimal DNS Changes When Transitioning from One Environment to Another 

CMS must be able to transition applications from development to test, to implementation, and to production environments with minimal DNS configuration changes, i.e., no hard coding of domain names. 

BR-DNS-4: Use Stealth Primary Name Servers 

Stealth Primary name servers must reside in the Management Zone, must not have direct access to the Internet, and must not answer DNS queries, and must not answer DNS queries. 

BR-DNS-5: Secure Means Must Be Used for Zone Transfers 

BR-DNS-6: Provide Segregated DNS Resolution of Internal and External Queries 

Implement split zones to provide segregated resolution of internal and external queries. 

BR-DNS-7: Inter-Data Center Hosting of Secondary DNS Servers Is Required 

To ensure redundancy and geographic dispersion, at least one CMS data center DNS server must host or act as a secondary DNS server for another CMS data center’s DNS server. 

BR-DNS-8: Place Primary and Secondary Name Servers in Each CMS Data Center 

To improve fault tolerance and load sharing, strategically place primary and secondary name servers in each CMS Data Center. 

BR-DNS-9: Every “A” Record Defined Must Have a PTR Record 

This includes “AAAA” records for IPv6. 

BR-DNS-10: (Retired after TRA 2016R1): CMS must limit DNS zone data updates to trusted sources 

BR-DNS-11: (Retired after TRA 2018R1): DNS Protocol Must Be BIND 9 or Above and IPv6-capable (for Future Use) 

BR-DNS-12: CMS DNS Design Must Minimize Impact to the WAN 

BR-DNS-13: All DNS Changes Must Be Subject to CMS’s Change Management Procedures 

Information about DNS requests is provided at CMS Hybrid Cloud: Introduction to DNS request. 

BR-DNS-14: DNS servers deployed within ATO(ed) environments must use DNSSEC 

Per OMB Memorandum M-23-22, Delivering a Digital-First Public Experience, September 23, 2023, Domain Name System Security (DNSSEC) is required for all federal information systems. See Zone Transfers and Domain Name System Security for more information. While this practice is required in CMS production environments, maintaining signatures for DNSSEC zones adds some operational complexity that may not be justified for the lower environments of a FISMA system. 

 PREFERRED - Find detailed information about CMS DNS at CMS Cloud: Introduction to DNS request.  

DNS Services Infrastructure and Design Considerations 

The DNS design considerations presented in this topic supply the necessary engineering details used to implement the existing CMS DNS infrastructure. The requirements described in this topic ensure a simplified name space, provide a secure DNS query scheme among users in different contexts, and establish efficient and logical placement of DNS servers in appropriate locations within CMS infrastructure. 

These DNS requirements are consistent with the CMS Information Security Policy Standard and Guidelines Handbook, the CMS ARS, and NIST SP 800-81-2, Secure Domain Name System (DNS) Deployment Guide. 

CMS Domain Name Space Design Approach 

CMS DNS services are available to the Internet; all CMS data centers; the intranet; and CMS Presentation Zone, Application Zone, Data Zone, and Management /Security Zones within the standard CMS TRA Multi-Zone Architecture. The CMS DNS design is hierarchical, which helps achieve better security, ease of management, and scalability for future business needs. The name space design includes DNS functionality throughout CMS’s multi-zone infrastructure. 

Split DNS is used for CMS DNS Architecture. This means that the same domain name is used both internally and externally for accessible resources. Split DNS is implemented via views to separate external and internal records and provides differing access to the enterprise DNS for internal and external hosts. Splitting the internal and the external DNS servers also protects internal hosts and permits certain hosts to advertise to the Internet. Views facilitate the DNS administrator’s ability to define how a name server will answer queries based on the location of hosts and name servers requesting the information. 

See CMS Domain Name Spaces for more details 

CMS Internal DNS Services Infrastructure 

To support CMS’s internal networks and VDCs, CMS DRaaS-CACHE manages an enterprise-wide distributed system for IP Address Management (IPAM), DNS management, and IPv6 services. The system features: 

  • DNS and DHCP services for CMS workstations at the main campus and Regional Offices (RO) 
  • Integration with Windows Active Directory (AD) DNS services 
  • Servers and virtual appliances to provide replicated DNS resolution in all CMS environments, including AWS commercial and GovCloud; Microsoft Azure Government (MAG); and CMS on-premises data centers 
  • DNS firewall to detect and block malicious activity 
  • Centralized IPAM and DNS management 
  • Enterprise Failover Resiliency, Local Survivability Failover, and Disaster Recovery 

The system is configured to provide services through three network tiers: 

Tier 1 – Mission Control 

Authoritative name servers providing the actual answer to DNS queries such as mail server IP address or website IP address (a resource record) 

Tier 2 – VDCs and CMS HQ, Regional Offices, and Local Offices 

Acting as Internal Service Zone for internal DNS queries and Office Automation (OA). Tier 2 name servers will not respond to queries from external networks and will be configured for bi-directional synchronization with Windows AD. 

Tier 2 is further divided into separate network “grids” for OA and for each VDC. The OA grid includes CMS HQ, Regional Offices, and Local Offices. 

Tier 3 – HQ Caching Layer DNS Proxy to the Internet 

Cache-only DNS servers for public-facing DNS queries and DNS Firewall services. 

CMS Domain Name Spaces 

CMS maintains different domain name spaces for use in different CMS processing regions. These are: 

  • Public-facing (e.g., www.cms.gov, *.healthcare.gov, *.medicare.gov). These are generally program-specific and managed by the Office of Communications (OC) with Akamai services. The OC Web Help Service Desk supports these. CMS also has internal zones for some public domains. 
  • CMS Cloud internal (*.cmscloud.local) and The CMS Enterprise Cloud Service Desk supports these requests. 
  • CMS private (all other *cms.local). Requests for these should contact the CMS IT Service Desk through CMS Connect. 

CMS Internal DNS Name Space 

Primary and secondary DNS zones are distributed among CMS data centers to allow application and name resolution sharing across multiple sites. 

The CMS DRaaS-CACHE service maintains a distributed set of servers and virtual appliances across all CMS enclaves to provide continuous, redundant service. For Windows workstation environment. DNS services follow these strategic principles: 

  • DNS services for the Central Office and ROs are provided by CMS AD. 
  • Each “AD domain controller” is a DNS server and provides DNS services for its respective regional office. 
  • DNS zone transfers take place through AD replication that allows for secure and encrypted communication. 
  • Microsoft DNS integrates seamlessly with DHCP to provide encrypted dynamic registration of resource records for workstations and servers. 
  • At least one (1) domain controller / DNS server is allocated per RO. 
  • All clients point to a secondary DNS server for name resolution in the event of a failure, as follows: 
  • CO users point to a secondary DNS server in the CO. 
  • RO users point back to the CO for a secondary DNS server. 

Production DNS Name Space 

In production, CMS DRaaS-CACHE provides ubiquitous, redundant service across all enclaves. 

CMS External (Internet-Facing) DNS Name Space 

External DNS services are managed by the Office of Communications (OC) and provided by an external DNS service provider (Akamai). 

CMS Tier 3 name servers are configured to transfer their zone information to other name servers located on the intranet for external name space queries. 

CMS Development / Test / Implementation DNS Name Space 

The Development, Test, and Implementation Environments maintain their own DNS name servers to reduce management complexity. The extra zones ensure that developers and operational staff connect to non-ATO(ed) environments during Development, Test, and Implementation test activities. CMS DNS infrastructure is designed to ensure that business applications do not need to change configurations while transitioning between Development, Test, Implementation, and Production Environments. 

Multi-Zone Infrastructure 

Domain Name Services are available to the Presentation Zone, Application Zone, Data Zone, and Management / Security Zones within the standard CMS TRA Multi-Zone Architecture. External DNS services are available to the Presentation Zone for resolving connections to and from the private networks or other CMS ATO(ed) environments. The multi-zone DNS is segmented into multiple DNS sub-domains, such as ALOM and HIDS, to improve security. Internal DNS servers are located in the Management Zone for internal resolution of connections originating from the ALOM and HIDS network segments.  

Guidelines and Best Practices for DNS Implementation 

This topic presents the guidelines and operational best practices for execution of the CMS-approved DNS implementations to ensure efficient, reliable, and consistent DNS services for domain names and delegation. 

The following guidelines, representing industry and government best practices, apply to the CMS DNS infrastructure in general. This includes name space operations as well as management and delegation standards when deploying and maintaining the DNS infrastructure of the CMS ATO(ed) environments. 

Change Management 

CMS data center contractors manage all DNS zones, configurations, and tools in accordance with CMS Change Management and Configuration Management processes. For example, CMS requires a service request for any modification of static DNS record entries or DNS server configuration changes. CMS production environment contractors track all updates to these systems for consistency across the CMS ATO(ed) environments. To allow quick response and mitigate DNS security exploitation, it is good practice for the CMS CISO to retain the ability to change critical DNS records from a single point of control over top-level records. 

Management and Reporting 

The CMS production environment contractors currently provide integrated DNS system and error logs into other existing network management facilities to enable a real-time view of the Enterprise DNS for CMS Operations staff. 

Availability and Resilience 

CMS employs highly available zone transfers between data centers to facilitate availability and resilience in the overall DNS Architecture. Internal and external name servers in a given data center perform zone transfers to corresponding internal and external name servers at another data center to provide redundancy in the event of a primary name server failure. For example, a primary name server hosted by a CMS contractor transfers its zone records to secondary name servers located on CMS premises to improve reduce round-trip time (RTT) and provide redundancy in the event of primary server failure. CMS highly recommends that each data center contractor provide regularly tested DNS disaster recovery plans that address all DNS outage scenarios. 

Guidelines for DNS Configuration Requirements 

  • Authoritative Name Servers 

CMS’s DNS Architecture employs two types of authoritative name servers: primary and secondary. To improve fault tolerance, both types of name servers are deployed in each CMS data center. It is important to note that while a name server acts as primary for one zone, the same physical server may also receive transferred zone information from other servers and act as secondary name server for the zone transferred. 

  • Name Servers 

A primary name server contains zone files created and edited manually by the zone administrator. Sometimes a primary name server allows the zone file to be dynamically updated by authorized DNS clients. A name server configured with this feature usually is called the primary name server. CMS currently employs dynamic DNS update features. For security reasons, the primary name servers hosting records for ALOM and HIDS zones must be deployed in each data center’s Management Zone . 

CMS also employs secondary name servers to facilitate lowest RTT, availability, DR, and COOP. Primary name servers are configured with the DNS NOTIFY feature to facilitate automatic zone transfers to secondary name servers. 

  • Stealth Primary Name Servers 

One name server feature that CMS should consider deploying is a Stealth Primary name server, which must be located in each data center’s Management Zone. The main purpose of a Stealth Primary name server is to provide a central point of administration for all externally facing DNS records that CMS has not delegated to a CMS production environment contractor. The main requirements for Stealth Primary name servers are that they: 

  • Must not have direct access to the Internet. 
  • Must not answer DNS queries. 
  • Must protect or hide internal hosts from the public in situations of either interrogation (query or zone transfer) or if the DNS service is compromised. 

Zone Transfers and Domain Name System Security 

Per OMB Memorandum M-23-22, Delivering a Digital-First Public Experience, September 23, 2023, Domain Name System Security (DNSSEC) is required for all federal information systems. DNSSEC provides cryptographic protections to DNS communication exchanges, thereby removing threats of DNS-based attacks and improving the overall integrity and authenticity of information processed over the Internet. CMS requires DNSSEC for DNS servers deployed within CMS ATO(ed) environments. 

In accordance with best practices, the DNS name servers must comply with the following standards from Requests for Comment submitted by the Internet Engineering Task Force (IETF): 

While this practice is required in CMS production environments, maintaining signatures for DNSSEC zones adds some operational complexity that may not be justified for the lower environments of a FISMA system. 

DNS security objectives focus on the protection of the DNS host platform, DNS software, and DNS data integrity from availability threats. The security measures CMS implements must be commensurate with the function of the DNS server, i.e., Internet-facing and authoritative DNS servers will require the highest level of security relative to other types of DNS servers. CMS requires adherence to the recommendations in NIST SP 800-81-3, Secure Domain Name System (DNS) Deployment Guide. 

DNS content exploitation focuses on malicious manipulation of zone transfer files. Mitigation techniques concentrate on restricting exchanges to only authorized entities and implementing integrity checks on zone files. Using secure means for zone transfers is mandatory and helps protect data integrity during exchanges. 

INFRASTRUCTURE SERVICES

Virtualization

Introduction to Virtualization

Scope

This chapter defines the Virtualization guidelines and standards for the Centers for Medicare & Medicaid Services (CMS) Processing Environments. It notes where these intersect with CMS Cloud capabilities and services, and provides links. It does not cover Virtual Desktop Integration (VDI) or other desktop virtualization technologies.

Objectives

CMS’s overarching goal for virtualization is to facilitate the Agency’s transition from managing individual data center components to managing pools of computing resources. The following goals are common to all server virtualization implementations:

  • Cost savings
    • Consolidating applications to run on fewer physical servers
    • Improving the efficiency of data center operations by increasing the utilization of computing resources
    • Simplifying the management and maintenance of infrastructure
    • Reducing “server sprawl,” which diminishes power requirements and server hardware footprint
  • Increased scalability and elasticity capabilities
  • Enhanced business continuity and disaster recovery (DR) capabilities

Server Virtualization

Server virtualization is crucial to ensuring the success of server consolidation projects that maintain the isolation of separate systems. This topic addresses the principles of server virtualization and applicable business rules (BR). It identifies the design and implementation requirements for consistency with the CMS Technical Reference Architecture (CMS TRA) and describes implementation options.

Although virtualization may be generally applicable within the CMS Processing Environments, there are a few situations for which virtualization may not be appropriate:

  • Need for specialized hardware
  • Need for extreme performance
  • Need for higher security

These exceptions are becoming less common over time.

CMS Cloud Server Virtualization

 PREFERRED

The CMS-preferred solution for virtual servers is to use AWS Elastic Compute Cloud (EC2) (password required) instances or Microsoft Azure VMs (password required), running Windows or Linux. For single-purpose services, it is often more cost-effective and flexible to use AWS or MAG integrated services.

One CMS strategic solution is the CMS Cloud Gold Image. Gold Image AMIs are standard, baseline operating system (OS) images for most common Linux and Windows versions. Updated monthly, these contain:

  • Current approved patches to resolve vulnerabilities
  • Latest DISA STIG or CIS configurations to ensure configuration compliance
  • Pre-installed Shared Service applications and agents that are utilized in the CMS AWS environment

Terminology

Despite the large numbers of products available for virtualization, there is significant commonality in the underlying architectural and operational requirements for the deployment of server virtualization. This commonality stems from the common goals of and architectural elements employed by server virtualization implementations. The terminology in the table below applies to all server virtualization implementations.

Table - Server Virtualization-Specific Terms
TermDefinition
ProductionRefers to a CMS Authorization to Operate (ATO)(ed) environment, or to components in a CMS ATO(ed) environment.
Virtualization ImplementationA method for providing virtualized computing facilities to an application. For the purposes of this chapter, it means UNIX, Windows, or mainframe virtualization.
Virtualization Host ServerA single physical hardware instance that supports multiple virtual machines running independently under a hypervisor.
ContainerA process running under the control of a container engine. Containers are not virtualized, but rather, operate under strict control of the operating system (OS) hypervisor (stricter than normal processes). Containers run on a common kernel and operating system, unlike virtual machines.
Container EngineA control process that manages operating system resources to create, destroy, and supervise containers. The container engine performs analogously to a hypervisor, but without the use of Central Processing Unit (CPU) virtualization technology.
Resource

Architectural elements required by a virtualization host server. Resources can be divided into the following four basic types—Memory, Interfaces, Processors, and Storage:

  • Memory – Volatile data storage that can be read, written, and erased by applications. Memory usage quotas can be enacted to limit memory usage on a per-application basis.
  • Interfaces
    • Network Interface – An interface to an Internet Protocol (IP)-based data network. This is as opposed to an interface to a Storage Area Network (SAN) or directly connected peripheral. Network interfaces can be divided into virtual and physical interfaces.
    • Virtual network interfaces share not only network interface cards (NIC) but the IP stack, which communicates with an IP data network external to the virtualization server. This single stack supports network address translation (NAT) of virtual applications and the shared NIC and IP stack. The virtual interface may be associated with a NIC or a Virtual Local Area Network (VLAN) on a NIC.
    • Physical network interfaces share only the NICs. Each application has its own IP stack and data link layer instance. The data link layer instance may be associated with a NIC or a VLAN on a NIC.
  • Processors – This resource type includes shared CPUs and special use processors,
  • Storage – Permanent data storage that can be read, written, or erased via an OS either directly through an OS interface or by other applications using an application programming interface (API) to the OS. Storage is divided into file systems and further subdivided into directory structures and their individual files. Access to these file systems, directories, or files can be controlled based on user or group. In addition, these file systems can have per-user or group-usage quotas enacted to limit file system usage. Storage can be backed up and restored either via networked or storage or backup servers, e.g., tape backup devices.

In some instances, OS resources can also be a constraint, such as user IDs, devices, IP semaphores, message queues, and other limited resources. These typically do not limit virtualization but may impact containerization.

InstanceAn individual physical or virtual instantiation of a resource, e.g., a single CPU on a virtualization server or a single NIC.
PoolA set of resource instances of similar type on a single virtualization server, e.g., a set of CPUs that can be used by an application.
OS VariantA specific make and version of an operating system, e.g., Solaris 10.
Oversubscribed Pooled ResourceA pooled resource is oversubscribed when the sum of all minimum subscriptions is greater than 100 percent of the total available resources in the pool. For example, if five (5) virtual machines (VM) share a processor pool and their minimum processor utilization is set to 25 percent, then the processor pool is 25 percent oversubscribed.
Hypervisor (a.k.a. Virtualization Controller)An OS variant that can be used to configure virtual instances of other OS variants and their global resource quotas, as well as to enforce these quotas. Global resource quotas pertain to all instances of all resources on the virtualization server. As a result, only one instance of the virtualization controller may be running on a single virtualization server.
Virtual Machine (a.k.a. Virtual Machine instance, Supervised OS, Guest OS)An instance of an OS variant running under the control of a hypervisor.
Management AccessAccess to a VM or hypervisor using an account with management privileges.
User AccessAccess to a VM using an account with user privileges.
Management ApplicationAn application that can only be run with management access. This includes management programs and OS.
User ApplicationAn application that can be run with either user or management access. These applications include business and office automation applications.
ProjectWork done to develop a single application. This may be done by multiple contractors; however, the project’s scope should be covered by a single statement of work.

Despite this commonality, there are significant differences between Solaris, Linux, Windows, and IBM mainframe virtualization implementations. Specifically, there are substantial differences in the hardware, software, and security architectures employed by these implementations. These architectural differences require identification of the hypervisor and network security boundaries with implementation-specific architectural elements. The table Virtualization Implementation-Specific Security Boundaries defines the terms for implementation-specific security boundaries.

Table - Virtualization Implementation-Specific Security Boundaries
Example Virtualization ImplementationHypervisor Security BoundaryNetwork Security Boundary
AWS Elastic Compute Cloud (EC2),
Microsoft Azure VM
Virtual Instance
  • User Data – Each virtual NIC is isolated by security group(s).
  • Management Segment – Virtual NIC for a Management Zone is isolated by cloud-specific rules.
  • Security Segment – Virtual NIC on Security Zone is isolated by cloud-specific rules.
VMware vSphere,
HP Virtualization Infrastructure,
IBM z/OS,
IBM z/VM
Blade Server or Mainframe
  • User Data – Virtual NIC on User VLAN mapped to physical cBlade NIC for inter-zone communications through firewalls.
  • Management Segment – Virtual NIC on internal Management VLAN mapped to physical cBlade NIC in external Management VLAN.
  • Security Segment – Virtual NIC on Security VLAN mapped to physical cBlade NIC in external Security VLAN.
Citrix XEN,
VMware VM Workstation,
Linux KVM,
Oracle VirtualBox,
Microsoft Hyper-V
Physical server
  • User Data – Virtual NIC on User VLAN mapped to physical NIC for inter-zone communications through firewalls.
  • Management Segment – Virtual NIC on internal Management VLAN mapped to physical NIC in external Management VLAN.
  • Security Segment – Virtual NIC on Security VLAN mapped to physical NIC in external Security VLAN.

Hardening

All CMS infrastructure must be hardened to CMS standards. The CMS policy on hardening is specified in CMS ARS Security Control CM-6.

Business Rules

Although server virtualization (SV) technologies are evolving at a rapid pace, CMS has established the following business rules and recommended practices (RP) to help the implementation of these technologies meet CMS’s needs. The following server virtualization BRs support the consistent implementations of these technologies.

BR-SV-1: Apply Separation of Duties to Virtualization Administration

CMS will maintain separation of administrative duties between the hypervisor administration and VM administration.

Related CMS ARS Security Controls include: AC-5 - Separation of Duties.

Rationale:

The responsibilities of the hypervisor administrator are different from the responsibilities of VM administration. Because the hypervisor administrator can create and destroy VM instances (and other virtual resources), their role is different and separate from the day-to-day administration of virtual machine instances.

In a Cloud Service Provider (CSP) environment, the CSP typically controls the hypervisor. In this case, the CSP application programming interface (API) allows authorized users to create and destroy virtual resources, effectively acquiring hypervisor administrator privileges.

BR-SV-2: Provide Hypervisor Root Access Only to Specific Administrative Accounts

CMS will not grant unrestricted root (or system administrator) access to production hypervisor servers to the developers or end-users. Instead, root access should be restricted to a specified list of administrative accounts.

Related CMS ARS Security Controls include: CM-05 – Access Restrictions for Change and AC-06 – Least Privilege.

Rationale:

Giving out unrestricted full root access to hypervisors allows users to alter system configurations and system audit logging controls, which CMS ARS Security Control CM-6 specifically prohibits. Limiting access helps limit the risk due to the highly sensitive nature of these accounts. (Please refer to National Institute for Standards and Technology [NIST] Special Publication [SP] 800-125.)

BR-SV-3: Different Administration Account on Blade Controllers and Hypervisors

Administration accounts on the blade controllers and hypervisors must be different from those used to administer VM instances.

Rationale:

Using different accounts to administer blade controllers and hypervisors helps to enforce separation of duties.

Related CMS ARS Security Controls include: CM-5 - Least Privilege and AC-5 - Separation of Duties.

BR-SV-4: Configure UserIDs and GroupIDs to Be Unique across the Processing Environment

User groups and usernames must be unique across all file systems to prevent unintended inter-VM, hypervisors, and (when possible) Host access.

Rationale:

Network-based file servers use the UNIX userid and group--id when accessing files. It is possible for two different users to have the same userid number when on two different machines. Files created by the first user on a network share would be accessible by the second user. Ensuring different user-ids avoids accidental disclosure through this mechanism. In addition, the use of consistent UNIX ids allows for more consistent audit log correlation.

When a userid is no longer needed (such as user leaving, etc.), the userids should be retired and not reused.

RP-SV-5: Maintenance Window Planning

Applications that have conflicting maintenance windows and uptime requirements should not be hosted on the same VM or hypervisor.

Rationale:

During planned maintenance, it may be necessary to reboot virtual machines. If applications have different maintenance windows, it may be very difficult to schedule maintenance. Similarly, if the hypervisor or VM software needs upgrades or patches (for example), it is important to determine the impact to applications before undertaking this maintenance. Testing and Training systems can have similar constraints, which makes sharing difficult.

RP-SV-6: Consider High-Availability Configuration

All virtualization servers and VMs should consider configuring high-availability and disaster recovery services to mitigate failure and meet business owner-defined availability limits.

Rationale:

High-availability configurations, such as the use of clustering technology or redundant servers, should be used to meet business-owner-defined availability limits. If no such requirements exist or the limits are sufficiently low, a non-redundant configuration may be used.

BR-SV-7: No Co-Hosting on Production and Non-Production Hypervisors

In a physical data center, production and Non-Production environment virtualization servers, and their associated VMs, must reside on separate physical server hardware or logical partitions (LPAR).

Rationale:

CMS prohibits commingled workloads, thereby jeopardizing production workloads. Instead, host non-production workloads on different hypervisors.

BR-SV-8: Do Not Oversubscribe ATO(ed) Environments and Management Zones

All VMs in ATO(ed) environments or in Management Zones must have adequate resource constraints (upper and lower bounds) to ensure that they do not adversely impact the performance of other VMs on the same virtualization server. Over-subscription of pooled or shared resources is not permitted.

Please refer to BR-SV-20 and BR-SV-21 for additional related guidance.

Rationale:

Oversubscription can jeopardize meeting Service Level Agreements (SLA). In Production applications, this can result in sub-par performance and potentially impact application security. This is critical because any system processing CMS data is a Production system.

Please refer to CMS TRA Foundation, CMS Processing Environments for a formal definition.

BR-SV-9: Use Storage Quotas for Virtual Machines

Storage quotas must be enacted on a per-user and group basis.

Rationale:

Using quotas helps prevent filling file systems and object storage, which can lead to service failure and raise vulnerability to a denial-of-service attack.

BR-SV-10: ATO(ed) Environment VMs and Resource Pools Must Not Be Shared between Zones

Rationale:

A virtual machine can only exist in one zone—otherwise compromise of one virtual machine can lead to compromise of multiple zones, nullifying the advantage of defense-in-depth architecture.

RP-SV-11: Collect Virtualization Performance Metrics

CMS must gather virtualization server and VM performance metrics on a consistent basis to ensure maintenance of proper resource allocation.

Rationale:

Along with traditional server performance monitoring, virtualization technology offers additional performance metrics that may be useful to monitor.

BR-SV-12: Perform Asset Management of Virtual Instances

All virtual instance resources must be tracked in an asset management database, provided either by CMS or the hosting provider, and must provide a cross-reference to the host server hardware and host operating system.

Rationale:

Virtual instances must be identified in a durable form (such as Object Identifier [OID]) to track the use of assets and perform crosswalks between security and audit logs and virtual instances.

Related CMS ARS Security Controls include: CM-8 - Information System Component Inventory.

BR-SV-13: Keep Forensic Evidence per CMS Security Rules

Before deleting or disposing of virtual machine instances, determine if image copies should be retained for forensic purposes.

Rationale:

A copy of the instance can be used for forensic analysis by CMS. CMS may require adherence to specific rules for evidence gathering. Consult with CMS security to determine proper chain of custody rules.

RP-SV-14: Use VM Configuration Templates

Rather than generating VM instances as one-off configurations, consider using virtualization templating technology to define and create instances from repeatable configurations.

Rationale:

Repeatability is desirable because it facilitates change and configuration management and allows for replication of configurations in other processing environments. Unlike Graphical User Interface (GUI) forms, templates can be stored in source control systems, allowing for better auditing and change control.

BR-SV-15: Production Management Zone VMs May Not Use IP Multipathing

IP Multipathing (IPMP) on UNIX/Linux/Solaris or Multipath Input/Output (MPIO) on Windows Host OS are not permitted in the CMS Production environment.

Rationale:

Multipathing is discouraged because Intrusion Detection Systems (IDS)/Intrusion Protection Systems (IPS) are sometimes not available to re-assemble a full multipath Transmission Control Protocol (TCP) session and, therefore, are unable to properly inspect traffic. As a result, multipathing can be used to obfuscate attacks.

BR-SV-16: Originate Administrator Access to Blade Controllers and Hypervisors from the Management Zone

All administration access to the Blade controllers and hypervisors must originate from the Management Zone only and must use a CMS-approved secure access method.

Rationale:

Ensuring that network access to blade controllers and hypervisors originates from the Management Zone helps thwart administrator access from within the Application Zone.

Related CMS ARS Security Controls include: CM-5 - Least Privilege and AC-5 Separation of Duties.

BR-SV-17: All Management Traffic Must Originate or Terminate in the Management Zone and Use Only Isolated and Protected Interfaces

  • All management network traffic must originate or terminate in the Management Zone.
  • Management traffic must only use management interfaces and VLANs.
  • Management traffic must be protected from inspection and tampering by other users.
  • Production environment virtualization server NICs must be Media Access Control (MAC) address locked to prevent spoofing.

Rationale:

This rule ensures that management access is exclusively performed from the Management Zone. It also reduces the likelihood that unauthorized access could be performed on the business application network. Management protocols may be used between a Zone and the Management Zone only. For example, one cannot run a local syslog server in the Application Zone because syslog is, presumably, a management protocol.

BR-SV-18: Hypervisor Access Is Permitted Only Via the Management Interface

Access to the hypervisor must be allowed only via the management interface on the hypervisor host.

Rationale:

Management of the hypervisor is the sole responsibility of the operations team, which uses the Management Zone to initiate access to hypervisors. Limiting access from the Management Zone on the management interface prevents malicious actors from accessing the hypervisor from within the Application Zone.

BR-SV-19: Separate Security Segment from All Other Management Zone Segments

Management interfaces serving the Security segment must be isolated from those serving all other Management segments, e.g., Backup segments.

When Application and Presentation Zone VMs use virtual network interfaces:

  • They must route to each other through the external switch, i.e., no loopback routing; and
  • They must have compatible ports, protocols, and services policies that can be enforced by an external switch and firewall.

Related CMS ARS Security Controls include: AC-06 - Least Privilege, SC-03 - Security Function Isolation (High), SC-03(02) - Supplemental: Access/Flow Control Functions, and SC-07 - Boundary Protection.

Rationale:

Network traffic between Zones must be available for inspection by security services not running on the hypervisor, such as Network-based Intrusion Detection System (NIDS), and must be filtered by a firewall or equivalent.

BR-SV-20: Oversubscription of Non-Production Instances Is Permitted

Oversubscription of non-production instances is permitted, with business owner approval.

Please refer to BR-SV-8 for additional related guidance.

Rationale:

Oversubscription of non-production instances is helpful in reducing the operational cost of lower environments. It is understood that concurrent use could result in sub-par performance.

BR-SV-21: No Business Applications May Run on the Hypervisor’s Host OS

The hypervisor must only be used for the operation and management of virtual machines.

Rationale:

The hypervisor is a high-value, high-impact information technology (IT) asset. Operating business applications on the hypervisor exposes the hypervisor to a larger attack surface and additional sources of instability. By focusing the hypervisor on a single task (virtualization management), it becomes more robust and secure.

BR-SV-22: Operate Applications under Application-Specific System Accounts

Each application running on an operating system instance must operate under an application-unique system account userid. In particular, applications may not operate under an end user ID or RACF ID. Likewise, end users must not log in using an application system account userid.

System accounts must have only the minimum necessary permissions and should not have permission to log in.

Related CMS ARS Security Controls include: AC-02(09) - Supplemental: Restrictions on Use of Shared Groups/Accounts, AU-10 - Non-Repudiation (High), AC-6 - Least Privilege (High and Moderate), and CM-7 - Least Functionality (High, Moderate, Low).

Rationale:

By employing application-specific system account userids, it is possible to use operating system controls to impose access control and have better specificity in audit logs. If users can log in using system IDs, this limits non-repudiation for actions taken while operating as that userid. 

Containerization

The use of operating system containers, hereafter simply “containers,” is becoming the de facto standard for application deployment. Containers are a different form of process isolation than virtual machines and play a somewhat different role.

Containerization is a class of technology that permits isolation to occur either on physical or virtual machines. Containers do not perform CPU virtualization but rather use the facilities of the operating system kernel to sub-divide the machine into containers. Roughly speaking, each container shares a portion of the operating system with the host but has control over their own resources such as filesystems, networking, and memory. Thus, a container has many of the same properties as a real machine but is implemented using the lower overhead of the operating system’s process mechanism instead of CPU emulation, simulation, or virtualization.

Since containers operate within virtual machines and are, in effect, inheriting many of the properties of a process, they have no special status within the CMS TRA. Containers must be deployed within the CMS TRA Multi-Zone Architecture. A single container may not incorporate more than one zone. Container technology must be hardened like any technology within the CMS Processing Environment consistent with the guidance of CMS ARS Security Control CM-6.

Containers are covered in more detail in CMS TRA Application Development, Containers and Microservices.

Network Virtualization

The server virtualization techniques described in this chapter offer numerous benefits. To achieve the full rewards of virtualization, CMS must also virtualize the network links supporting communications between the virtual and physical servers. This topic describes the virtualization techniques to accomplish that end.

One of the key benefits of network virtualization (NV) is that it provides efficient utilization of network resources through logical segmentation of a single physical network. Logical, secure segmentation helps CMS comply with regulations for resource and information security. CMS’s goal is to reduce total cost of ownership (TCO) by sharing network resources while still maintaining secure separation between distinct network segments either within the Agency (e.g., web, application and data zones) or among business partners.

Overview of Network Virtualization

As a general principle, CMS employs best-of-breed network components. The primary technologies involved in delivering network path isolation and virtualization techniques that helps maintain secure logical network segmentation are:

  • Generic Routing Encapsulation (GRE)
  • Virtual Routing and Forwarding (VRF) Lite (VRF-Lite)
  • Multiprotocol Label Switching (MPLS)
  • Virtual Local Area Network
  • Virtual Device Context (VDC)
  • Virtual Switching Systems (VSS)

These technologies help provide solutions that preserve the benefits of existing network design while introducing the capability of logical segmentation. They help carve the network into secure virtual networks by overlaying Virtual Private Network (VPN) mechanisms onto the existing Wide Area Network (WAN) or Local Area Network (LAN). These path isolation and virtualization techniques can also help address issues associated with deploying network-based services and security policies in a distributed manner.

Business Rules for Network Virtualization

CMS has established the following network virtualization business rules to accomplish consistent implementations of these technologies in accordance with CMS TRA guidance.

BR-NV-1: Use Highly Available Network Services to Implement Zone Separation

Each zone must use highly available (HA) network services to implement and maintain the network separation, security, and performance architecture, as prescribed in the CMS TRA for the specific zone.

Virtualization techniques discussed in this chapter, such as VRF-lite, MPLS, VLAN, or VDC, may also be used for path isolation and virtualization.

The following network services must be independently managed components (as defined in CMS TRA Foundation, Processing Environments):

  • Border routing
  • Network switching
  • Load balancing
  • Intrusion Detection and Prevention
  • Firewalling

These network services cannot be shared with other tenants of the data center.

Rationale:

CMS requires independently managed components to retain operational control and monitoring of the network. The CMS TRA recognizes that underlying technology may be shared, such as in the case of CMS Clouds; however, the configurations, logs, and operational data are CMS sensitive and may not be shared. In addition, CMS must maintain the ability to perform reconfiguration at will without interference with or interference from other tenants.

Mainframe Virtualization

The CMS strategy for Mainframe Consolidation and Virtualization requires a highly centralized and highly consolidated environment that employs server virtualization technology. This strategy will achieve a large reduction in infrastructure and support costs by improving resource utilization and consolidating management functions. The mainframe virtualization figure below provides an overview of the basic components that comprise mainframe server virtualization solutions.

Mainframe Virtualization (page 17)

Business Rules for Mainframe Virtualization

CMS derived the following mainframe virtualization (MV)-specific BRs from various sources, including the CMS ARS and CMS TRA – Foundation. Not all of the Server Virtualization Business Rules apply to the mainframe environment; therefore, only those Business Virtualization rules that do apply appear below.

BR-MV-1: z/VM Sole Configuration

z/VM must be the sole configuration for all logical partitions (hereafter referred to as an LPAR-z/VM).

BR-MV-2: z/Linux as Sole Guest for z/VM

z/Linux (hereafter referred to as a Virtual Machine) must be the sole guest operating system for all z/VMs.

BR-MV-3: No Nested z/VM Instances

Nested z/VM instances must not run on z/VMs.

BR-MV-4: Multiple Zones for a Single Application

Applications are not permitted to run n-zones (web front end / business logic / database) virtualized on a single LPAR-z/VM Production environment.

Mainframe Virtualization-Specific Terminology

The Mainframe Virtualization-Specific Terms table defines the following terms that are specific to virtualization of Mainframe server implementations hosting CMS applications.

Table - Mainframe Virtualization-Specific Terms
TermDefinition
Authorized Program Facility (APF)Selectively permits or denies program access to sensitive system functions.
ChannelA communication path from the mainframe to an external device such as disk or tape. I/O devices are attached to the channel subsystem through the control unit.
tChannel PathThe connection between the channel subsystem and an I/O control unit.
Channels Path Identifiers (CHPID)Uniquely identify channel paths and can be assigned to one or more LPARs.
Control BlocksA data structure that serves as a vehicle for communication in z/OS.
Direct Access Storage Device (DASD)A storage device (e.g., a disk drive) that is directly connected to and controlled by the Mainframe. Disk I/O functions are controlled by the Mainframe operating system.
Logical Partition (LPAR)A set of processor, memory and I/O interface pools on a mainframe that acts as a virtual hardware platform for operating systems or VMs.
LPAR HypervisorA native hypervisor that runs directly on the mainframe hardware as part of the Processor Resource/Systems Manager.
Multi-Level Security (MLS)A security policy that allows the classification of data and users on a system of hierarchical security levels combined with a system of non-hierarchical security categories.
Processor Resource / Systems Manager (PR/SM)The feature that allows the processor to use several z/OS images simultaneously and provides logical partitioning capability.
Resource Access Control Facility (RACF)An IBM Security Center application that protects resources by granting access only to authorized users of the protected resources.
Security Category (SC)A non-hierarchical category that corresponds to some grouping within an organization that has similar security access authorization within a security level.
Security LabelA combination of a hierarchical level of classification (security level) and a set of zero or more non-hierarchical categories (security category).
Security LevelA positive integer that defines the hierarchical degree of sensitivity of the data. Higher numbers correspond to higher sensitivity. z/OS System 10 supports 256 levels.
Storage Area NetworkA Storage Area Network is a network specifically dedicated to the task of transporting data for storage and retrieval. The SAN provides an architecture to attach remote computer storage devices to servers such that the devices appear as locally attached to the operating system. The remote system accesses storage as though it were a local disk.
z/VMA software hypervisor that runs as an OS instance in an LPAR.
zLinuxLinux operating system compiled to run on zOS LPARs or as VMs in a z/VM, simply denoted VM.

Mainframe Server Implementation and Security

The following describes the design and implementation rules for consistency with the CMS TRA.

Mainframe Virtualization Implementation Rules

The general rules apply to all environments and zones within an environment. In addition, rules for each environment and each zone within an environment are provided when required.

General Mainframe Virtualization Implementation Rules

The following implementation rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in all environments and zones:

  • All management access to VMs (e.g., administrator login), must be limited to a single physical management interface or management interface pool that is not accessible by user processes in a VM.
  • All network management access to z/VM instances must be from the Management Zone via a CMS-approved secure access method.
  • All network management access to the LPAR hypervisor must be from the Management Zone via a CMS-approved secure access method, e.g., via VPN into attached consoles.
  • Management accounts for LPAR-z/VM management access must be different from those used to access VMs.
  • Multiple environments must not share LPAR-z/VMs.
  • Management interfaces on LPARs must be configured to be in a unique LPAR Management VLAN accessible only from the Management Zone.
  • Management interfaces for z/VM instances must be configured to be in a unique z/VM Management VLAN accessible only from the Management Zone.
  • Management interfaces serving the Security segment must be physically separate from those serving all other Management segments (e.g., LPAR-z/VM management segments).

Non-ATO(ed) Environment Mainframe Virtualization Implementation Rules

The following implementation requirements apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in Non-ATO(ed) environments, i.e., Test, Development, and Implementation environments, and their zones. The general rules apply to all zones within the environments. In addition, rules for each zone within the environment are provided when required.

  • When multiple development projects share the same LPAR-z/VM, each project must reside in a different VM.

Data Zone Rules

The following implementation rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in non-ATO(ed) environments and their Data Zones:

  • Non-Production Data Zone servers must reside in unique VMs.
  • Non-Production Data Zone VMs must use VLAN virtual network interfaces mapped to individual or pooled NICs.

Application and Presentation Zone Rules

The following implementation rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in non-ATO(ed) environments and their Application and Presentation Zones:

  • Non-Production Application and Presentation Zone servers may share VMs.
  • Non-Production Application and Presentation Zone VMs may use shared VLAN virtual network interfaces.

ATO(ed) Environment Mainframe Virtualization Implementation Rules

The following implementation rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in ATO(ed) environments. The general rules apply to all zones within the environment. In addition, rules for each zone within the environment are provided when required.

  • Production LPAR-z/VMs must not be shared between zones.
  • Each Production application server must reside in its own dedicated VM.
  • All Production VMs must use VLAN-dedicated physical or pooled network interfaces.
  • Management interfaces serving the Security segment must be physically separate from those serving all other Management segments, e.g., Backup segments.

Mainframe Virtualization Security Rules

CMS derived the following Mainframe-specific security rules from security requirements presented in CMS TRA Network Services and the CMS ARS. The general rules apply to all environments and zones within an environment. In addition, rules for each environment and each zone within an environment are provided when required.

General Mainframe Virtualization Security Rules

The following security rules apply to LPAR- z/VMs and VMs running on LPAR-z/VMs in all environments and zones:

  • Management interfaces must use dedicated interfaces or interface pools with VLAN separation.
  • Management traffic must only use management interfaces and VLANs.
  • Management traffic for VMs on the same LPAR-z/VM may share the same VLAN.
  • Management traffic for LPAR-z/VMs must use a separate, dedicated VLAN, distinct from that used for VMs.
  • Security traffic must use a separate physical interface from other management functions, e.g., backup.
  • All management traffic must be encrypted and must only use SSH version 2 or Transport Layer Security (TLS), never telnet. Authentication must be provided by a Role-based Security (RBS) with authentication, authorization, and accounting (AAA) functions provided by a common AAA server. Encryption must be accomplished using Federal Information Processing Standards (FIPS) 140-2 certified modules (please refer to CMS ARS Security Control SC-13).
  • All management traffic must originate or terminate in the Management Zone.
  • VMs must have individual access rules and roles for authentication through an RBS.
  • All accesses and configurations must be in accordance with the CMS ARS.
  • Management accounts with administrative privileges in VMs must not have access privileges to LPAR-z/VMs.
  • All VMs and LPAR-z/VMs must be assigned unique IP addresses.
  • All VMs on a single LPAR-z/VM must be assigned IP addresses from a subnet unique to the LPAR-z/VM.

Non-ATO(ed) Environment Mainframe Virtualization Security Rules

The following security rules apply to LPAR-z/VMs instances and VMs running on LPAR-z/VMs in Non-ATO(ed) environments and their zones. The general rules apply to all zones within an environment. In addition, rules for each zone within an environment are provided when required.

  • Information associated with different projects must use project- and user-unique keys for all information protected by cryptographic hashes, rather than those inherited from VMs on which the business applications run.
  • Access to file systems associated with a project must be controlled through project-unique security categories to prevent direct access by members of other projects. These security categories must map to UNIX Group Identifiers (GID), with the User Identifiers (UID) of group members associated with the unique GID.

Data Zone Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in non-ATO(ed) environments and their Data Zones:

  • Non-Production Data Zone application servers must use separate VMs.
  • When Non-Production Data Zone VMs reside on the same LPAR-z/VM, RBS AAA functions must be provided by a common AAA server to ensure that management and user roles are unique for each VM.
  • Non-Production Data Zone VMs on the same LPAR-z/VM may route internally between each other on their z/VM instance’s subnet.
  • Non-Production Data Zone LPAR-z/VMs must be configured to filter all traffic at the IP subnet level.

Application and Presentation Zone Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in non-ATO(ed) environments and their Application and Presentation Zones:

When Non-Production Application and Presentation Zone application servers share VMs,

  1. They must use separate file systems; and
  2. They must use different UIDs and GIDs with
    1. RBS AAA functions provided by a common AAA server to ensure that management and user roles are common throughout a zone to prevent accidental or unauthorized access; or
    2. Application account passwords that are unique to the VM instance, e.g., apache account, to enforce separation of roles.

ATO(ed) Environment Mainframe Virtualization Security Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in ATO(ed) environments and their zones. The general rules apply to all zones within an environment. In addition, requirements for each zone within an environment are provided when required.

Production VMs must use:

  • RBS AAA functions provided by a common AAA server to ensure that management and user roles are common throughout a zone to prevent accidental or unauthorized access; or
  • Application account passwords that are unique to the VM instance, e.g., apache account, must be used to enforce separation of roles; and
  • Unique GIDs and UIDs across all file systems to prevent inter-VM access.
  • Crash dumps residing on a common log server must be protected to prevent accidental destruction or unauthorized access.
  • Each VM must implement the CMS ARS Security Controls as required by the CMS Security Level of the resident application (please refer to CMS Risk Management Framework (RMF): Categorize Step, Task C-2: Security categorization
  • Resident applications with different CMS Security Levels or information types may not share a VM.
  • When required by CMS ARS requirements on the resident application, each VM will implement a HIPS.
  • All LPAR-z/VM access, and management traffic to VMs, must be encrypted using FIPS 140-2 certified modules (please refer to CMS ARS Security Control SC-13).
  • When required by CMS ARS requirements on the resident application, each VM will implement a firewall configured to allow communication only with authorized VMs in other zones over approved ports, protocols, and services.
  • Access to file systems associated with a project must be controlled through project-unique security categories to prevent direct access by members of other projects. These security categories must map to UNIX GIDs, with the UIDs of group members associated with the unique GID.
  • Production LPAR-z/VMs must run at a higher security level than Non-Production environment LPARs.
  • Production and Implementation LPAR-z/VMs must not be accessible from Test or Development LPAR-z/VMs.
  • Firewall rules must be implemented to deny intra-zone, inter-VM network traffic by default and must allow only ports, protocols, and services authorized for a source and destination IP address pair to pass.

Data Zone Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in ATO(ed) environments and their Data Zones:

  • Production Data Zone LPAR-z/VMs must authenticate file systems before remount after VM crash.
  • GID and UIDs must be unique to each VM in a Production Data Zone.
  • Production Data Zone LPAR-z/VMs must run at a higher security level than all other LPAR-z/VMs on a Mainframe, except when Management Zone LPAR-z/VMs are present.
  • Production Data Zone application servers must reside on unique VMs.
  • Production Data Zone data sets, e.g., databases, must have unique security categories.
  • Processes associated with user applications must not be allowed to use APF.
  • Production Data Zone applications accessing Personally Identifiable Information (PII) or Protected Health Information (PHI) must run with a higher security category than those not requiring access to PII or PHI.

Application Zone Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in ATO(ed) environments and their Application Zones:

  • Production Application Zone LPAR-z/VMs must authenticate file systems before remount after VM crash.
  • Production Application VMs must employ RBS AAA functions that must be provided by a common AAA server to ensure that management and user roles are common throughout a zone to prevent accidental or unauthorized access.
  • An application account’s passwords for Production Application VMs must be unique for each VM instance, e.g., an apache account must be used to enforce separation of roles.
  • Production Application VMs’ GID and UIDs must be unique across all file systems to prevent inter-VM access.

Presentation Zone Rules

The following security rules apply to LPAR-z/VMs and VMs running on LPAR-z/VMs in ATO(ed) environments and their Presentation Zones:

  • Production Presentation Zone LPAR-z/VMs must authenticate file systems before remount after VM crash.
  • Production Presentation VMs’ RBS AAA functions must be provided by a common AAA server to ensure that management and user roles are unique for each VM.
  • Production Presentation VMs’ GID and UIDs must be unique across all file systems to prevent inter-VM access.
  • Production Presentation Zone VMs providing intranet and extranet access must not reside on the same LPAR-z/VM.
  • Production Presentation Zone LPAR-z/VMs providing extranet access can run at a lower security level than all other LPAR-z/VMs on the mainframe except those LPAR-z/VMs supporting publicly available Internet web sites.
  • Production Presentation Zone LPAR-z/VMs supporting a publicly available Internet web site can run at a lower security level than an extranet.

Cloud IaaS and PaaS Infrastructure Introduction

Forward

Topics in this chapter describe the definitions, principles, and rules surrounding cloud infrastructure. CMS promotes its own strategic implementations of managed cloud environments that are based on these principles and rules. These topics discuss aspects of CMS Cloud, and refer to relevant parts of this preferred implementation. Readers may find these a useful way to better understand CMS Cloud.

Background

As identified in NIST Special Publication (SP) 800-145, The NIST Definition of Cloud Computing, September 2011, the cloud operating model has the following essential characteristics:

  • On-demand self-service – A CMS business owner can provision computing capabilities, such as server processing capability and network storage, as needed automatically without requiring human interaction with each service’s provider.
  • Resource pooling – The provider’s computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to consumer demand. Examples of resources include storage, processing, memory, network bandwidth, and server virtual machines.
  • Rapid elasticity – Capabilities can be rapidly and elastically provisioned, in some cases automatically, to quickly scale out and rapidly released to quickly scale back. To consumers, the capabilities available for provisioning often appear to be unlimited and can be purchased in any quantity at any time.
  • Measured service – Cloud systems automatically control and optimize resource use by leveraging a metering capability at some level of abstraction appropriate to the type of service provisioned (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
  • Broad Network Access – Capabilities are available over the network and accessed through standard mechanisms that promote use by heterogeneous thin- or thick-client platforms (e.g., mobile phones, laptops, and PDAs).

Leveraging these cloud characteristics obligates CMS to provide its IT community appropriate rules and guidance. The CMS TRA has historically addressed the definitions and operations of the Virtual Data Centers (VDC). This chapter expands the CMS TRA data center hosting and managed services architecture to address evolving cloud environments.

Scope

As identified in the CMS Information Security and Privacy Group (ISPG) Cloud Computing Standard, the scope of CMS cloud standards and guidance spans the cloud deployment models of public, private, community, and hybrid clouds as well as the cloud service models of IaaS, Platform-as-a-Service (PaaS), and Software-as-a-Service (SaaS). This chapter, however, is limited to IaaS and PaaS Clouds.

The original Federal Cloud Computing Strategy (Kundra, 2011) defines a decision framework that should be considered for determining whether a CMS business application is a candidate for inclusion in a cloud environment. Some value drivers include:

  • Efficiency
    • Higher computer resource utilization (through virtualization)
    • Lowering labor costs
  • Agility
    • Speed of deployment
    • Rapid provisioning
  • Access to innovation

It is important to recognize that both value and readiness are provided as dimensions to help plan cloud migrations.

It is highly recommended that CMS business owners familiarize themselves with the Federal Cloud Computing Strategy as well as the Department of Health and Human Services (HHS) document, HHS Cloud Computing Tactical Implementation and Transition, v1.2.2b, Department of Health and Human Services, July 2012 because this CMS TRA chapter does not provide distinct selection criteria or guidance for deploying into a cloud environment.

PREFERRED

However, the CMS strategic solution is to use CMS Hybrid Cloud Services, which are implemented in compliance with the CMS TRA. Whether Amazon Web Services (AWS) or Microsoft Azure Government (MAG), solutions based on CMS Cloud are preferred. Platform & Hosting Solutions provides an overview of the current offerings. CMS has announced initiatives to support additional cloud service providers.

Items in Scope

This chapter represents the CMS guidelines and standards to be used by CMS and CMS / Contractor partners for CMS Processing Environments that are implemented and operated using private or community clouds.

This chapter contains CMS’s architectural standards employed in all CMS Processing Environments that operate in a private or community cloud, as defined by NIST SP 800-146, Cloud Computing Synopsis and Recommendations, and that are using existing and approved CMS cloud infrastructure (IaaS) and platform (PaaS) managed services.

This chapter complements the CMS ISPG Risk Management Handbook Volume III: CMS Cloud Computing Standard and NIST SP 800-144, Guidelines on Security and Privacy in Public Cloud Computing.

Any specific “requirements” mentioned within this chapter apply only to specific items of technology and reference architecture and do not eliminate, replace, supersede, override, or nullify published CMS minimum security or privacy requirements. The security or privacy requirements stipulated in the CMS ARS and RMH have precedence over any perceived conflict among requirements.

Items Out of Scope

It is important to note that selection criteria, assessment, and acquisition of Cloud Service Providers are out of scope for this chapter.

CMS business owners and their IT advisors must perform due diligence and assessment before choosing to host and operate their business application within a CMS data center or the cloud, based upon the application’s business requirements and the application’s security risk profile. The original Federal Cloud Computing Strategy (Kundra, 2011) as well as the HHS Cloud Computing Tactical Implementation and Transition document provide some guidance for determining cloud deployment suitability. The Department has provided some guidance regarding acquisition in section 4.6 of the HHS Cloud Computing Tactical Implementation and Transition.

This chapter does not focus on how to implement a cloud infrastructure, because this is the CSP’s responsibility. Rather, it addresses how CMS can work with a third-party CSP to securely and effectively leverage business value from the CSP’s services. It is anticipated that many of the best practices pioneered by CSPs will eventually find their way into traditional data centers.

Finally, this chapter does not address the use of Software-as-a-Service CSPs.

Cloud IaaS and PaaS Infrastructure

Cloud Environment Business Rules

This topic provides a core set of Cloud Infrastructure (CI) business rules that are binding on all CMS business applications hosted and operated in a cloud solution.

BR-CI-1: Choosing a Cloud Deployment and Service Model

CMS business owners will use the CMS Cloud Computing Standard, CMS Office of the Chief Information Security Officer, Risk Management Handbook Volume III, Version 1.0, May 3, 2011, as a point of reference in (a) choosing the cloud deployment model (private, community, hybrid) and service model (IaaS, PaaS, or SaaS) based upon the business application’s identified security risk; and (b) employing CMS Chief Information Officer (CIO)-approved Cloud Service Providers.

BR-CI-3: Document the Impact of Cloud Deployment

A business application identified for deployment to and operation in a cloud environment must include in its Requirements Document those specific functional, technical, performance, service, and support requirements of the cloud environment as mandated by the CMS Target Life Cycle (TLC). The business owner must also include documentation of the Analysis of Alternatives used to determine the suitability-for-the-cloud assessment in accordance with the guidance of HHS Cloud Computing Tactical Implementation and Transition.

BR-CI-4: Engage TRB Consulting for CMS-Owned Equipment

Business owners should consult with the CMS TRB when planning to install and operate CMS-provided Government-Furnished Equipment (GFE) hardware at the cloud service provider’s facility.

Rationale:

This rule is necessary because of the increased costs and complexity in managing physical devices at a cloud provider, which run counter to cloud efficiency principles. The increased complexity also increases security risk. An analysis of alternatives should be conducted to communicate and document choices.

Examples of CMS-owned GFE include networking equipment used to connect the cloud service provider’s network to the CMSNet.

BR-CI-5: Acquisition of New IaaS or PaaS Cloud Service Providers

The acquisition and procurement of an IaaS or PaaS Cloud Service Provider must follow the guidelines set forth in Federal Risk and Authorization Management Program (FedRAMP), Office of Management and Budget (OMB) memoranda, and other appropriate federal, HHS, and CMS guidelines. Business owners are not permitted to acquire such services directly.

Only a CMS CIO-designated Cloud Management organization or HHS Office of the CIO cloud authority may procure and manage cloud environments.

Related CMS ARS Security Controls include: CA-6 - Authorization.

Rationale:

Only the CMS CIO can issue an ATO for a CMS IT processing environment.

BR-CI-6: Define Data Backup and Contingency Plans in a Cloud

All CMS business applications must evaluate and establish that defined backup and contingency plans, including the desired Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO), are achievable and specified within established or planned Service Level Agreements (SLA).

Related CMS ARS Security Controls include: CP-6 - Alternate Storage Site, CP-7 - Alternative Processing Site, and CP-9 - System BackupCP-9(8) - Cryptographic Protection.

Rationale:

While it is true that many CSPs provide services that maintain multiple copies in multiple data centers, this does not constitute data backup but rather addresses reliability and availability concerns. Hosting in a cloud environment does not inherently provide data backup or contingency planning.

BR-CI-7: Cloud Resource Capacity Planning

Anticipated resource capacity needs, along with the expected need for elasticity, must be documented and communicated both in the TLC-specified document templates (such as the System Design Document) and to the CSP if these capacity levels are required. Automated actions required for exceeding those established ceilings must be established either through SLAs or with automated management rules/responses with the CMS management tools and previously agreed to by CMS Cloud as well as the business owner. This may incur additional charges to which business owners must agree. Capacity planning must be included as part of contingency plans as well.

Related CMS ARS Security Controls include: CP-2 - Capacity Planning.

Rationale:

Clouds are capable of fulfilling requests rapidly, which can potentially commit the government to unexpected or unforeseen charges. Capacity planning is necessary to maintain control over existing usage and forecast future usage, to inform business owners of costs and budgetary impacts. Additional guidance is available in NIST SP 800-34.

BR-CI-8: Separate Production, Management, and Non-Production Resource Clusters

Production and Non-Production environments must use distinct resource clusters, as defined above in Cloud Architecture, Resource Aggregation – Resource Clusters and Resource Pools. In addition, the Management Zone (for both Production and Non-Production) must have at least one resource cluster distinct from the Production and Non-Production environments.

Related CMS ARS Security Controls include: SC-6 - Resource Availability

Discussion:

This means there are at least three (3) resource clusters in any CMS TRA-compliant IaaS or PaaS cloud.

Rationale:

The Management Zone cluster is distinct from the Production and Non-Production clusters to ensure that there is no contention for resources.

BR-CI-9: Applicability of Multi-Zone Architecture

IaaS clouds must implement the CMS Multi-Zone Architecture as defined in the CMS TRA. PaaS clouds must be based on an infrastructure that meets the CMS Multi-Zone Architecture.

Rationale:

The CMS Multi-Zone Architecture can and has been implemented in IaaS clouds. The benefits of defense-in-depth architecture can be achieved using cloud technologies, even if the architecture is not an exact one-for-one match with virtualized data centers. For example, clouds may offer different services to perform the role of the firewall in the multi-zone architecture. System developers utilizing cloud technologies are encouraged to consult with the TRB to ensure TRA and defense-in-depth principles are being addressed. Business Rule CI-11 enables the use of such services.

Consider leveraging FedRAMP to make it easier to embrace tested and commercially available cloud services. IaaS clouds must separate information flows logically or physically using FedRAMP-approved mechanisms and/or techniques to accomplish required separations.

BR-CI-10: (Rule Withdrawn after TRA 2016R1): Required Physical Components

BR-CI-11: Cloud Services Covered by a CMS ATO May Be Used in Lieu of Virtual Network Elements

The CMS TRA Multi-Zone Architecture permits the use of certain cloud services instead of virtual machines or applications only if those services hold a CMS ATO.

Rationale:

Some CSPs may offer services rather than virtual machines to provide critical CMS services, such as firewalls, load-balancing, or XML acceleration.

Cloud Architecture

As an extension of the CMS Common Enterprise Infrastructure environment, the cloud infrastructure must conform to the same guidelines and policies as the other CMS operating environments. Cloud architectures at CMS are built on:

  • Cloud infrastructure components
  • Cloud multi-zone architecture
  • Cloud-suitable applications

The infrastructure components serve as the architecture’s “building blocks.” The cloud multi-zone architecture outlines the “city plan” for constructing with those building blocks. Like a city, there are shared services that are defined for clouds to provide certain capabilities that are above and beyond traditional data centers. Application software can leverage the cloud architecture to provide business services that meet cloud computing expectations of scalability, elasticity, and self-service. The following subtopics address these architectural aspects.

Cloud Infrastructure Components

Some Cloud Service Providers may offer all three types of cloud service models (IaaS, PaaS, and SaaS). The delineation between IaaS and PaaS is often blurred depending on the CSP’s service offering as described in its cloud service catalog. More common are CSPs that coordinate with third-party providers or system integrators to supply and operate SaaS and/or PaaS independent of the IaaS offerings.

For purposes of discussion within this chapter and as defined by NIST, IaaS represents the core computing resources of CPU, memory, and virtual machine hypervisor instances; storage capacity; LAN throughput and Internet connectivity; and management software for managing these resources. Additional IaaS services often include backup storage for off-line / off-site storage, infrastructure and hypervisor security services, additional logging and reporting services, load balancing services, relational database and unstructured database storage services, and cloud administrator account management services to allow monitoring and managing of PaaS services. Other services such as database administration or web hosting typically fall under the PaaS category.

Role of Virtual Machines in Cloud

The core infrastructure component of the IaaS cloud is a virtual machine. A VM is a set of resources, provided by one or more physical devices, configured to provide x compute cores, y memory, z disk storage, and t network interfaces. These values for (x, y, z, and t) vary by provider. The VM can run any general-purpose operating system certified by the IaaS provider. In addition, the VM may also support special purpose components such as routers, firewalls, Graphics Processing Units (GPU), and Extensible Markup Language (XML) appliances, which use manufacturer-hardened OS, for virtual security appliance roles.

Not all flavors of OS are supported by the specific hypervisor(s) in use at a CSP. Most CSPs offer a catalog of the OS images their infrastructure supports. In addition, CMS may provide additional security-hardened images as needed, especially if virtual security appliances are necessary.

Whether a platform is IaaS or PaaS, the VM is the core technology that facilitates on-demand provisioning. The primary difference from the standpoint of the cloud user is that in IaaS, the VMs are provisioned on demand while such details are typically not transparent to PaaS users.

Types of Virtual Machines

IaaS Clouds use two kinds of VMs: application and infrastructure. This intentionally mimics the physical machine configurations typically found in data centers.

Application VMs operate application software such as:

  • Database servers (both relational and NoSQL)
  • Web servers
  • Java 2 Platform Enterprise Edition (J2EE) application servers
  • Commercial Off-the-Shelf (COTS) applications

Infrastructure VMs provide virtual equivalents of physical data center infrastructure components, including:

  • Virtual firewalls and switches
  • Virtual caching proxy servers
  • Virtual load balancers
  • Virtual network intrusion detection system devices
  • Virtual XML acceleration processors
  • Virtual machine management servers
    • Virtual Lightweight Directory Access Protocol (LDAP) servers
    • Virtual Domain Name Services (DNS) servers

The IaaS cloud configuration combines both kinds of VMs into a multi-zone architecture that meets the security and privacy requirements mandated for CMS systems.

In some situations, the IaaS cloud vendor may offer physical instances for some of the devices listed. There are valid reasons to choose physical devices over virtual devices if the choice is available. For example, using a physical firewall rather than a virtual appliance may improve performance or have a lesser exposure to security vulnerabilities. Due to the rapidly changing cloud infrastructure market, it is important to assess these options within the context of required CMS security and business requirements.

In other instances, the IaaS cloud vendor may offer these capabilities as services, without specifying virtual machines. For example, Amazon Web Services Security Groups and Elastic Load Balancers are not defined as virtual machines.

Services covered by a CMS ATO (above and beyond the FedRAMP Authorization Process) may be used in lieu of virtual machines in the CMS Multi-Zone Architecture. Satisfying the CMS ARS is a higher bar than the FedRAMP controls. In other words, it is not sufficient to satisfy FedRAMP to meet CMS’s requirements. See CMS SaaS Governance for more information.

Hosting CMS Government-Furnished Equipment at a CSP

Depending on the CSP, it may be possible to place CMS GFE into a cloud environment. Typically, the CSP would cage off and use this equipment for very specific purposes. For cost reasons, this is generally prohibited, as reinforced by Business Rule CI-4.

There are, however, several legitimate considerations for hosting GFE in a cloud environment. The following considerations are the most common:

  • Non-virtualized hardware may be more suitable in some cases than virtual hardware. This is true for applications that may leverage specific hardware capabilities or that are extremely performance intensive.
  • Some applications may not be suitable for a virtualized environment.
  • Specialized, non-general-purpose hardware may be required by CMS in the cloud environment.
  • Some data must be stored on GFE rather than third-party storage hardware.
  • Greater security may be realized with specialized appliances than general-purpose VMs.

Use of GFE constrains available cloud capabilities, such as bursting, elasticity, and managed services. By hosting GFE in the cloud, CMS loses many of the benefits of cloud computing.

Cloud Multi-Zone Architecture

The CMS TRA defines multiple zones in the architecture, including the Presentation Zone, Application Zone, Data Zone, a Transport Zone, and a Management Zone. In a traditional hosted managed services environment, these zones would be physically separated, with redundant paths and multiple (layered) security devices along the paths.

The cloud environment must be adapted to fit the CMS Multi-Zone Architecture. Since each change introduced to the standardized structure results in higher cost and management overhead for CMS, one benefit of cloud architectures is the standardized structure that employs common equipment, network interconnects, and management servers. The tradeoff is the need to provide sufficient detection and preventive security controls to meet CMS requirements.

In an IaaS / PaaS cloud environment, the devices that comprise the various zones (including servers and storage) are usually virtualized instances. The network itself is often a combination of services as well as virtualized and physical devices making extensive use of VLANs.

There are, of course, certain risks to be considered in cloud environments, including:

  • High-performance applications may experience performance degradation due to virtualization overhead.
  • Maintaining compliance with licensing contracts can be challenging due to relative ease of provisioning, elasticity, and bursting. These features may impact licensing costs, and not all software licenses are cloud compatible.
  • Not all software is equally suitable for deployment in a cloud environment.
  • For PaaS, the platform is typically upgraded for all customers simultaneously with no option to stay on back-level software.
  • The hypervisor introduces some additional security risk because it lies behind the cloud’s private infrastructure.
  • Ease of creation of virtual machines may lead to VM sprawl. Although VMs are efficient users of hardware, a VM requires the same amount of technical labor as a physical machine for services above the IaaS layer. Thus, costs can and will grow with elasticity.
  • For PaaS, the platform is typically proprietary. Consequently, applications written for this platform will not be portable to other platforms (PaaS or otherwise). There is a very real danger of vendor lock-in when dealing with PaaS.

Resource Aggregation – Resource Clusters and Resource Pools

Cloud computing involves two levels of resource aggregation: resource clusters and resource pools. A resource cluster is defined as a group of connected servers that work together so that, from the point of view of aggregating resources such as CPU processing and memory, they can be viewed as though they were a single computer. For example, a resource cluster might have 30 CPUs and 120GB of RAM. In a community cloud, a CSP builds resource clusters and assigns them to the community. The resource cluster is only available to members of the community.

A resource pool is a set of available CPU and memory resources, and generally consists of multiple “like” resources, i.e., of the same type. In a community cloud, once the resource cluster(s) are created, resource pools are defined and allocated to members of the community. Resource pools may be organized in a hierarchy in which resource pools are subdivided into subordinate resource pools. For example, a resource pool could be defined as “Administrative” and have sub-pools of “HR” and “Payroll.” At CMS, the rule is to use one resource pool per business application.

CMS Reference Architecture

Virtualized Multi-Zone Cloud Infrastructure depicts a conceptual CSP cloud implementation for a CMS reference architecture. Virtualized Multi-Zone Cloud Infrastructure illustrates several key concepts when deploying systems on a cloud infrastructure. To provide a defense-in-depth posture, the firewall and network devices managing Internet communications must be implemented as physical devices. If they were implemented on the same shared VM host, a single exploitation of the hypervisor would compromise the entire infrastructure.

The enclosing boxes in Virtualized Multi-Zone Cloud Infrastructure represent three different physical sets of resources, presented clockwise as the CMS-provided Equipment Cage (at 11 o’clock), the Production Zone Resource cluster at 1 o’clock, and the Management Zone Resource cluster, at 8 o’clock. The single physical firewall at the center of the diagram and the virtual switches are generally part of the CSP IaaS. The addition of virtual firewalls between processing zones meets CMS security requirements regardless of whether CMS manages only the firewall rules or the entire firewall instance.

For technologies that require shared responsibilities, CMS and the CSP must document the demarcations (e.g., firewall configurations managed by the CSP and firewall rules managed by CMS). Separate resource clusters must be defined for production, non-production, and Management Zones. Resource pools within a resource cluster are defined for each business owner or application.

Virtualized Multi-Zone Cloud Infrastructure (page 11)

The cloud environment offers additional architectural flexibility. Some CSPs provide a service through which customers can host agency-provided equipment at the CSP in a caged environment. CMS may choose to implement COTS or Government-Off-The-Shelf (GOTS) products, additional security monitoring equipment, or specialized equipment that handles CMS’s most sensitive data—e.g., Protected Health Information or Federal tax Information (FTI). Business owners should carefully consider costs when using the CSP’s customer hosting services. CMS or the CSP may be responsible for monitoring and managing this equipment.

Note: The diagram Virtualized Multi-Zone Cloud Infrastructure greatly simplifies the details of the Internet connection, especially the requirement for Trusted Internet Connection (TIC) as required by FedRAMP. The CMS ARS provides the details for wide area networking.

Cloud Services Architecture

A cloud can be represented functionally as the combination of many managed services. The following services are needed in a CMS Cloud. It is unlikely that any one commercial cloud offering would contain all such services; however, applications destined for the Cloud and with a DevOps focus need to ensure that these capabilities are available.

Here are cloud services organized from the perspective of the primary user:

Customer View

  • Self-Service portal
  • Cloud Service Provider Help Desk services

System Programmer View

  • CSP Application Programming Interface (to communicate with the CSP services and cloud management infrastructure)

Systems View

Virtualized machine resources

  • Virtualized storage resources
  • VM configuration management (different from Software Configuration Management)
  • Resource and Capacity Management services
  • Application Performance Management services
  • Storage Management services
  • Virtual Machine Image library
  • License-tracking services
  • Run-book automation services

Security View

  • Security services (including firewalls, audit logging, security monitoring, and other such services)
  • Remote Log-in services

Network Operations View

  • Backup and Archival services
  • Virtual Network services
    • VLAN Management services
    • Internet Protocol Address Management services
    • Dynamic Host Control Protocol (DHCP) / DNS services
    • Virtual Firewall services
    • Authenticated Time Servers
  • Application Performance Monitoring services

Cloud Service Provider View

  • Billing and Payment services
  • Hypervisor infrastructure

Data centers commonly use many of these tools today. The highly automated nature of cloud environments now makes their use mandatory.

Cloud Environment Governance and Management

As noted in NIST SP 800-146, Cloud Computing Synopsis and Recommendations:

Attempts to describe cloud computing in general terms have been problematic because cloud computing is not a single kind of system, but instead spans a spectrum of underlying technologies, configuration possibilities, service models, and deployment models.

Many factors influence the definition, benefits, risks, and implementation of a cloud infrastructure. From a pragmatic standpoint, deploying into a cloud infrastructure relies on the same underlying technologies, principles, and processes of a managed services virtualized data center. A cloud infrastructure differs from a virtualized data center by providing the following additional capabilities:

  • Self-service, allowing authorized users to provision servers and applications on demand, on behalf of business owners, using streamlined processes
  • Automation, beyond that of virtualized and physical data centers
  • Rapid elasticity, allowing applications to meet short- and long-term growth goals in a regulated but automated fashion
  • Multi-tenancy (in Community Clouds), allowing multiple cloud customers to share resources while securely separating resources for one customer from those of another
  • Metered billing, which allows CMS to lease rather than own the cloud resources and pay only for what CMS uses from the CSP

Utilizing both the proper controls, as defined by applicable CMS and federal security mandates, and the requirements for the handling and processing of CMS information and information systems, will help assure a trusted, shared environment that provides the benefits and efficiencies of the cloud environment and cloud solution providers.

 PREFERRED

The CMS strategic implementation that supports these requirements is CMS Cloud. Governance and management facilities include:

Cloud Environment Management

Controlling and managing the cloud environment requires deployment of cloud management tools.

In aggregate, these tools must provide the following capabilities:

  • Cloud virtual resource management
  • Cost management
  • Capacity management
  • License compliance management
  • Security management
  • Automation of virtual resource configuration management

The following subtopics address each of these cloud management tool capabilities.

Cloud Virtual Resource Management

Cloud resources are virtual resources temporarily allocated to the business owners on demand and by request. At some point when a business owner determines that it no longer requires these resources, the business owner may release the resources back to the CSP’s available resource pool. If the business owner fails to do this, the organization will continue to pay for resources that are no longer utilized. Even though the cloud resource capacity may seem endless, the applications need to be designed and configured to operate with a specific minimum and maximum capacity in mind. The capacity consumed should be monitored carefully to ensure that costs remain bounded within planned parameters. The CSP’s cloud resource management tools and reports are designed for these purposes. At a minimum, these tools must provide:

  • Resource provisioning
  • Resource de-provisioning
  • Resource accounting and inventory
  • Capacity management

Cloud Resource Life Cycle

There are six states in the cloud resource life cycle:

  • Provisioned virtual resources – have been created using resources either previously cleansed and made available for use or are archived resources that must be restored to operational status. Before making any resources operational, preparatory steps (such as enabling encryption) may be necessary.
  • Operational resources – have been provisioned and are running normally.
  • Elastic resources – have already been provisioned but may be bursted or made elastic (as needed) by adding or removing capacity to the resources. The last atomic unit of a resource may not be removed except during resource de-provisioning.
  • Archived resources – no longer occupy execution resources but do occupy storage. These resources may be provisioned again in the future.
  • Cleansed resources – were previously used by CMS and have been wiped based on CMS ARS requirements. Where wiping is impractical, such as with Solid-State Drive (SSD) technology, resources must have been previously encrypted and data handled in accordance with CMS ARS requirements.
  • De-provisioned resources – have been fully reclaimed by the CSP’s hypervisor systems.

Resource Provisioning

Once CMS has contracted for cloud services, actual provisioning of IaaS and PaaS services within the cloud must adhere to the CMS TRA, the CMS ARS, the CMS ISPG Cloud Computing Standard (RMH Volume III, Version 1.0), and any other federally mandated policies.

Table - Resource Provisioning Methods presents the three common methods for provisioning resources in existing clouds. Not all vendors accommodate every method. Some of the methods may be accommodated through different means and capability levels.

Table - Resource Provisioning Methods
MethodBenefitsRisks
Manually through the CSP’s Web interface
  • Improved controls and awareness of services
  • Better control over resource costs
  • Cannot provision instantaneously; short-term performance degradation may occur
  • Inability to respond effectively to bursts in activity due to manual monitoring
Programmatically through cloud API calls embedded in scripts
  • Can quickly meet demand and scale accordingly (up or down)
  • Cost efficient
  • Must be synchronized with the CSP’s capabilities
  • Increases vendor lock-in
  • Must be performed securely to prevent external exploitation of CSP resources and CMS data at CMS expense
Automatically by a cloud management system based on rules
  • Can quickly meet demand and scale accordingly (up or down)
  • Cost efficient
  • No potential conflict due to configuration management
  • The CSP is responsible for managing resources and SLAs
  • Automatic provisioning can lead to spiraling costs

Whether a CSP or a CMS application owner manages the resource provisioning, the capability to monitor and manage the resources’ performance characteristics must be made available to both parties. This helps ensure the expectations of and consistency in quality and levels of service.

It is essential to carefully control the use of APIs for resource provisioning. Some form of logging, auditing, and notification should support the automated provisioning of resources. Uncontrolled use of automated provisioning routines could have devastating consequences in service performance and cost escalation. Thus, rigorous control is mandatory for the capability to perform automated provisioning and de-provisioning.

CSPs offer tools to gauge performance of their service offerings. Additional tools may be considered to help CMS measure performance characteristics of a CSP’s full services. Aggregation and analysis of CSP data along with CMS data is a likely use case.

System owners must clearly understand the CSP provisioning policies and the physical location of the resources. CSPs have different allocation policies and practices for normal provisioning and elasticity capabilities. Some CSPs may support elasticity within the confines of a specific data center, and only migrate to alternative sites based on complete outages; others may spread provisioned resources over several data centers as a matter of course.

Resource De-Provisioning

Disposing of technical resources after retiring a business capability is not always a clear and simple process. It is even less so in a cloud, where previously allocated resources return to a pool to be reused minutes or even seconds later. Identifying all dependent resources can be particularly difficult in a cloud environment. In private data centers, these types of remnants do not typically have the significant impact on resources or expenses that they do in a private or community cloud environment, where services are billed based on metered consumption and resources are shared among the cloud’s tenants.

CSPs supply resource management tools to manage a business owner’s resource consumption. Depending on the CSP’s infrastructure capabilities, pay-as-you-go services can be very dynamic in nature and difficult to track. In cloud environments, where resources may be manually provisioned for temporary use through bursting or elasticity, the business owners must remember to de-provision the resources when they are no longer required. A follow-up audit may be necessary because CMS may not be aware that copies of data, via mirroring and similar redundancy services, may reside on different logical units of the CSP’s infrastructure.

Although de-provisioning can be performed in either manual or automated solutions, automated de-provisioning is generally recommended. Manual de-provisioning is not properly aligned with the security architecture principle of decreased system footprint. Manual de-provisioning also increases risk of unnecessary costs and is likely to be more error prone than automated. To operate successfully, both automated and manual de-provisioning require reliable system inventory and architectural information.

License Management

Provisioning and de-provisioning virtual machines in an IaaS cloud may generate software license issues. The business application must have sufficient licenses to cover its fully elastic configuration (unless previously negotiated). If bursting was enabled, there may be implications to licensed software that is licensed per CPU or core.

Although open-source software at CMS may be freely employed, there can be associated licensing fees for support, which may limit the number of supported instances of that software.

Automated Virtualized Environments and APIs

CSPs operate an advanced, virtualized environment, using tools to manage multi-tenancy, self-service, and resource consumption tracking that provide expected cloud benefits. Due to the economics of operating a cloud data center, CSPs rely heavily on automation to reduce their labor costs and permit rapid bursting and elasticity.

The need for highly automated environments has driven CSPs to provide APIs to their environments. Essentially, a cloud API allows a cloud consumer to control cloud resources via a programmatic interface instead of using a manual, cloud consumer portal. For example, APIs typically allow for such activities as starting VMs and expanding capacity.

The cloud industry has not yet standardized on APIs for managing cloud environments. Instead, other providers are supporting de facto standards such as the Amazon Elastic Compute (EC2) APIs in both commercial and open-source products (such as Eucalyptus). Currently, it is too early to standardize on a vendor API because of the immature state of the cloud industry. As a result, programming to a cloud vendor’s API involves vendor lock-in. Architectural design techniques may mitigate this risk to some degree, and business owners should consider such mitigations when deploying business applications to the cloud. For example, Apache libcloud provides API wrappers for proprietary cloud APIs, allowing consumers to program to a compatibility layer instead of vendor-specific APIs. Although this approach may be preferable, using API wrappers requires thorough technical analysis of the business application’s requirements.

Cost Management Implications

Given that dynamic resource management is a key element of the cloud value proposition, cost management and its architectural impact is a critical element of cloud resource management.

The cloud’s cost management component must be able to provide detailed resource utilization billing, allowing CMS data center operations to identify the costs associated with CMS business applications. In an IaaS cloud, this means accounting for CPU, memory, network bandwidth, and other agreed-upon billable resources. In a PaaS cloud, the billable resources may be number of user accounts, number of customer accounts, or other metered resources.

Pay-as-you-go billing may have an architectural impact to custom business applications. For example, it is possible that certain resource tradeoffs (such as memory vs. disk I/O) will lower the overall costs. CMS application developers and contractors should provide alternative architectures, when appropriate, to reduce overall costs.

Elasticity, while an attractive concept, has financial impacts that have not been fully explored at CMS. CMS business owners will need to include the cost of resource elasticity in their estimated operational cost for a business application. Cost models should be capable of translating pay-as-you-go billing to existing billing models until pay-as-you-go becomes acceptable in the CMS financial environment. For example, it may be feasible to have a model that uses prior-year average (or even moving average) billing to determine future costs once applications achieve a resource utilization steady state.

Resource elasticity does not normally include the licensing cost associated with COTS software. Provisioning a VM does not mean that the OS and application software are properly licensed. It is CMS’s responsibility to ensure compliance with licenses. Although some CSPs offer license management, it is typically not the CSP’s responsibility to do so.

Capacity Management

The management tools must provide authorized users a snapshot or “dashboard” view, at any given time, of the capacity management statistics, such as percentage of resources utilized, available capacity, growth trends, etc., for a range of duration (day, week, month, year). These tools should have threshold settings that take some automated action when limits are exceeded. These actions could range from simple notifications (alarms) to automatically provisioning more VM resources or VMs (at predetermined allocation increments). Many of these features depend on the CSP and its offered tool set. Different CSPs offer a variety of tool sets based on the cloud framework and vendor-specific technology in use. In addition to thresholds, “caps” should be considered. Caps are soft thresholds put in place on various cloud resources (such as CPUs, CPU cycles, Terabytes of storage, memory, and bandwidth) that may not be exceeded without a certain level of CMS approval. This approach would prevent a spiraling of costs during an unusual surge situation, much like a circuit breaker on an electrical circuit.

Current CMS virtualization business rules prohibit over-subscription of virtual machines and resources in CMS Production Environments. This means that CSPs may not overburden physical machines under the assumption that workload overlap is unlikely. Each production VM must have enough capacity to run at 100 percent load (peak capacity) without impacting another production VM.

Cloud Service Management Support

Cloud operations and support should be the same as the ongoing operational support of any other system / technology in the CMS enterprise. In fact, cloud support should leverage existing support infrastructure and managed services, such as operations staff, help desks, and service management processes. The key distinction between data center and cloud service management support is the varying degrees of managed services and possibly different service providers / contractors for the IaaS and PaaS. This also means there is potentially a large set of stakeholders involved in communications. The larger communications component necessitates more coordination across processes and reporting to help CMS make informed decisions and provide necessary oversight to the CSP and associated contractor(s) regarding their contracted performance.

Service Level Management

All ongoing operations and support should follow a formal process framework, be adequately documented, and be part of a continuous process improvement regimen. SLAs with the CSP should be no different from SLAs with any other CMS data center hosting vendor that provides integrated support and services to CMS. With some cloud providers, IaaS SLAs may be the CSP’s responsibility while PaaS SLAs are the responsibility of an associated integration contractor or broker.

This separation of concerns requires better CMS oversight to assure appropriate management of service levels. The details of service level management are outside of the scope of this chapter, but are addressed by CMS Cloud: Defining Service Level Objectives (SLOs).

IT Processes and Best Practices

Using a CSP makes the integration of standard IT processes and procedures more complex. Problem management, incident management, change management, and configuration management remain critical processes, but often their responsibility is shared between the CSP and associated integration and application contractors. IaaS changes, patches, and configuration adjustments are the CSP’s responsibility. PaaS changes and configuration settings are typically the responsibility of the integration or application contractor. CMS is responsible for ensuring smooth identification, reporting, and resolution of issues.

Cloud Security Requirements

This chapter incorporates the best practices discussed in the Cloud Security Alliance (CSA) Security Guidance for Critical Areas of Focus in Cloud Computing, Version 5.0, and the security control areas defined in the CSA Cloud Controls Matrix (CCM) for managing security risks associated with cloud computing. The security focus of this chapter is on Private and Private Community clouds, which CMS expects to host operational environments that have Low or Moderate system categorizations.

The CSA CCM contains a comprehensive list of control areas for consideration, mapped to the relevant NIST SP 800-53 controls. CMS has already provided guidance for most control areas in the CMS ARS. This topic sets forth CMS standards and guidance for implementing a subset of CCM controls areas that have special relevance in cloud environments. These areas are crucial to the security assessment of CMS cloud systems.

Note: Only the Federal Information Security Modernization Act (FISMA)-designated Authorizing Authority (the CMS CIO) can accept risks above and beyond the minimum federal requirements required under FISMA in NIST SP 800-53. For the most part, physical devices identified in the CMS TRA are replaced by logical (or virtual) devices when implemented in the cloud. Replacing physical devices with logical equivalents must be accompanied with compensating controls to avoid compromising security.

Cloud computing presents additional risks over traditional IT environments because of the virtualization of computing resources that must be properly managed to ensure the confidentiality, integrity, and availability (CIA) of CMS data. CMS requires a comprehensive approach that pays appropriate attention to the specific security challenges inherent in managing risk in cloud computing environments to ensure the CIA of CMS information and information systems.

 PREFERRED

The CMS strategic implementation that supports these requirements is CMS Hybrid Cloud. CMS Cloud maintains and secures its environments, leaving the application with primary responsibility for its Authority To Operate (ATO). CMS Cloud provides security and compliance capabilities that include:

Cloud Security Objectives

Regardless of environment, protecting sensitive data requires implementing and enforcing security controls to reduce the risk to the CIA of that data. Because of the technology for implementing clouds, the use of shared physical equipment, lack of visibility into the security management of the system, and reliance on the maturity of vendor processes and procedures, the risk in implementing a system in the vendor-operated federal cloud is likely higher than in an agency’s dedicated computer center.

Similar to traditional computing environments, cloud computing implementations are subject to local physical threats as well as external threats. These threat sources include accidents, natural disasters, external loss of service, hostile governments, criminal organizations, and terrorist groups. Additional threat sources are intentional or unintentional introduction of vulnerabilities through internal or external authorized or unauthorized human and system access, including but not limited to, employees, contractors, and intruders. The characteristics of cloud computing, including multi-tenancy and the variety of service and deployment models, underscore the need to consider data and systems protection within the context of logical as well as physical boundaries.

Cloud computing implementations include the following major security objectives:

  • Preventing unauthorized access to cloud computing infrastructure resources. This includes implementing security domains that have logical separation between computing resources, e.g., logical separation of customer workloads running on the same server by VM monitors (hypervisors) in a multi-tenant environment and using secure-by-default configurations.
  • Managing hypervisor threat vectors. Hypervisor attack threat vectors include CSP users and CSP employees. Hypervisor vulnerabilities include poor configurations, missed or delayed security patching, or unauthorized activities from a privileged user. Because hypervisor exploitation can have disastrous results, CMS mandates that all hypervisor solutions must be Type 1 (native / bare metal)-based products. Industry best practices should be implemented where available or the vendor’s own published guidelines should be used.
  • Minimizing shared network access. Most CSPs have some common network infrastructure components between various cloud customers; these shared network infrastructure components present significant risk because a single breach of a shared component could compromise all users of a CSP’s service. Thus, configurations must adhere to best practices, and exceptions must be well understood, documented, and accepted by users of a CSP’s service.
  • Managing privileged user access. CSP privileged users with access to the hypervisor should be kept at a minimum. CSPs should ensure the timely removal of access based on a CSP employee termination event or when access is no longer required (e.g., a job transfer). CSP targets for managing privileged user access in both cases should be real-time removal of access, audit records should be available to CMS auditors that establish when a request was made, and the actual removal of privilege.
  • Ensuring that appropriate security safeguards are deployed at the CSP. CMS should conduct independent assessments to verify that appropriate safeguards are in place. This includes traditional perimeter security measures in combination with the additional safeguards required for cloud computing.
  • Defining trust boundaries between CSPs and the CMS consumers. It is crucial to clearly document the responsibility for providing security and that the consequences for non-deployment of agreed-upon security controls by the CSP are well-defined in contracts and SLAs.

CMS Cloud Roles and Responsibilities

CMS has established its approach for implementing the guidance in NIST SP 800-37, Risk Management Framework to Federal Information Systems and Organizations. This topic describes how CMS applies this framework in the life cycle of cloud-based systems.

Operating in the cloud may introduce ambiguity about who is responsible for specific activities. Since ambiguity may lead to risk, CMS defines accountability for managing its CSP provider(s). The CMS Cloud Manager is responsible for managing the CSP. Only a CIO-designed cloud management organization may procure a cloud.

The table Table - Roles and Responsibilities in the Cloud (NIST Risk Management Framework) summarizes the cloud-specific guidance associated with managing a CSP.

Table - Roles and Responsibilities in the Cloud (NIST Risk Management Framework)
ActivitiesCloud-Specific Roles / Guidance
Conduct Federal Information Processing Standards (FIPS) 199 security categorizationCMS CM, CMS business owner, and CMS ISPG.
Identify regulatory compliance requirementsCMS Chief Information Security Officer communicates all compliance requirements to the CMS CM. The CMS ARS and FedRAMP define all of the security requirements.
Analyze cloud deployment model options (Public, Private, Community, Hybrid) based on security categorization, HHS Cloud Computing, and CISO Cloud Standards guidance

CMS CM works with CMS business owners, ISPG, and other appropriate CMS organizations to define CSP requirements. CMS CM defines and acquires the appropriate cloud deployment model.

CMS CM, Chief Information Security Officer (CISO), and CSP evaluate cloud architecture, and the CMS ARS controls selection.

Design, test, and integrate

In coordination with the CMS CM and ISPG, the CSP implements, tests, and documents the controls, including artifacts of compliance to regulatory requirements.

Connectivity between a CSP and other CMS data centers will require a secure network solution and all required technical documentation. If both interconnecting systems have the same Authorizing Official (AO) or same primary operational IT infrastructure manager, an Interconnection Security Agreement is not required.

Conduct independent security test and evaluationCSP hires a FedRAMP-certified third-party organization (3PAO) to conduct assessment of FedRAMP security controls related to the deployment model and additional control requirements mandated by CMS (such as, but not limited to, Health Insurance Portability and Accountability Act (HIPAA), Health Information Technology for Economic and Clinical Health (HITECH), etc.); CMS CM and ISPG perform an independent validation.
Government stakeholders may conduct audits—e.g., audits for Federal Tax Information, HIPAA, or the Office of the Inspector General (OIG)CMS CM, along with ISPG, works with the CSP to determine how to support CMS and external assessments and audits.
Identify corrective actions for risks

CMS ISPG and the applicable FedRAMP Authorizing Organization records and manages the Plan of Action Milestones (POA&M)

CSP supports actions required by the POA&M.

Manage authorization

Business owner submits Authority to Operate package.

CMS ISPG makes risk-based ATO recommendation. CMS CIO makes risk tolerance-based ATO decision.

CMS may leverage FedRAMP and HHS cloud authorizations where applicable.

Perform security continuous monitoring activities, including asset management, configuration management, vulnerability / patch management, and intrusion detectionCSPs must align their continuous monitoring to support CMS’s reporting and oversight requirements.

The FedRAMP and CMS ISPG govern all CMS policy and processes necessary for obtaining an Authorization to Operate (ATO) in or as a cloud environment at CMS. The CMS Information Systems Security & Privacy Policy (IS2P2) and other CMS Policies and Guidance provide direction on achieving an ATO in the CMS enterprise.

In managing risks within the cloud environment, the security of a platform (PaaS) depends on the controls’ effectiveness within the underlying infrastructure (IaaS); similarly, the security of an application depends on the effectiveness of the underlying platform and infrastructure. For this reason, the approval of the PaaS is contingent on whether the PaaS resides in an approved infrastructure, and the approval of a cloud-based application is contingent on whether the application resides in an approved PaaS and IaaS. In all cases, the PaaS security controls are the same whether the infrastructure supporting the PaaS resides in a traditional data center or is provided by a CSP.

FISMA reporting requirements apply to all systems that store or process government data. CMS will manage these requirements for the virtual instances where their services will run. CSPs will be required to provide Security Content Automation Protocol (SCAP)-compliant FISMA reporting information for the underlying cloud infrastructure and support systems, including hypervisors and all physical hosts that make up the cloud resources. These reports will include CMS-related cloud devices only and be delivered in the format and schedule dictated by FISMA reporting requirements.

Encryption and Key Management

Many CSPs offer encryption services to their customers. As CMS business owners require this service, the following requirements must be met:

  • Where the CSP owns the encryption key(s), the encryption key must be different from CSP keys used to manage other CSP customers.
  • The CSP CMS keys must be escrowed with CMS.
  • CSPs should be able to demonstrate to CMS how the keys are managed and protected.
  • The CSP must ensure that the assignment of a key manager adheres to segregation of duties (per the CMS ARS) and provides an option for CMS to perform its own key management.

Resource Traceability and Records Management

In a traditional physical data center deployment, system managers know precisely where data resides by physical server and storage device. In a cloud, data and applications will be allocated across several hardware units, some of which may be geographically disparate.

A CSP may move data securely between its various data center environments providing the CSP complies with CMS’s contractual agreement and their compliance has been documented a priori during the security authorization process. In addition, the CSP must have logging enabled to record where data resources are currently located and where they have been deployed. Depending on the data’s sensitivity (e.g., PII, PHI, and FTI) and records management requirements, CMS data stewards may have to account for where the data currently resides and where it has been to demonstrate the proper application of data life-cycle management.

In addition to the requirements placed on CSPs, business applications have federal records management requirements for data life cycles as well as data retention and records destruction.

Data Availability

Data management policy may require alternate solutions for managing data and keeping sensitive data outside of the cloud environment. Data management may require traditional hosting solutions with controls in place for following the chain of data authority and integrity. A review and understanding of the controls provided in the cloud environment for managing data must be clear and well documented.

Moving data into or out of the cloud requires careful planning and execution. Although data moving into or out of a cloud can use conventional methods for small data repositories, such as through files or programmatically, large amounts of data may require physical media that would need to be passed through the CSP to load into the cloud or be removed from the cloud. CSPs offer this type of service, which requires advanced coordination and planning.

Availability of Data

Archived data must meet National Archives and Records Administration (NARA) requirements and be stored in a format capable of meeting e-discovery needs. (Please refer to CMS Cloud Computing Standards for details.)

CSPs must meet the data needs for security forensic analysis by following CMS ARS Audit and Accountability (AU) requirements.

Data Protection

CMS must require the CSP to safeguard any confidential information (such as PII, PHI, and FTI) in compliance with Agency standards and to comply with all applicable regulatory reporting requirements.

In a multi-tenant environment, confidential data must be protected using a combination of access control, contractual liability, and encryption.

Solutions are needed for protecting data under the following circumstances. Depending on the type of cloud service, the solution can be supplied either by the CSP or by CMS:

  • Backup to tape or disk (IaaS or PaaS)
  • Archived data (PaaS)
  • Removal of disk for repair (IaaS)
  • Protection of data between and in disaster recovery sites (IaaS)
  • Data extracts sent to partners (PaaS, SaaS, or CMS)
  • Shared / consolidated storage used by multiple organizations (IaaS or PaaS)
  • Protection from insider theft (service provider employees) (IaaS, PaaS, or SaaS)

Detective Controls for Mitigating Cloud-Specific Risks

The architecture of cloud computing environments presents challenges that do not exist in traditional computing environments. Cloud implementations may simulate the standard multi-tier architecture; however, this is only a simulation of a defense-in-depth strategy implemented at the infrastructure and hypervisor levels. The ability to maintain the security profile of this zone simulation requires both good configuration management, which reduces the likelihood of misconfigurations that would expose the system and data to risk, and appropriate responses to exploits that would compromise the pieces of the architecture that enforce access control and separation between tiers.

The major threats to the confidentiality, integrity, and availability of CMS systems and data are:

  • Misconfigurations
    • Accidental misconfigurations
    • Delayed security patching
    • Insider threats
  • Process and procedural immaturity
    • Operational procedures
  • Management oversight

Several detection techniques can be implemented to reduce the risk of these threats to better realize the potential benefits of using the cloud. The ensuing subtopics address proposed requirements for detective techniques to mitigate the risk of operating a multi-tier system in a cloud environment.

The multi-tier architecture is implemented as multiple VLANs through a network interface on the cloud infrastructure; zone separation is provided by either physical firewalls or by virtual firewalls.

A cloud implementation of a multi-tier architecture is dependent on the infrastructure components and the hypervisor to enforce the architecture. Misconfiguration of shared CSP devices (i.e., devices shared with non-CMS customers and/or devices not under the direct control of CMS) or the hypervisor makes the CMS systems vulnerable. Thus, detection techniques focus on ways to quickly identify and remediate misconfiguration of these key components, whether the result of unintentional errors or malicious activity.

The following technical measures are intended to detect component misconfiguration.

Firewall Modification Detection

The following requirements are for firewalls not under CMS’s direct control:

  • Any changes to the firewall (rules, configuration, etc.) must automatically generate an alert.
  • Real-time security-relevant events on the firewall must be audited.
  • Audit logs for indications of inappropriate or unusual activity must be reviewed in a timely manner.
  • Unauthorized firewall modifications must be reported to the CMS security team within 60 minutes (as required by OMB Memorandum M-25-04) if these modifications have not been identified as a false-positive result.

Hypervisor Modification Detection

The following requirements are for hypervisor oversight:

  • Any changes to the networking configuration of the hypervisor must automatically generate an alert.
  • Real-time security-relevant events on the hypervisor must be audited.
  • Hypervisor audit logs must be audited weekly for indications of inappropriate or unusual activity.
  • Unauthorized modifications to the hypervisor must be reported to the CMS security team within 60 minutes (as required by OMB Memorandum M-25-04) if they have not been identified as a false-positive result.

Network Infrastructure Modification Detection

The following requirements are for network devices not under the direct control of CMS:

  • Security-relevant events of the underlying infrastructure and network IDS must be audited in real time for indications of inappropriate or unusual activity
  • Audit logs of network infrastructure devices, such as network IDS monitors, alert logs, etc., must be reviewed weekly for indications of inappropriate or unusual activity.
  • Unauthorized modifications to network infrastructure must be reported to the CMS security team within 60 minutes (as required by OMB Memorandum M-25-04) if they have not been identified as a false-positive result.

Administrative Network Port Scans

The following requirements are for network devices not under the direct control of CMS:

  • Cloud management infrastructure devices must be scanned daily to identify any undocumented or unauthorized ports.
  • Unauthorized ports and protocols must be reported to the CMS security team within 60 minutes (as required by OMB Memorandum M-25-04) if they have not been identified as false-positive results.

Management Detective Controls

Management oversight by the CSPs provides the government with insight into how its cloud is managed and confidence that it is managed securely. The detection management controls permit the government to identify patterns or weaknesses that could compromise the security of its cloud.

Weekly Verification Report

On a weekly basis, summary reports must be provided to CMS to document infrastructure and hypervisor modifications, to confirm that change requests have been verified, and to report any unaccounted for or unapproved modifications with all corresponding corrective actions taken.

Definitions

This chapter relies on the NIST SP 800-145, The NIST Definition of Cloud Computing, September 2011, as the authoritative source for cloud terminology. Visual Model of NIST Working Definition of Cloud Computing shows a visual model of the NIST Cloud model.

Visual Model of NIST Working Definition of Cloud Computing (page 12)

Table - NIST Definitions identifies the key terms and NIST definitions for Cloud computing.

Table - NIST Definitions
TermDefinition
Private CloudThe cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party and may exist on premises or off premises.
Community CloudThe cloud infrastructure is shared by several organizations and supports a specific community that has shared concerns (e.g., mission, security requirements, policy, and compliance considerations). It may be managed by the organizations or a third party and may exist on premises or off premises.
Public CloudThe cloud infrastructure is made available to the general public or a large industry group and is owned by an organization selling cloud services.
Hybrid CloudThe cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technology that enables data and application portability (e.g., for load balancing between clouds).
Infrastructure as a Service (IaaS)The capability provided to the consumer is to provision processing, storage, networks, and other fundamental computing resources where the consumer is able to deploy and run arbitrary software, which can include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure but has control over operating systems, storage, deployed applications, and possibly limited control of select networking components (e.g., host firewalls).
Platform as a Service (PaaS)The capability provided to the consumer is to deploy onto the cloud infrastructure consumer-created or acquired applications created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including network, servers, operating systems, or storage, but has control over the deployed applications and possibly application hosting environment configurations.
Software as a Service (SaaS)The capability provided to the consumer is to use the provider’s applications running on a cloud infrastructure. The applications are accessible from various client devices through a thin client interface such as a Web browser (e.g., Web-based email). The consumer does not manage or control the underlying cloud infrastructure, including network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

Other terms commonly used when addressing cloud and virtualized environments include:

  • Virtual Machine – An instance of an operating system variant (e.g., Linux) running under the control of a hypervisor.
  • Bursting – The ability to increase VM resources, such as processor, memory, or storage capability, provided as a service by the CSP.
  • Elasticity – The ability to add or subtract VMs for use by a business application, implemented through configuration of the application platform (PaaS).

IT Performance Management

Introduction to IT Performance Management

Background

IT Performance Management (IT PM) is the practice of using performance monitoring to collect, identify, assess, forecast, and correct performance problems. It is a key function of IT Service Management.

CMS data centers perform monitoring within the CMS Processing Environments for various purposes, including, but not limited to:

  • Security monitoring (intrusion detection system, firewalls, logs, incidents, etc.)
  • IT Compliance monitoring (configuration management, policies, etc.)
  • Resource monitoring (utilization, capacity, aging, etc.)
  • Business Activity Monitoring (BAM) (monitoring of business process performance)
  • Infrastructure Performance Monitoring (IPM) (availability, outages, throughput, latency, etc.)
  • Application Performance Monitoring (APM) (transaction response time, throughput, integrity, etc.)

All of these purposes inform the monitoring framework for inspecting and controlling IT systems at CMS. Because the same tools and procedures often address more than one purpose concurrently, there is potential for significant overlap in monitoring for each of these purposes.

Purpose

The guidance in this chapter introduces performance monitoring concepts, differentiates performance monitoring from performance testing, and identifies key business rules for the CMS Processing Environments.

IT PM should be applied to all performance-critical business or infrastructure applications. This chapter provides guidance to application owners, developers, and maintainers in the following areas:

  • Deciding which application behaviors to instrument based on business objectives
  • Determining how to instrument application behaviors
  • Defining application performance metrics and standards

Scope

This chapter represents the IT PM Architecture standards that should be used by CMS and CMS / Contractor partners for CMS Processing Environments. Application Performance Monitoring is a key focus of this chapter. Monitoring requirements for other purposes, specifically security and operations monitoring, are covered in CMS ARS and in other CMS TRA chapters.

IT PM applies best to IaaS Clouds and traditional virtualized data centers. IT PM can be performed on PaaS Clouds and even SaaS Clouds, but there will be challenges due to the lack of access to the hosting platform.

Business Drivers

Business operations and business goals depend on the performance of applications. An application’s performance, as experienced by business end users, is the result of capabilities provided by multiple IT components and services from contractors managed by CMS.

Goals

The business goals for using IT PM are to continuously:

  • Improve uptime, responsiveness, accuracy, quality, security, and timeliness of the customer experience.
  • Process CMS business workloads efficiently and effectively.
  • Enhance and improve the use of IT resources to control and reduce costs over time.

IT PM supports the business by providing insight into the operation of IT systems. The information gathered through IT PM allows CMS to:

  • Translate business goals into IT operational level agreements (OLA). Performance standards in OLAs must be grounded in the requirements of business operations and business goals.
  • Support government oversight of SLAs with applicable IT PM data and reports when the collection of the IT PM data is practical and unobtrusive.
  • Use IT PM data in planning and allocation of application resources to avoid over-provisioning when acquiring or allocating hardware and network resources.

Common Capabilities

IT PM includes tools and methods that measure the end-to-end performance of an application and each of its components. End-to-end performance measurements inform IT and business management whether end users are currently or are likely to be experiencing performance issues with an application. IT PM tools and methods provide capabilities that can elevate the quality of IT services in several key ways. For example, IT PM:

  • Provides operators with real-time status of an application’s monitored performance metrics
  • Presents information by application, not just by data center components
  • Measures and reports on service levels from an end-user perspective
  • Enables operators to discern the service- and business-level impact of events
  • Gives operators the cross-platform visibility and impact analysis capabilities needed to prioritize responses and improve service availability
  • Allows correlation of multiple performance events and their user impact to business processes
  • Allows detection of developing application and service performance problems before they result in wide impact to end users
  • Enables forecasting near-term behavior that might have an adverse impact on performance

Concepts and Terminology

Layered Architecture

Two layers of monitoring exist: infrastructure / systems and service / application monitoring. These layers represent the dichotomy between technology and business.

The lowest level is infrastructure and systems monitoring, which seeks to quantify the performance of network elements such as load balancers, switches, routers, firewalls, and other such devices. Metrics tend to be very technical and translate to the technical capabilities of the networking infrastructure. In clouds, this kind of data may be difficult to obtain.

The next highest level is resource monitoring, where computing and storage elements are quantified for their respective performance.

The highest level is service or application-level monitoring, where the performance of the application or service is quantified. At this level, metrics are business focused and reflect the link between technology and business function. The challenge is in defining the services and metrics because these will not be expressed in universal terms as is true with technical metrics. The benefit is that the business can reason about the performance in business terms, which is how customers think of services.

Supporting Tools

IT PM can be defined as the process of monitoring and reporting on the availability and performance of a service, including the detection and diagnosis of performance issues. IT PM tools include:

  • Agents and other mechanisms, which collect or generate metric data relating to application performance and status
  • IT PM servers, which store, compile, and communicate the performance metric data
  • Analysis tools, which analyze and report on the metric data
  • A dashboard portal, which provides data center operators with status information

IT PM analysis tools track and monitor response time or other performance metrics. The analysis tools can be configured by IT PM administrators using Event Situation Rules to detect and respond automatically to Performance Events as they occur based on metric data.

Configuration

Configuring the IT PM analysis tools involves settings that define metrics, condition thresholds, events, rules, and actions. Configuration relies on:

  • Metrics refer to data reported to the IT PM analysis tool concerning specific measures of an application’s performance, such as response time.
  • Thresholds are defined for each metric. A threshold may be a stated value or range of values for a given metric that, if exceeded, indicate a certain condition is true.
  • A Condition is true when a monitored metric data value exceeds a given threshold.
  • A Performance Event occurs when some predefined combination of conditions is true. Each type of Performance Event has its own criteria.
  • A set of Event Situation Rules consist of criteria for a given type of Performance Event and the actions to take when the Performance Event occurs.
  • An Alert is triggered when the Event Situation rules are met.

When the predefined conditions of a rule are met, a Performance Event is logged and the IT PM analysis tools automatically perform the rule’s predefined set of actions such as displaying an alert, sending a notification, or generating a ticket. The metrics, thresholds, and rules used to monitor a CMS application are defined through collaboration between the business application owner, Office of Information Technology (OIT), and the IT PM team.

Other Common Terms

Business Service Metrics

Measures of the availability or performance of a business service as provided by an application.

Event Correlation

Process of correlating multiple events using decision logic to determine a measurement of a multiple-event process or to detect a pattern indicating a performance issue.

Event Management

Activity of monitoring and taking action on events.

Event Enrichment

Information and analysis added to the description of an event defined by a set of Event Situation Rules. An example of Event Enrichment might be a description of the business impact of that event, or a note that the Event may be affected by cache size.

Performance Testing vs. Performance Monitoring

Performance and Stress Testing occurs in an Implementation environment and seeks to:

  • Verify that an application system’s performance meets system baseline performance requirements
  • Baseline the application’s performance and resource use under different workloads to establish initial norms
  • Stress an application system to identify failure points and inform design decisions before service implementation

In the CMS Processing Environment, Performance Monitoring seeks to:

  • Verify that a service is performing within prescribed control limits
  • Detect when a service is not performing within pre-defined standards so that corrective steps can be taken
  • Forecast when a service may exceed prescribed control limits (or thresholds) so that corrective action can begin

Although these activities use the same metrics and sometimes the same tools, their implementation differs, as shown in Table - Implementation Differences between Performance Testing and Performance Monitoring.

Table - Implementation Differences between Performance Testing and Performance Monitoring
Performance and Stress TestingPerformance Monitoring
Takes place in the lower environments (development, testing, integration), not productionTakes place in any of the CMS Processing Environments
Tests business and system functions that may affect performanceMonitors only critical functions
Uses simulationUses sampling
May introduce a heavy load to stress the systemApplies a minimal load to avoid impact on system
May simulate many usersMonitors from multiple locations, each simulating only one user
Test data, often designed to cause problemsLive data, reflecting actual use in that environment

Leveraging Performance and Stress Testing

By providing common guidelines for Performance Monitoring across all applications and leveraging performance and stress testing in the Implementation environment, CMS seeks to benefit from:

  • Reuse of test scripts
  • Cross-training on common performance tools
  • Use of a single, end-to-end process with smoother handoffs
  • Sharing of best practices
  • Accurate baseline testing scripts for more accurate results
  • Greater efficiency / reduced cost

Additional Testing vs. Monitoring Considerations

It is to CMS’s advantage to leverage Performance & Stress Testing within its approach to Performance Monitoring. To accomplish this, there must be consistency in performance metrics, performance requirements, and naming conventions in test scripts and transaction properties.

Tests developed for Performance Testing may not be directly reusable in production performance monitoring because of additional constraints such as these:

  • Capability differences may exist between testing tools in the Implementation environment and monitoring tools in the CMS Production Environments at different CMS data centers.
  • The presence of production data for Performance Monitoring may constrain the types of transactions that may be used or the metric data that may be collected.
  • Logging too much data or too often in the Production Environment could present unintended performance consequences and storage issues.
  • Access to performance logs in the Production Environment is controlled.

Business Rules for Application Performance Monitoring

To guide IT PM within the CMS Production Environments, CMS developed the following business rules for performance monitoring and measurement (PMM).

BR-PMM-1: Performance Monitoring Data Is FOUO

IT PM data, generated from CMS data and the performance of CMS IT resources, will be treated as For Official Use Only (FOUO).

Related CMS ARS Security Controls include: SI-4 - System-Generated Alerts.

Rationale:

IT PM data contains information about the performance of government applications, which should be used only for ensuring and improving the performance of CMS processing systems. This data is also used to update the CPIC Annual Operational Analysis Report. IT PM data must be kept long enough to inform such reports as well as meet federal and CMS data retention and records management guidelines.

BR-PMM-2: Control Access to Performance Monitoring Data

Access to the performance monitoring data must be controlled.

Related CMS ARS Security Controls include: SI-4 - System-Generated Alerts.

Rationale:

Performance monitoring data is FOUO and may contain CMS data.

BR-PMM-3: Monitor Production Environments

All applications in Production must be monitored for performance. This includes infrastructure and application performance monitoring.

In addition to mandated security logging, applications must also log performance and troubleshooting data, such as transaction paths taken and timing.

Rationale:

Production applications must be monitored for performance to ensure that CMS services are provided with known quality.

INFRASTRUCTURE SERVICES

IT Performance Management

Best Practices and Recommendations

RP-PMM-1: Coordinate Application Changes with Monitoring Operations

Significant application changes must be coordinated with the existing IT PM infrastructure to avoid negative impact to IT PM services or to the monitored application.

Rationale:

Changes to an application may invalidate some of the rules, transactions, or resources used for IT PM. The result may degrade the ability to monitor application performance or possibly the performance of the application itself or a related application. Coordinate changes, typically via the project or data center change control board (CCB), can help mitigate the risk of negatively impacting IT PM.

RP-PMM-2: Consider Monitoring Lower Environments

Monitoring applications in the Development, Test, and Implementation environments is a business decision. Applications in Production must be monitored for application performance.

Rationale:

The additional cost of IT PM may be considered unjustified by business owners who may chose not to monitor lower environments. Instrumenting and monitoring applications in lower environments may help identify problem code prior to promoting code to production. In addition, it makes lower environments execute similarly to production, which typically helps smooth migration.

RP-PMM-3: Provide Performance Data to CMS NOC

If requested, CMS data centers must provide performance event data to the CMS Network Operations Center where it can be monitored.

Rationale:

By having an integrated, enterprise-wide view of performance, CMS can better control and manage resource usage. This can be accomplished using either open standards-based products or the CMS-sanctioned IT PM gateways.

RP-PMM-4: Conduct Performance Management Planning

All CMS application maintainers must provide and maintain a Performance Management Plan approved by the data center operations contractor(s).

Rationale:

A Performance Management Plan identifies specific application monitoring requirements, business service metrics, key performance indicators (KPI), and thresholds for alerts and SLA violations. This plan also identifies contacts to be notified of performance degradation issues.

RP-PMM-5: Use a Trouble Ticketing System to Track Performance Problems

Use the CMS enterprise trouble ticket system to track resolution of incidence of performance issues in production systems.

Related CMS ARS Security Controls include: IR-4 - Incident Handling, IR-5 - Incident Monitoring, and IR-6 - Incident Reporting.

Rationale:

An enterprise trouble ticketing system allows for trouble tickets to be tracked, prioritized, and dispositioned according to business and technical priorities.

RP-PMM-6: At Least One End-to-End Test

Before going into production, applications hosted in CMS data centers must have at least one test to verify end-to-end operational status of the performance monitoring system as part of Production Readiness testing from within the Production environment.

Rationale:

This rule ensures that applications processing CMS data have implemented and validated the performance monitoring system.

RP-PMM-7: Identify Unmonitorable Components as a Risk

Application owners and maintainers should identify any component that cannot be monitored with existing CMS IT PM tools as part of the application project risk register.

Rationale:

Application components that cannot be monitored represent an application performance risk and need to be tracked in the risk register. Rather than being wholly unmonitorable, such components may require different monitoring capabilities than the infrastructure possesses. Risk mitigations must be provided.

RP-PMM-8: Support Standards-Based Tools

Performance-critical applications hosted in CMS data centers should support standards-based centralized monitoring, alerting, and event management. It is more important, however, to integrate with the existing CMS monitoring infrastructure than to deviate to meet a standard.

Rationale:

This applies to all performance-critical production applications providing business or infrastructure services. A performance-critical application is one for which significant degradation of the application’s performance or availability may have an immediate impact—either directly or by impacting another application—on end users or business operations.

Exceptions may be required for some existing or COTS applications; however, data centers contain tools for monitoring most of these applications and their underlying platforms without requiring modification of the application code.

RP-PMM-9: Define Services in a Service Catalog

Services must be defined and collected into a Service Catalog before use. The Service Catalog should also contain service level agreements and definitions of metrics to ensure that all service consumers may understand service constraints.

Rationale:

Without service definitions, service level agreements cannot be defined. There are usually additional constraints (such as capacity) and dependencies (such as other subordinate services or resources) that should be described in the service definition as well. The Service Catalog is the collection of active service definitions.

CMS Performance Management Systems

This topic provides guidance on using and configuring Performance Management Systems.

Additional requirements and guidance for system logging and alerting, including enterprise logging tools, may be found in the TRA Network Services, Information Security Monitoring.

Event Rules and Situations

Event situations are predefined rules consisting of metrics, thresholds, conditions, and actions that automatically trigger and generate an event. When the conditions of a situation have been met, an event occurs and the IT PM monitoring system triggers predefined actions such as displaying an event indicator on the IT PM portal dashboard.

For each event situation, rule sets may be defined using the following elements:

  • System Platform Distribution specifies the systems to which the situation rules apply.
  • Formulas for conditions being tested may include one or more pairings of metrics and thresholds in a given condition formula.
  • Expert Advice includes comments or instructions the IT PM system will display along with an event in the Event Results workspace.
  • Actions may include:
    • Displaying an alert
    • Generating an issue ticket
    • Sending alert messages
    • Sending commands to managed systems
  • Until conditions close the event after a period of time or when another situation is true.

Performance Data Collection Methods

CMS employs multiple approaches for monitoring applications in the CMS Processing Environment. These approaches are broadly classified as Platform Monitoring and Transaction Monitoring.

Platform Monitoring uses methods and tools specifically developed to monitor operating systems and specific server platform and infrastructure products, including:

  • Operating systems and logs
  • Application servers
  • CICS on z/OS
  • Database management systems (DBMS)
  • Cluster managers and load balancers
  • Messaging servers
  • Business integration middleware to service-enable legacy business applications
  • Energy management devices, server instrumentation, and resources
  • Virtual servers
  • Web infrastructure

Platform Monitoring may be used to monitor the status, resource utilization, and performance of servers and infrastructure products. Platform monitoring methods may also support Transaction Monitoring.

Transaction Monitoring methods measure the response time and availability of transactions, including the end-to-end response time and availability as experienced by end-users of an application as well as the response time and availability of individual components.

Many metrics can be measured using COTS IT PM capabilities. Some application-level metrics, such as the end-to-end response time experienced by end-users of an application, may require application-specific methods, which in some cases may be built into the application. Such application-specific methods may be implemented using various techniques, including:

  • Monitoring web services and web servers directly
  • Executing real or synthetic transactions from multiple locations to test a business service end to end. The application owner and developer must define these transactions. In some cases, the application may have to be modified to ensure the test transaction does not modify or disclose production data. See RP-SC-8: Consider Synthetic Transactions. CMS Cloud supports New Relic synthetics.
  • Enabling an application to report its own performance data to a monitoring server using the standard Application Response Measurement (ARM) protocol, which may require modification of the application software

Security

IT PM data and services require the same protections as other system management data and services. These protections are described in the CMS ARS and the CMS TRA and include Management Zone compliance, encryption requirements, access control policies, and data retention policies.

IT PM data transiting the network requires protection to reduce the risk of unauthorized access or disclosure of sensitive application and business data that could be collected as part of the IT PM process.

It is important to secure IT PM data at REST.

Related CMS ARS Security Controls include: MP-4 - Media Storage and SC-28 - Protection of Information at Rest.

Guidelines for Performance Monitoring Planning

This topic provides CMS guidance on developing a Performance Monitoring Plan for an application. The suggested methodology begins with considering business service performance, formulating a Performance Monitoring Strategy, developing a Logical Application Performance Monitoring Plan, and finally, producing the Detailed Application Performance Monitoring Plan, which becomes part of the application’s Operations and Maintenance Manual (OM&M). Of these, only the OM&M and the Detailed Application Performance Monitoring Plan are required of all applications. It is strongly suggested that a Performance Monitoring Strategy be presented at the Preliminary Design Review (PDR).

Throughout the development of an Applications Performance Management Plan, the business owners and application maintainers should work closely with the CMS Enterprise Monitoring and Management Team.

The Enterprise Monitoring & Management Matrix below also discusses the Enterprise Monitoring & Management Matrix Form that application owners should use to request the setting of thresholds and actions.

Performance Monitoring Strategy

The Performance Monitoring Plan should implement a performance monitoring strategy based on the following factors:

  • Business needs, which drive business service performance requirements.
  • Technical needs, which drive technical performance requirements.
  • Impact of monitoring, which drives cost and performance tradeoffs.

The next subtopics address each of these factors in steps representing the development of a Performance Monitoring Strategy.

Business Service Performance Considerations

An application’s business services exist to support business operations. These operations have performance and capacity requirements based on the business mission, such as the need to process and store a specific number of claims per week. A poorly performing application can bog down business operations, reduce customer satisfaction, and even cause additional work or penalties for business operations. Therefore, some of the performance standards (and metrics) for application performance must be derived from the business operations’ performance and capacity requirements.

For each business application service, a business owner or operations manager must consider:

  • What is the slowest permissible service response time before users or operations are significantly impacted?
  • How often can the service be unavailable before users or operations are significantly impacted?
  • For how long can a service be unavailable before users or operations are significantly impacted?

A “significant impact” to users or operations is when one of the following occurs:

  • Operations are slowed or delayed to the point where costs increase, backlogs occur, or time-critical deadlines for business processes are missed.
  • A user’s productivity is reduced. If the user cannot work as quickly due to waiting for the system to respond, that represents a significant impact. In time-sensitive business operations such as a call center, slow system response can directly translate into higher operations costs, negatively impact staff personnel performance ratings, and (in a call center), directly increase call costs.
  • A user becomes frustrated with the service performance. For interactive services accessed directly by the public, higher costs or productivity impacts from poor service performance may not be obvious. For example, a frustrated user may make mistakes, give up on his or her online request, or file complaints—all of which have negative consequences for CMS.

The service fails to meet user performance expectations. Today’s users expect computer applications to perform rapidly, at continuous, maximum “internet speeds” and with very high availability. When CMS does not meet these expectations, this reflects poorly on the Agency and may cause users to seek older, less efficient, more costly methods of interacting with the Agency, such as paper and postal mail.

Step 1: Define Service Performance Objectives

For IT PM to be an effective tool that avoids the above significant impacts, application performance must be monitored for compliance with business-driven performance standards, and data center IT must understand the impact of sub-standard performance on business operations and mission. After evaluating the foregoing Business Service Performance Considerations, the application owner or operations manager should produce a list of Service Performance Objectives. For each objective, the following information should be provided:

  • A clear, measurable definition
  • An Acceptable Performance Standard as a minimum, maximum, or range of performance
  • A delineation of where the standard applies (everywhere, regional offices, public internet, etc.)
  • A monitoring threshold as a minimum, maximum, or range of performance at which the business owner must be notified whenever the application’s performance is degraded beyond the threshold.
  • A brief description of the business service and its role in business operations
  • A brief explanation of the potential impact to business operations if the business service is degraded or unavailable for a few seconds, minutes, or hours.

Armed with the above information, data center IT staff can plan for and monitor application performance, proactively anticipate performance problems, and prioritize restoration of services when problems occur. The Service Performance Objectives should be reviewed with the data center operators and may be used to define Business Service Performance Standards in application hosting service agreements.

Step 2: Map Service Performance Objectives to Components and Transactions

Application developers and maintainers can identify additional critical performance metrics. They should begin by reviewing the Service Performance Objectives identified by the business owner and the operations manager and mapping the objectives to specific child transactions and application components. External services such as database operations, remote procedure calls, and remote queues should be included as components addressed by the Service Performance Objectives. Additional Service Performance Objectives may be defined as needed, based on non-functional requirements, component dependencies, or other technical requirements.

Step 3: Develop a Performance Monitoring Strategy

The Performance Monitoring Strategy identifies which transactions, resources, and application performance data will be routinely monitored in production.

Prioritize Service Performance Objectives

The application developer and maintainer should identify which Service Performance Objectives represent the highest risk in probability of performance degradation and impact of performance degradation. Higher-risk objectives require more frequent monitoring, while lower-risk objectives may be unmonitored or monitored only when required for testing changes or diagnosing problems.

Identify Transactions and Resources for Monitoring

Once it has been decided which Service Performance Objectives will be monitored, the application developer and maintainer should identify the simplest, most minimal, lowest-impact set of transactions and resources for routine monitoring to determine whether those objectives are being met. The set can comprise any combination of application resources, specific application transactions, or application-provided data.

Cost and Performance Tradeoffs

Over-monitoring an application can have negative impacts. Over-monitoring can degrade performance, increase resource utilization, produce volumes of performance data, and increase costs and complexity. By carefully selecting which transactions will be monitored and when they will be monitored, the negative impacts of monitoring can be controlled and minimized. The list of proposed transactions for monitoring should be refined accordingly.

Logical IT PM Plan

The Performance Monitoring Strategy is the basis for the initial Logical IT PM Plan. This plan should contain:

  • A list of application resources to be monitored.
  • A list of specific application transactions to be monitored.
  • Definitions of application-provided performance data.
  • Recommendations about where monitored transactions should originate (data center, call center, regional office, etc.) and when or how often they should be measured.

These details should be based on the standards defined for the Service Performance Objectives, the risks associated with the Service Performance Objectives, and the possible impact of monitoring the transaction. To minimize the number of items monitored, the plan should avoid any unnecessary or duplicate application resources, specific application transactions, or application-provided data.

When complete, the Logical IT PM Plan should include tables or matrices providing traceability between Service Performance Objectives and the resources, transactions, or data for routine monitoring. It is unnecessary to include any items that will not be monitored routinely.

For each Service Performance Objective, the Logical IT PM Plan should identify some combination of application resources, specific application transactions, or application-provided data that may be employed to ensure that Service Performance Objective is being met.

It is important for the application developer to consult with the data center operators because many elements may already be monitored with existing tools, without requiring modifications to the application.

Detailed IT PM Plan

The application developer or maintainer works closely with data center operators and the CMS Enterprise Monitoring and Management Team to create the Detailed Application Performance Monitoring Plan by defining the properties for each resource, transaction, or data element in the Logical Application Performance Monitoring Plan.

The resources, transactions, or data may each have different types of properties, including:

  • Identity properties – that define the type or class of an application or transaction.
  • Context properties – that may be unique for each instance of a given application or transaction. Examples include userid, processid, or business information, such as a claim number.
  • Relationship properties – that show how one transaction relates to another. Typically, these are parent / child relationships.
  • Metric properties – that describe measurements of the resource, transaction, or data elements. Typical examples include status, response time, time of day, blocked time, number of bytes, number of records, number of threads, and queue size.

Once approved, the Detailed Application Performance Monitoring Plan becomes part of the application’s OM&M.

Enterprise Monitoring & Management Matrix

An Enterprise Monitoring & Management Matrix Form is used to request the setting of thresholds and actions. The matrix form must be submitted to the CMS Enterprise Monitoring and Management Team that will review and approve the request and implement the settings.

Table - Contents of the Enterprise Monitoring & Matrix Form summarizes the major columns in an Enterprise Monitoring & Management Matrix Form.

Table - Contents of the Enterprise Monitoring & Matrix Form
Form ColumnDescription
Identity Properties
Resource Name

Names the resource to be monitored. A resource may include a physical or virtual resource (memory, bandwidth, etc.), a process starting or dying, a queue, or a transaction.

EXAMPLE: Disk space

Monitored Resource Description

Contains a short statement that explains the reason for monitoring the resource.

EXAMPLE: Warning threshold exceeded for disk space usage 85%

Context Properties
Host NameIdentifies the host name or device name to be monitored.
Relationship Properties
Metric or Event DependencyIdentifies any dependencies on other events/thresholds. Usually there are NONE.
Metric Properties
Condition/Threshold Reached

Lists the specific metric value (in bold) or the specific indication the event has occurred.

EXAMPLE: Maximum Disk Space Percent >=85

Sampling Interval

Identifies the interval between samples for the metric or event.

EXAMPLE: 15-minute intervals

Occurrences within Sampling Interval

Identifies situations in which alerts are needed when a specified number of metric / event occurrences takes place within the Sampling Interval.

EXAMPLE: Alert if more than 3 occurrences during a Sampling Interval

Criticality

Lists the criticality of the detected situation, such as Critical, Warning, or Informational.

EXAMPLE: Warning (Yellow)

Action to be Taken

Identifies the action to be taken when metric or event is detected.

EXAMPLE: Send email notification to support team at <email address>

Situation/Monitor Name

Names the monitor performing the monitoring task.

EXAMPLE: UNIX_DiskSpace_Monitor

Software as a Service (SaaS)

Software as a Service Introduction

Software-as-a-Service (SaaS) is a cloud-based software distribution model that delivers applications to end users remotely over the internet, rather than within the enterprise. Applications hosted by third-party service providers and delivered as subscription services are increasingly popular and encouraged at CMS. Instead of requiring traditional application development and maintenance, SaaS services rely on users to configure and customize their applications. As a result, CMS business owners may find that SaaS Clouds offer faster implementation and lower cost than other forms of Cloud computing.

CMS uses the NIST SP 800-145 definition of software as a service. It is recognized that mapping commercial or government products to this model is not always straightforward because many providers produce clouds that have characteristics of one or more of these models.

Government IT systems have additional requirements that are often overlooked in commercial SaaS products. Before utilizing any SaaS, it is important to understand the policies of the service provider, their security model, and the security requirements that remain the customer's responsibility. SaaS, like all cloud services, operates with a shared security responsibility model where the provider and the customer each retain important security roles and responsibilities. The FedRAMP program establishes security baseline standards for cloud services and specifies agencies’ responsibilities when using these services. However, FedRAMP does not guarantee these services are suitable for any given CMS workload. An Authority to Operate (ATO) is still required for any cloud services used. See BR-SAAS-2.

To help CMS teams understand and manage SaaS risk and make good business decisions around SaaS usage, CMS has developed a SaaS Governance Program. Through this program, CMS teams can obtain information about SaaS applications that are already approved for use within the Agency as well as request help in reviewing new SaaS providers. The CMS SaaSG Dashboard (password required) shows  the applications already in use and under review. Further information is available at the CMS SaaS Governance website (see CMS SaaS Governance).

Data collection for official government business must be done on systems granted a CMS Authorization to Operate (ATO) (see CMS ATO for additional information) and that includes all clouds of all types. The responsibilities for securing the confidentiality, integrity, and availability of data throughout its life cycle remain the same regardless of implementation. SaaS services, like all CMS applications, must comply with the CMS ARS and CMS TRA, including CMS management of access policies, authorization, and user identities.

When using a SaaS to host public CMS websites and web services, it is important that those sites and services be represented as “.gov” domain names with valid SSL certificates (per Business Rule BR-F-18). As with all government websites, the HTTPS-Only Standard per OMB Memorandum M-15-13, Policy to Require Secure Connections across Federal Websites and Web Services, June 8, 2015, applies. This gives visitors to those sites the confidence that they are dealing with a legitimate CMS business site and that their information is encrypted in transit with “the strongest privacy and integrity protection currently available for public web connections.”

Business Rules for Software as a Service

SaaS Business Rules

BR-SAAS-1: SaaS Clouds Are Defined by NIST SP-800-145

CMS uses the NIST SP 800-145 definition of software as a service. It is recognized that mapping commercial or government products to this model is not always straightforward because many providers produce clouds that have characteristics of one or more of these models. If uncertain, consult with the CMS TRB.

Rationale:

The HHS designation of “Third Party Web Site” (TPWS) has been confused with SaaS. Some have argued that a given website is a TPWS and not a cloud (therefore not requiring FedRAMP certification). TPWS applies only to the analysis for privacy concerns. That is, when a CMS website redirects a user to a TPWS, it is instructing the user’s browser to shift its attention to another address. The user will be interacting with a website that is not operated on behalf of CMS, does not have a contractual relationship with CMS, and may use data in ways that CMS cannot control. Examples would include popular social media and news sites.

A service that collects official government data must be used under contract and with an ATO, even if it is possible to acquire the service without payment. The privacy and security of CMS data is at stake.

OMB Memorandum M-10-23, Guidance for Agency Use of Third-Party Websites and Applications, states that:

The term “third-party websites or applications” refers to web-based technologies that are not exclusively operated or controlled by a government entity, or web-based technologies that involve significant participation of a nongovernment entity. Often these technologies are located on a “.com” website or other location that is not part of an official government domain. However, third-party applications can also be embedded or incorporated on an agency’s official website.

There may also be unknown licensing issues with using what appears to be a free service. Many services that are free to an individual are not free to an organization. Integrating these capabilities into CMS systems places CMS and the government at risk of violating copyright laws.

BR-SAAS-2: SaaS Must Have a CMS ATO

Projects may only use SaaS if they hold a CMS Authority to Operate (ATO) issued by the CMS Designated Authority, nominally the CMS CIO. Although not absolutely required, the SaaS should have FedRAMP Certification, especially in scenarios where the SaaS may store or process sensitive information. Without FedRAMP Certification to inherit security controls from, the business owner must address the controls and risks in the SaaS environment, which may be a burdensome undertaking. If the SaaS environment will be storing or processing CMS data, it is a CMS Processing Environment (as defined in CMS TRA Foundation, Processing Environments) and must comply with the CMS ARS, CMS TRA, RMH and other relevant CMS and federal requirements.

Related CMS IS2P2 Cloud Computing requirements include: CMS-CLD-1, CMS-CLD-1.1, and CMS-CLD-2.

Rationale:

A CMS ATO is required by FISMA, OMB Memorandum, Security Authorization of Information Systems in Cloud Computing Environments, December 8, 2011; OMB Memorandum M-24-15, Modernizing-the-Federal-Risk-and-Authorization-Management-Program, July 25, 2024; and CMS IS2P2 with Cloud Computing requirements: CMS-CLD-1 and CMS-CLD-2. Projects may only use SaaS if they hold a CMS ATO issued by the CMS Designated Authority, nominally the CMS CIO. Also, per CMS-CLD-1.1, a SaaS product without current FedRAMP authorization, a CMS Provisional Authority to Operate (P-ATO) may be granted following a Rapid Cloud Review (RCR).

BR-SAAS-3: Ensure CMS Security May Perform Periodic Security Assessments

Projects must coordinate with the cloud and CMS security to allow designated CMS security personnel to perform non-destructive security testing and assessment, with prior communication with the SaaS service, including penetration testing.

Related CMS ARS Security Controls include: CA-07 - Continuous Monitoring, CA-7(4) - Risk Monitoring, CA-08 - Penetration Testing (High and Moderate HVA Systems), and CM-07- Least Functionality.

Rationale:

Ensuring that CMS data and services are secure cannot be satisfied by a one-time check. CMS security must perform periodic security assessments to identify risks to CMS data and services. Testing of cloud services will require concurrence from the Cloud Service Provider. Assessments are performed by ISPG, another federal government agency, or a contractor working on behalf of ISPG or another federal government agency.

BR-SAAS-4: Plan for Data Archival to Comply with Federal Records Management

Projects must establish how CMS data will be archived (for long-term records management purposes). Data within the SaaS Cloud must be catalogued by a CMS data steward and appropriate disposition recorded to ensure data is stored for the appropriate time period and in the appropriate manner.

Related CMS ARS Security Controls include: AU-11 - Audit Record Retention, MP-6 - Media Sanitation, SI-12 - Information Management and Retention, SI-12(3) - Information Disposal

Rationale:

Commercial SaaS providers do not typically retain data after contract expiration. Federal Records Management guidelines require specific handling of records at system disposition. Archival is typically performed to storage outside of the SaaS cloud.

BR-SAAS-5: Establish SaaS-Specific Contingency Program

Projects using SaaS Cloud providers must establish a contingency program plan, independent of the SaaS Cloud Service Provider, that meets the business owner’s RTO / RPO requirements.

Projects must ensure that gaps in the SaaS Cloud provider’s recovery plan be accounted for and handled. Disaster recovery must also include recovering any integration between the CSP and CMS, such as CMS Cybersecurity Integration Center (CCIC) integration, EUA integration, or CMSNet.

Contingency Planning and Disaster Recovery may be conducted within the same Cloud Service Provider (typical) or another provider (atypical.)

There are several critical requirements associated with the contingency program that must be addressed:

  • Business recovery processes may include the following documents: Business Continuity Plans, Disaster Recovery Plans, Continuity of Operations Plans (COOP), Crisis Communications Plans, Critical Infrastructure Plans, Cyber Incident Response Plans, Insider Threat Implementation Plan, and Occupant Emergency Plans.
  • It is essential all recovery-related plans be coordinated.
  • It is essential that testing of all recovery-related plans be coordinated.
  • It is essential all recovery-related stakeholders (and personnel) be periodically trained on responsibilities.

Business owners should also anticipate and plan for disruption of service such as voluntary or involuntary transfer from a SaaS service to an alternative.

Related CMS ARS Security Controls include: CP family of controls, in particular, CP-2 - Contingency Plan and CP-2(1) - Coordinate with Related Plans.

Rationale:

Although many Cloud Service Providers have contingency plans for their own purposes, the responsibility for contingency planning for CMS services lies with the project team and CMS business owner. Disaster Recovery is just one element of a Contingency Program.

BR-SAAS-6: Perform Configuration Management

All configuration settings and code changes must be placed under configuration control by project teams. Configuration Change control must be also implemented. The baseline configuration and settings must also be managed. Specialized tools focused on SaaS Security Posture Management (SSPM) have become available in the market and CMS is currently evaluating various tool options. See the CMS Cybergeek SaaS Governance Program web page (SaaS Governance) for more information and how to contact the SaaS Governance team to get current information on SSPM tools

Related CMS ARS Security Controls include: the entire CM family of controls.

Rationale:

SaaS Clouds often eliminate the need for projects to manage the software that runs a service. There may still be configuration settings and changes performed to set up a product before use by CMS users. Managing configurations is essential to ensure reproducible configurations.

There may also be configuration-specific data items, such as SSL / TLS keys or SSH keys that must be installed and managed (during key expiration, for example), as well as user privileges and data access restrictions.

Projects must establish how CMS life-cycle reviews and system promotion will be conducted to maintain orderly promotion of code from development to production. Configurations and code should be stored in a CMS-approved source control system.

Operating SaaS applications does not absolve project teams from these responsibilities.

BR-SAAS-7: Integrate with CMS CCIC

Projects using clouds must integrate with the CMS Cybersecurity Integration Center. This means transferring event and incident logs periodically to the CCIC.

Related CMS ARS Security Controls include: AU-02 - Event Logging, AU-03 - Content of Audit Records, AU-06 - Audit Record Review, Analysis, and Reporting, CA-02 - Control Assessments, CA-06 - Authorization, CM-03 - Configuration Change Control, and CM-07 - Least Functionality.

Rationale:

CMS cloud-based systems must be integrated with CCIC to participate in the global CMS security monitoring. Because of the wide variety of different SaaS clouds and their ability to integrate with external services, the exact form of integration will vary. The objective should be to participate with the CCIC in sharing information about security and application status of SaaS-hosted applications.

To meet CCIC requirements, it may be necessary to obtain direct feeds from the SaaS service or, lacking that, indirect feeds using external monitoring and polling. Please refer to CMS TRA Network Services, Security Services and CCIC Integration chapters.

BR-SAAS-8: CMS Data Must Always Reside in the U.S.

CMS system owners must ensure that CMS data is not processed, transmitted, transferred, or stored outside the United States and its territories. This maintains the jurisdiction of the Privacy Act of 1974 and the Health Insurance Portability and Accountability Act of 1996 (HIPAA).

Rationale:

CMS data must be stored, transferred, and processed entirely (and exclusively) within the U.S. to eliminate the possibility that foreign powers might obtain access to CMS data and information.

SaaS Recommended Practices

RP-SAAS-1: Integrate with CMSNet

Projects using clouds should integrate with CMSNet to reach other CMS processing environments and transfer data securely.

Rationale:

CMSNet is a secure transport between most CMS data centers and cloud enclaves. Using CMSNet reduces the possibility that data could be intercepted during transit because, in addition to using encryption, CMSNet is a private network with limited access.

RP-SAAS-2: Comply with CMS Defense-in-Depth Architecture

Projects using clouds should follow the defense-in-depth practices when possible.

Rationale:

It is understood that SaaS providers often do not reveal or discuss their internal architectures, which makes complying with the multi-zone architecture difficult. The following steps can provide comparable risk reduction offered by a Defense-in-Depth strategy like the CMS Multi-Zone Architecture:

  • Use a Web Application Firewall (WAF)
  • Use CMS-issued certificates
  • Use a Content Distribution Network with DDoS protection
  • Ask the CSP to provide documentation for these steps and services

The CMS IS2P2 section 4.1.3 provides guidance for ensuring compliance with the CMS Multi-Zone Architecture. In addition, IS2P2 System and Communications Protection (SC) requirements SC-1.1.3, SC-1.1.3.1, SC-1.1.3.2, and SC-1.1.3.3 should be consulted.

RP-SAAS-3: Continuous Monitoring

Projects using SaaS should integrate with CMS continuous monitoring initiatives.

Related CMS ARS Security Controls include: CA-07 - Continuous Monitoring and CA-7(4) - Risk Monitoring.

Rationale:

NIST requires continuous monitoring to address security impacts. This is complementary to, but different from, continuous operational monitoring, which seeks to assess the operational condition of an IT system.

RP-SAAS-4: Integrate SaaS with CMS Identity Management Systems

Projects must establish how CMS’s Identity Management Systems will be integrated with the SaaS (if possible).

If not available directly, projects are responsible for establishing how they will conduct user onboarding / offboarding to ensure that only active, authorized users are allowed in and CMS is not paying for excess capacity/accounts, etc.

Rationale:

The CMS Identity Management Systems (such as EUA) are used across CMS systems to ensure a common set of processes for onboarding and offboarding users as well as authentication. This assures CMS, for example, that employees who leave and contractors whose contracts expire are quickly and reliably removed from authorization lists. Using such a system also ensures that employees and contractors can benefit from reduced or single sign-in (R/SSO), greatly reducing the burden on managing credentials.

Keys and Secrets Management

Keys and Secrets Management Introduction

Key and Secrets Management (KSM) is the practice of managing the full life cycle of cryptographic and digital data to which access must be strictly controlled. KSM facilitates central management in operating applications and access to data and applies automation and digital records tracking to the problem of key proliferation and data access. KSM can also help with replacing credentials more frequently, which reduces the risk of compromise.

A secret is a digital authentication credential to be presented to a system, while a key is a specific cryptographic artifact (as well as a type of secret). Consider some of the secrets that applications must manage:

  • User passwords
  • Root passwords
  • Application and database passwords
  • Auto-generated encryption keys
  • Private encryption keys
  • Application Programming Interface (API) keys
  • Application keys
  • Secure Shell (SSH) keys
  • IAM secret keys
  • Programmatic access keys (project / client IDs)
  • Authorization tokens
  • Bearer tokens
  • Certificates
  • Private certificates (e.g., Transport Layer Security (TLS), Secure Sockets Layer (SSL))
  • RSA and other one-time password devices
  • Account tags
  • Passphrases
  • Any other application tokens that are deemed confidential

In addition, many of these secrets are registered and managed in external systems, with corresponding expiration dates, renewal dates, and other timelines that must be managed lest these credentials expire and cause an outage. There are also compliance issues with key management. It is crucial to ensure adherence to the Least Privilege principle while managing credentials and access control. Increasingly, keys must be issued and rotated automatically across many applications. This burden exceeds the capability for manual tracking with spreadsheets, desktop databases, or paper. It is a complex problem.

Key management differs from secrets management. The basic life cycle steps for each are below.

  • Key Management Life Cycle
    • Generation
    • Use and distribution
    • Storage
    • Escrow and Backup
    • Accountability and audit
    • Key Rotation
    • Key Revocation
  • Secrets Management Life Cycle
    • Store a secret.
    • Modify secret attributes.
    • Associate a user or application with a secret.
    • Issue a secret to a user (on demand).
    • Destroy a secret.
    • Audit usage.

In key management, particularly when implemented with Hardware Security Modules (HSM), the keys never leave the system. This differs from secrets management where the secrets may be retrieved from the system by approved users / roles. A secrets management system is designed to store any kind of secret, while key management system is used to specifically manage keys for encryption.

PREFERRED

Per the CMS TRA’s recommended practice RP-KSM-1, “Applications Should Use a KSM to Manage Keys and Secrets.” CMS Cloud recommends use of the AWS Secrets Manager.

The ISPG Key Management Handbook summarizes practices related to secrets management as well as key management. In particular, the section Key Management Lifecycle Best Practices provides a comprehensive guide.

References:

Business Rules for Key and Secrets Management

KSM Business Rules

BR-KSM-1: KSM Auditing Must Be Enabled and Connected to CMS Logging Infrastructure

KSM systems must be configured with full auditing enabled, ensuring that all access to secrets and keys is recorded in tamper-proof logs.

Related CMS ARS Security Controls include: AU-2 - Event Logging.

Rationale:

One of the key advantages of using a KSM is to obtain a clear history of who accessed what secrets and when. A KSM audit log gives this history. The KSM audit log should be connected via CMS enterprise logging tools to the CMS log management infrastructure.

BR-KSM-2: KSM Must Use RBAC with Separate Roles by CMS Application Environment

KSM must be configured with separate credentials for each CMS application environment (Development, Test, Implementation, and Production).

Related CMS ARS Security Controls include:

Rationale:

By segmenting secrets and credentials across environment boundaries, CMS reduces the likelihood that credentials may leak and be used out of context. This reduces the risk of accidental or malicious use.

BR-KSM-3: Ensure the KSM Has Sufficient Availability for Your Applications

KSM should provide sufficiently robust availability to avoid any constraints on applications from accessing key resources like databases and API-based services.

Related CMS ARS Security Controls include: CP-2 - Contingency Plan, CP-7(3) - Priority of Service, and PL-8 - Security and Privacy Architectures.

Rationale:

Given the critical role that the KSM plays in an application, a KSM can be a single point of failure for an application. Inability to obtain credentials reliably can lead to service disruption. If necessary, it may be necessary to operate multiple KSM instances to offer sufficient availability for all applications.

BR-KSM-4: Audit Logs for Credential Leaks

It is possible to leave debugging statements in place in software that cause a leak of credentials into logs. Automated scans can also be performed to search for logging statements that reference known fields containing secrets.

Related CMS ARS Security Controls include: AU-2 - Event Logging, and AU-3 - Content of Audit Records.

Rationale:

It is a common mistake to log credentials to audit logs. Even with a KSM, it is possible to log credentials that could be used to compromise systems, including the KSM. To mitigate this risk, perform code reviews on logging statements to prevent such occurrences.

KSM Recommended Practices

RP-KSM-1: Applications Should Use a KSM to Manage Keys and Secrets

Configure CMS applications to use KSM whenever possible as opposed to custom solutions.

Related CMS ARS Security Controls include:

Rationale:

Persistent data files that contain secrets create both ongoing maintenance overhead and security vulnerabilities. File encryption alone does not constitute a secrets management solution. Alternative approaches using files, environment variables, UNIX pipes, and others are not as comprehensive, secure, or scalable as a KSM. Using alternatives often leads to one-off solutions that might be appropriate for one application but will not scale across the organization, leading to training and data management problems.

RP-KSM-2: Configure Applications to Use Dynamic Credentials

Instead of static credentials, use dynamic credentials managed by the KSM to provide access.

Related CMS ARS Security Controls include:

Rationale:

Dynamic credentials are automatically rotated and managed by the KSM. This renders them only valid for short periods and sometimes only once (depending on the configuration), This makes a credential less likely to be used maliciously.

RP-KSM-3: Consider Credential Injection into VM Images during Startup

Some VM managers and clouds allow for credential injection to occur during VM startup. This has the advantage of not requiring VMs to request credentials. Instead, this guarantees that a VM receives only those secrets.

Related CMS ARS Security Controls include:

Rationale:

This recommended practice guarantees that a VM receive only the credentials injected during startup.

Mobile Devices and Applications

Introduction

This chapter provides guidance on CMS policies surrounding the use and management of mobile devices, as well as developing and managing CMS mobile applications and CMS services that communicate with mobile devices. This guidance is based on the following documents, as well as industry best practices and other guidance sources.

CMS policy aligns with HHS policy and thus references to HHS apply directly to CMS. CMS has the option to create more stringent policies in the future.

The HHS Mobile Applications Privacy Policy and HHS Policy for Mobile Devices and Removable Media provide guidance for managing and securing mobile devices, developing CMS mobile applications, and the protection of personal information (such as personally identifiable information [PII], sensitive personal information [SPI], and other sensitive information) on mobile devices.

  • The HHS Mobile Applications Privacy Policy focuses on providing guidelines to protect privacy in HHS mobile applications used by the public or HHS employees
  • The HHS Policy for Mobile Devices and Removable Media focuses on protecting information and information systems from risks related to the use of mobile devices for government businesses and the risks of using mobile devices to access HHS information systems remotely from outside of HHS facilities. The HHS Policy for Mobile Devices and Removable Media pertains to all HHS employees, contractors, and other personnel who use mobile devices (including authorized non-Government Furnished Equipment (GFE) mobile devices) or removable media to store, process, and/or transmit HHS information, or remotely access HHS information systems.

What Are Mobile Devices?

NIST SP 800-53 Rev. 5 provides the following definition for a mobile device:

A portable computing device that: (i) has a small form factor such that it can easily be carried by a single individual; (ii) is designed to operate without a physical connection (e.g., wirelessly transmit or receive information); (iii) possesses local, non-removable or removable data storage; and (iv) includes a self-contained power source. Mobile devices may also include voice communication capabilities, on-board sensors that allow the devices to capture information, and/or built-in features for synchronizing local data with remote locations.

Mobile devices include cell phones, smart phones, tablets, laptops, and other devices that store, process, and transmit HHS information. Note that the focus in this section will be non-laptop devices, as the security capabilities currently available for laptops are different than those available for smartphones, tablets, and other mobile device types. Also, mobile devices contain features that are not generally available in laptops (e.g., multiple wireless network interfaces, Global Positioning System, various sensors).

From a policy perspective, mobile devices fall into two categories: either government furnished equipment (GFE) or non-GFE devices, which may be personally owned or provided by a contractor or business partner. The HHS Policy for Mobile Devices and Removable Media provides details on the use of both GFE and non-GFE devices and what resources and HHS systems can be accessed.

Threat, Risks, and Vulnerabilities

Mobile devices pose threats, risks, and vulnerabilities if not provisioned or used correctly. Common threats related to mobile devices include:

  • Compromise of the mobile device from vulnerabilities in the device operating system or installed app
  • Device loss or theft
  • User credential theft through phishing, wireless eavesdropping, or social engineering
  • Device misconfiguration exposing enterprise information

All of these threats create risk to the confidentiality and integrity of CMS information. Mobile devices with remote access to sensitive data or CMS systems could be compromised to gain unauthorized access, which can put CMS information at risk and leave systems vulnerable to future attacks. Mobile devices that store, process, or transmit CMS information must be encrypted in compliance with CMS requirements. A failure to keep mobile software up to date can create a risk of compromise to the device.

To protect against these threats, the HHS Mobile Device and Removable Media Policy requires several core security principles be followed, including:

  • All mobile devices with access to HHS systems or information must be managed using a Mobile Device Management (MDM) solution, which enables:
    • management and control of device configuration
    • segregation of personal and federal information, including the ability to remotely sanitize federal information stored on the device
    • controlled access to CMS resources
    • monitoring of malware detection and analysis tools to ensure they are installed, implemented, and up to date
  • Mobile device data must be encrypted to protect against device theft and loss
  • Mobile devices must be explicitly authorized and tracked, including the mobile devices/device types and operating system/patch levels required for devices which access HHS systems or information
  • Multi-factor authentication must be implemented for the mobile device and for access to HHS information systems and resources including changing all vendor-supplied and default passwords to a complex password compliant with the HHS IS2P

The HHS Policy for Mobile Devices and Removable Media includes a threat model requirement that encourages operating divisions / organizations to assess the threat landscape in creating a mobile device program to address business specific risks. Requirements include developing threat models and performing risk assessments for each of the remote access methods, each type and ownership category of mobile devices, and the locations from which HHS information systems will be accessed remotely (e.g., user's home, contractor facilities, domestic travel locations, and international travel locations). The goal of this assessment is to create a tiered access approach which limits risk by permitting the most controlled and secure devices to have greater access to HHS information systems and resources, and restricting devices that have less control to have less access or no access.

The HHS Policy for Mobile Devices and Removable Media also incorporates security controls for use of mobile devices. Specifically:

  • Users must physically protect their mobile devices at all times
  • Users are restricted from syncing GFE mobile devices with non-GFE mobile devices and laptops of unapproved personal, vendor or commercial cloud services and external devices.
  • Users are required to read and adhere to the acceptable use policy for mobile devices and removable media
  • Users must use encrypted VPN communications to protect all federal information transferred to or from a mobile device

CMS Mobile Device Categories

Unmanaged Mobile Devices

Unmanaged mobile devices equate to personally owned mobile devices. These devices can access only public CMS systems and networks by default. Non-GFE mobile devices must not be allowed to access CMS systems and information if an MDM solution is not in place to assure complete segregation of personal and CMS information. Deviation from this requirement requires an approved risk assessment and a waiver granted.

CMS-Managed Mobile Devices

CMS-managed mobile devices shall be used only by the person authorized to use the device. CMS users must not load unauthorized software or illegal content onto their devices. CMS should monitor and log all wireless communication from mobile devices connected to the CMS network. CMS-managed mobile devices are subject to the following controls:

  • Email and all website communications are subject to routine or automated scans and can be removed based on threat.
  • Malware detection and analysis tools must be installed, configured, and kept up-to-date on both CMS-managed and Partner-managed mobile devices.
  • All CMS-managed mobile devices must be enrolled in the CMS MDM solution to enable the remote erasure of data in the event the mobile device is lost or stolen.
  • Users are not permitted to jailbreak their mobile device. Jailbroken or rooted mobile devices will be prevented access to CMS resources, information systems, and information.
  • Any additional or new connectivity to be provided to the mobile device via hardware, software or other methods must be controlled and approved by CMS.
  • All mobile devices, containers, and removable media must implement full-disk/full device encryption using FIPS 140-2 validated cryptographic modules in compliance with HHS encryption requirements
  • Users of mobile devices should enable Bluetooth and pair their devices only when it is needed and must disable Bluetooth when it is not being actively used.
  • Strong authentication for mobile devices must be enforced in accordance with the HHS Information Security and Privacy Policy (IS2P) by implementing the following:
    • Strong password authentication for accessing the mobile device
    • Two-factor authentication for accessing the container and when connecting to CMS information systems
    • Configuring mobile devices with appropriate access restrictions in compliance with CMS policies

CMS may identify and differentiate applications that are allowlisted or denylisted for use on CMS-managed devices and may allow apps to be loaded on the CMS-managed mobile devices only from CMS authorized app stores. These apps can only be obtained from CMS-authorized or CMS-managed app stores or an app vetting process that complies with FISMA requirements.

CMS Partner-Managed Mobile Devices

Mobile devices managed by a CMS partner (i.e., contractors and other agencies with a business or government relationship with CMS) shall be used only by the person authorized to use the device. CMS should monitor and log all wireless communication from mobile devices connected to the CMS network. The following controls apply to CMS Partner-managed mobile devices:

  • Any additional or new connectivity to be provided to the mobile device via hardware, software, or other methods must be controlled and approved by CMS.
  • MDM software and malware detection and analysis tools must be installed, configured, and kept up-to-date on both CMS-managed and Partner-managed mobile devices.
  • All Partner-managed mobile devices must have the capability to remotely erase all federal information stored on the device in the event the mobile device is lost or stolen.

Guest Mobile Devices on the CMS Guest Network

  • CMS-provided guest networks must comply with the HHS Policy for Mobile Devices and Removable Media and any applicable CMS ARS Security Controls including AC-18.
  • All data transmitted on these devices over a CMS guest network may be monitored, recorded, and disclosed at the discretion of CMS.
  • Users must not transmit sensitive information (e.g., PII, PHI, and CUI) and unencrypted federal information over guest wireless networks, including CMS guest wireless networks.

Mobile Software, Applications, and Data

Mobile applications can collect information from the device itself, such as location information and device identifiers. Operating divisions / organizations of CMS must have a baseline protection agreement for mobile applications comparable to the plan set up by HHS.

CMS applications for mobile devices must comply with the CMS TRA, CMS ARS, RMH, and other guidance, and should follow these best practices:

  • Use only CMS-authorized app stores to distribute apps
  • Collect only information necessary to achieve CMS’s mission
  • Use structured data entry methods, rather than freeform text entry, whenever possible to limit data collection and minimize data entry errors
  • Use standard best practices for mobile data encryption, recovery, and disposition
  • Ensure users have options to opt out and customize the mobile application’s features when appropriate, such as opting out of location-based services while still choosing to use other application services
  • Display a heads-up notification to users any time an action may impact PII
  • Leverage the OWASP Mobile Application Security Verification Standard (MASVS) as an industry standard for validating mobile application security
  • Mobile applications must go through a privacy risk assessment, and meet the privacy requirements outlined in the CMS Privacy Program Plan and other Privacy resources

Mobile Devices Business Rules and Recommended Practices

BR-MD-1: Mobile Devices (GFE and non-GFE) Must Be Authorized to Access CMS Systems and Government Data

To access non-public CMS systems and government data, GFE and non-GFE mobile devices require authorization.

Rationale:

Unauthorized mobile devices may lack important security controls and may be subject to security vulnerabilities which put the confidentiality and integrity of CMS information at risk. Mobile device vulnerabilities cold also be exploited by cybercriminals to access CMS systems. Authorization also enable CMS to track devices which have access to CMS systems, per HHS policy.

BR-MD-2: Mobile Devices Must Use Encrypted Communication to Access CMS Data

Mobile devices that access non-public CMS data must use encrypted communication channels between the device and the CMS services.

Rationale:

Mobile devices communicate via network paths that are subject to eavesdropping and interception, which can expose CMS sensitive information as well user credentials. Encrypting communications reduces the risk of data interception and ‘man-in-the-middle’ attacks.

BR-MD-3: Managed Mobile Devices Must Support Remotely Erasing All Stored Data If the Device Is Lost, Stolen, or Compromised

All mobile devices with access to CMS networks or systems must be managed via an MDM solution that can remotely sanitize mobile containers without impacting personal information on the mobile devices, and has the ability to remotely sanitize the entire mobile device when necessary. This applies to all mobile devices, both GFE and non-GFE, that are managed by CMS or CMS Partners.

Rationale:

Mobile devices are increasingly used for business purposes, increasing the likelihood that they contain sensitive data and applications. Sensitive information stored on mobile device could be compromised if the device falls into the wrong hands. Remote erase (also known as remote wipe) allows administrators to quickly remove sensitive information from the device.

BR-MD-4: Managed Mobile Devices Must Support Encryption for Internal and Removable Storage

This applies to all mobile devices, both GFE and non-GFE, that are managed by CMS or CMS Partners.

Rationale:

Mobile devices may contain sensitive personal or CMS business information that if breached, can cause significant problems for both the user and CMS. Encrypting data on mobile devices and any associated removable media renders the data unreadable if an unauthorized user gains access to the physical device.

BR-MD-5: Managed Mobile Devices Must Meet CMS Security Requirements

This includes, but is not limited to, CMS requirements for updated patches, complex passwords, smart card authentication, and collecting PII / PHI or sensitive data. This applies to all mobile devices, both GFE and non-GFE, that are managed by CMS or CMS Partners.

Rationale:

CMS security requirements for mobile devices are designed to reduce the risk of breach of CMS sensitive information through its remote access on a mobile device. These security requirements provide for protections as the user, access, application, and data levels to provide holistic protection for mobile devices and the data they contain.

BR-MD-6: Only CMS Authorized Applications May Be Installed on CMS-Managed Devices

CMS will identify and differentiate those applications that are allowlisted or denylisted for use on CMS approved devices. Apps that can be installed on the mobile devices are restricted to those within a CMS-authorized app store.

Rationale:

Apps downloaded through mobile device app stores are not guaranteed to be free of malware or other data-stealing capabilities. To reduce the risk of compromise of sensitive CMS or user data, or fraudulent access to CMS systems, mobile device users must only download approved apps from app-stores approved by CMS.

BR-MD-7: User Agreements Must Be In Place for Mobile Devices Accessing CMS Services and Data

Users are required to read and adhere the acceptable use policy for mobile devices and removable media, including the HHS Rules of Behavior for Use of HHS Information and IT Resources Policy. The agreements include requirements such as support for encryption, remote erasing, security controls, acceptable use, and trusted mobile applications.

Rationale:

Protecting CMS systems and information from the risks of mobile devices requires not only technical security controls but also the understanding and cooperation of mobile device users. User agreements create awareness of CMS security requirements, commitment to follow them, and identify the potential repercussions for non-compliance.

Internet of Things (IoT)

Internet of Things Introduction

The National Cybersecurity Center of Excellence (NCCoE), a part of the National Institute of Standards and Technology (NIST), develops cybersecurity solutions for different commercially available technology. One example where NCCoE applies standards and best practices is for Internet of Things (IoTs). IoTs are sometimes simply described as smart devices or systems that are connected to the Internet. According to NIST SP 1800-15, Securing Small-Business and Home Internet of Things (IoT) Devices: Mitigating Network-Based Attacks Using Manufacturer Usage Description (MUD), the term IoT is applied to the “aggregate of single-purpose, internet-connected devices” such as sensors, vehicles, thermostats, security monitors, lighting control systems, appliances, and smart televisions.

CMS guidance applies to IoT devices in CMS and CMS partner facilities, as well as external-facing CMS services that may interface with IoT devices of providers, contractors, or beneficiaries.

The NIST Cybersecurity for IoT Program provides guidelines and a framework for manufacturers of IoT devices, Federal Agencies, and consumer products that use MUD specifications. The MUD specifications require IoT devices to perform only their intended function and to include features to allow malicious attacks to be intercepted and mitigated. The MUD specifications and the NIST SP 1800-15 guidelines help ensure the security and integrity of data stored by the IoT devices, and the networks and systems that the IoT devices access. Unfortunately, CMS cannot assume all manufactured IoT devices, especially those already in the field, are compliant with the latest MUD specifications.

Medical devices that are designed to be integrated in the network to support healthcare are sometimes called mIoT devices. Networked medical devices have software that can become vulnerable to cybersecurity threats. These vulnerabilities pose threats to healthcare and require continued maintenance throughout the devices’ lifecycles to protect against malicious acts. Risk management of cybersecurity risks in mIoTs assists in reducing overall risks in healthcare. References include:

Securing Devices / Secure Data

Organizations under HHS, specifically the National Institutes of Health (NIH) and the Food and Drug Administration (FDA), have recommended security requirements for the use of IoTs within HHS.

Guidelines for IoT Platforms

IoT platforms are services, typically cloud-based, which connect IoT devices and include functions for device registration, device monitoring and management, and data transfer. IoT platforms often offer services for authentication and remote access control, data storage, and device data backup.

A recent article, Medical Internet of Things and Big Data in Healthcare from NIH discusses implementing IoT platforms. The NIH has observed that various parties, typically IoT device manufacturers, are trying to bundle the data streams of their IoT devices including data related to wearable devices and medical devices. This results in a massive influx of data from the sensors built into these IoT devices that can then be analyzed.

Beyond the NIST Cybersecurity for IoT Program, there are no CMS guidelines specifically for securing IoT platforms. However, any IoT platform that qualifies as a CMS Processing Environment (refer to the definition in TRA Foundation, Processing Environments) must comply with the CMS TRA, ARS, RMH, and all HHS and CMS security and privacy requirements.

Guidelines for IoT Devices

The FDA suggests that IoT device manufacturers should adopt the following practices: secure software or firmware updates by incorporating authentication to update processes, systematically update procedures for authorized users, and have a secure data transfer mechanism to and from the IoT devices.

CMS cannot depend on all IoT manufacturers to adopt these cybersecurity measures; however, for those IoT devices that are CMS managed, as a recommended practice CMS should seek at a minimum Manufacturer User Description IoT devices. Maintaining cybersecurity prevents the device from losing data, malfunctioning, open to security threats and general loss of data. Most importantly, maintaining cybersecurity keeps the patient healthy using a secure device.

CMS-Managed IoT Devices

Guidance for IoT devices managed by CMS or CMS partners should follow guidance as for CMS Managed mobile devices.

Unmanaged IoT Devices

Devices that are not managed by CMS or CMS partners are considered unmanaged and include personally owned IoT devices. Unmanaged IoT devices should follow similar guidance as for unmanaged mobile devices. That is, these devices can access non-public CMS services or networks if and only if a risk assessment has been approved, a waiver has been granted, and the device has been configured and provisioned so that there is an agreement with the owner that CMS has access to the data on the device and can use the data. These devices must have anti-virus and anti-malware software installed.

Other Challenges / Recommendations

IoT device manufacturers each have different proprietary protocols within the devices, which greatly complicates security, authentication, and data communication when using devices from different manufacturers. Encouraging the use of the draft Cybersecurity Practice Guide, NIST SP 1800-15 with MUD specifications as a framework by manufacturers is one way to assist with this communication issue and to reduce security concerns.

Risk Management

FDA considers the stakeholders for medical IoT devices to include the medical device manufacturers, the user, the Information Technology (IT) system integrator, Health IT developers, and an array of IT vendors that provide products that are not regulated by the FDA.

Mitigating cybersecurity threats to the user of the devices and to the operation of the device is one of the key concerns. Collaboration among its regulated stakeholders to alleviate concern for these cybersecurity threats is encouraged. To further improve risk management, the manufacturers of these IoT devices are encouraged to use the draft version of Cybersecurity Practice Guide, NIST SP 1800-15 when building the devices. The collaborative risk management approach strives for a consistent assessment and approach to the disruption of cybersecurity threats to device operations and users.

Standards and guidelines specifically for designing and securing IoT devices and platforms are still in an early state of evolution. Any IoT platform that qualifies as a CMS Processing Environment (please refer to definition in TRA Foundation, Processing Environments)must comply with the CMS TRA, ARS, RMH, and all HHS and CMS security and privacy requirements.

Business Rules and Recommended Practices for IoT

BR-IoT-1: CMS IoT Platforms Must Comply with CMS Requirements for CMS Processing Environments

Any IoT platform that qualifies as a CMS Processing Environment must comply with the CMS TRA, ARS, RMH, and all HHS and CMS security and privacy requirements.

RP-IoT-1: CMS-Managed IoT Devices Should Comply with NIST SP 1800-15 and the Latest MUD Specifications

Encouraging the use of draft Cybersecurity Practice Guide, NIST SP 1800-15 with MUD specifications as a framework by manufacturers is one way to assist with this communication issue and to reduce security concerns.

Disaster Recovery

Disaster Recovery Introduction

This chapter presents a general overview of CMS practices, services, and guidance for Disaster Recovery (DR), which should be implemented in accordance with appropriate security and CMS Continuity of Operations (COOP) requirements. CMS Emergency Preparedness and Response Operations (EPRO) manages all CMS COOP Planning including the assignment of CMS Mission Essential Functions (MEF), which influence DR parameters. This includes formulating guidance and establishing common objectives for CMS and its components to develop a viable, enterprise-wide, state of resilience for DR and COOP capability.

The primary focus of this chapter is DR as it relates to IT operations covering the following:

  • Hardware – Networks, servers, desktop and laptop computers, wireless devices and peripherals, etc.
  • Cloud configurations – Compute, data, virtual networks, load balancers, etc.
  • Connectivity to service provider – Fiber, cable, wireless, etc.
  • Software applications – Electronic data interchange, electronic mail, enterprise resource management, office productivity, etc.
  • Data and restoration – Backup and recovery
  • Environment – Secure computer room with climate control, conditioned, and backup power supply, etc.

Reference Documents

This chapter is not all inclusive. It complements and incorporates existing policies, standards, and procedures for CMS, HHS, DHS, and OMB, thereby offering an architectural view of the standards. Where there are conflicts, the following standards, and any successor documents, will take precedence:

COOP and DR:

Other related guidance

DR Key Concepts and Definitions

To ensure the audience fully understands the conceptual foundation for this guidance, the following terms, definitions, and explanations are presented in a sequence that progressively build the foundation for the guidance. The Glossary defines additional terms for the reader’s benefit.

Continuity Plan

A document that details how an individual organization will ensure it can continue to perform its essential functions during a wide range of events that impact normal operations. It is a common misconception that Information System Contingency Plans (ISCP) are synonymous with or a substitute for a continuity plan. ISCPs complement continuity plans, and the two plans should be coordinated. However, an ISCP does not account for how the organization will continue performing its essential functions during a disruption to normal operations. The ISCP impacts the organization’s continuity plans and operations by identifying recovery time objectives for key systems that support the performance of functions, including essential functions.

Continuity of Operations (COOP)

COOP is an effort within the Executive Office of the President and individual Departments and Agencies to ensure that essential functions continue to be performed during disruption of normal operations.

COOP focuses on restoring organization’s Mission Essential Functions (MEF) at an alternate site and performing those functions until normal operations can be resumed.

Disaster

Any event, whether human-caused, act of nature, or technology failure, that causes disruption to operations beyond acceptable time limits, thereby threatening the survival of critical services.

Disaster Recovery Plan

A plan which is executed that is designed to restore and resume normal IT operations after an event of a disaster within an acceptable timeframe as determined by the agency. Plans exist at both system and enterprise levels.

National Essential Functions (NEF)

The NEFs represent the overarching responsibilities of the Federal Government to lead and sustain the Nation and shall be the primary focus of the Federal Government leadership during a catastrophic emergency.

Primary Mission Essential Functions (PMEF)

PMEFs are those validated MEFs that must be performed to support or implement the uninterrupted performance of NEFs.

Mission Essential Functions (MEF)

MEFs are the essential functions directly related to accomplishing the organization’s mission as set forth in statutory or executive charter. MEFs may be unique to each organization. MEFs serve as key continuity planning factors for Department/Agencies to determine appropriate staffing, communications, information, facilities, training, and other requirements in an event of a disaster.

Essential Supporting Activities (ESA)

A subset of government functions that are determined to be critical activities. Functions that support performance of MEFs; and must be included in the organization’s DR planning process.

High Value Assets (HVA)

All MEFs’ supporting systems and ESAs are considered HVAs according to the Office of the Federal Chief Information Officer. HVAs are:

“Federal information systems, information, and data for which an unauthorized access, use, disclosure, disruption, modification, or destruction could cause a significant impact to the United States’ national security interests, foreign relations, economy, or to the public confidence, civil liberties, or public health and safety of the American people. HVAs may contain sensitive controls, instructions, data used in critical Federal operations, or unique collections of data (by size or content), or support an agency’s mission essential functions, making them of specific value to criminal, politically motivated, or state sponsored actor for either direct exploitation or to cause a loss of confidence in the U.S. Government.”

High Availability Assets (HAA)

HAAs are IT technology architecture where redundancy and failover processes are built into a system to maximize uptime and availability. The concept of HA is to achieve an uptime of 99.999 percent or higher, which equates to just a few minutes per year of downtime.

On Premises Hot Site

Hot sites are locations that operate 24 hours a day with fully operational equipment and capacity to immediately assume operations upon loss of the primary facility. A hot continuity facility requires on-site telecommunications, information, infrastructure, equipment, back-up data repositories, and personnel required to sustain essential functions.

On Premises Warm Site

Warm Sites are partially equipped office spaces that contain some or all of the system hardware, software, telecommunications, and power sources. To become active, a warm facility requires additional personnel, equipment, supplies, software, or customization. Warm sites generally possess the resources necessary to sustain critical mission/business processes but lack the capacity to activate all systems or components.

On Premises Cold Site

Cold Sites are typically facilities with adequate space and infrastructure (electric power, telecommunications connections, and environmental controls) to support information system recovery activities.  These facilities are typically neither staffed nor operational daily. Teams of specialized personnel must be deployed to activate the systems before the site can become operational. Although, basic infrastructure and environmental controls are present (e.g., electrical and heating, ventilation and air conditioning systems), systems are not continuously active.

Multi-Site Cloud

Multi-site cloud configuration, like hot sites, are duplicated fully operational environments. The environments are replicated in a different physical location and are load balanced to receive network traffic. This is also known as an active-active configuration.  In the event of a failure, traffic is rerouted immediately to the other identical environments minimizing interruptions. Most cloud service providers offer physical separation via larger geographic areas typically called regions and more granular delineations within the regions known as availability zones. Different cloud service providers may have slightly different terminology, but the concepts and functionality remain similar. Multi-site cloud environments require no additional personnel or equipment, as they are cloud based and failover is automatic. Additional cloud resources are needed to duplicate the environments in the cloud.

Multisite Cloud DR conceptual diagram (page 14)

Warm Standby Cloud 

Scaled down duplicates of the master resources are created in different geographical regions/zones but are in a minimalized operational (standby) state. Data is synchronized, key services are operational. However, network traffic is only redirected when a failure is detected, and the standby environment is activated. This is active-passive.

Warm Standby Cloud DR conceptual diagram (page 15)

Pilot Light Cloud

Absolute minimal duplicates of critical resources are kept alive, such as data synchronization. Once a failure is detected, duplicate resources are created from scratch using predefined configurations which are used to build the new environment.

The predefined configurations and Infrastructure as Code (IaaS), which are used to build the new environment after a failover, should be backed up both locally and remote where it can be accessed to build the new environment.

Cloud DR conceptual diagram (page 16)

DR Phases and Implementation

This topic provides high-level overview of the DR operational phases to support DR strategy for CMS. The scope of this topic is limited to IT systems and related components for DR that should be part of the overall COOP strategy. Please refer to CMS Continuity of Operations Plan for more details.

The DR process includes the following four phases:

  • Phase I - DR Readiness and Preparedness
  • Phase II - Activation
  • Phase III - Continuity Operations
  • Phase IV – Reconstitution

DR Readiness and Preparedness

Phase I includes preparatory activities that are performed in advance of a disruptive event to ensure CMS is ready to respond effectively and recover quickly in an event of a disaster.

DR Readiness is the ability of an organization to respond to a continuity activation. Readiness is an aspect of the planning and training activities, but ultimately CMS leadership is responsible for the overall determination of readiness. It must know that systems can perform disaster recovery operations, including essential functions before, during, and after emergencies.

CMS readiness activities are divided into two key areas:

  • Organization readiness and preparedness
  • Staff readiness and preparedness

The topics below provide overview of activities that are part of the Readiness and Preparedness phase.

DR Plan

The DR plan defines the hosting facilities processes for systems reestablishment and involves a set of policies, tools and procedures to enable the recovery or continuation of vital technology infrastructure and systems following a disruption. The DR Plan may sit at a data center or with the owner of a cloud virtual data center where applicable.

The DR plan:

  • Defines recovery from various emergencies, usually physical events, that result in disruption to service that inhibits access to primary facility infrastructure for an extended period of time.
  • Focused on restoring operations of an information system, application, or computer facility at an alternate location after an emergency.
  • May support Component Business Continuity Plan (BCP) to recover supporting systems for ESAs at an alternate facility once it has been established.
  • May support or more Information System Contingency Plans (ISCP) to recover individual systems.
  • Only addresses emergencies or information system disruptions that require relocation.

Test, Training and Exercises

  • Participate in CMS Test, Training Exercise (TTE&E) Program activities including corrective actions for incorporation into the following year’s continuity tests and exercises.
  • Conduct quarterly communication and IT testing of the CMS Continuity Facility at the Maryland Mission Support Center (MMSC).
  • Participate in annual Eagle Horizon (EH) exercise in accordance with the CMS TTE&E plan.

Communications

  • Conduct periodic testing of information technology (IT) and communication systems supporting continuity readiness.

Emergency Relocation and Teleworking

  • Ensure ERG members have logistical information to relocate to alternate facilities.
  • Monitor and maintain awareness of potential threats and developing situations through liaison with the HHS.
  • Ensure accounting for personnel and reporting procedures are in place.

Activation

DR Plan activation is a scenario-driven process that facilitates flexible and scalable responses to a full spectrum of emergencies and other events that could disrupt CMS operations. Activation is not required for all emergencies and disruptive situations that do not require relocation to an alternate site, since other actions may be deemed appropriate to maintain normal operations. A senior management official, such as the CIO, has the ultimate authority to activate the plan and to make decisions regarding spending levels, acceptable risk, and interagency coordination.

Reconstitution

Reconstitution Objectives

The overall objectives of the CMS Reconstitution Plan are to identify and outline the processes and procedures to return to normal operations once the Administrator or successor determines that reconstitution operations for resuming normal business operations can be initiated. Specific plan objectives are:

  • Provide an executable plan for transitioning back to efficient normal operational status from continuity operations or devolution status once a threat or disruption has passed.
  • Coordinate and pre-plan options for organization reconstitution regardless of the level of disruption that originally prompted the organization to implement its continuity plans. These options must include moving operations from the continuity facility or devolution site to the primary operating facility, a temporary operating facility, or a new or rebuilt operating facility.
  • Outline and execute the necessary procedures, whether under a standard continuity of operations event or under a devolution event, for conducting a smooth transition from the continuity facility to the primary operating facility, a temporary operating facility, or a new/rebuilt operating facility.
  • Ensure a safe location for organization staff to resume normal organization operations.
  • Reduce or mitigate disruptions to organization operations.
  • Ensure and validate reconstitution operations readiness through an integrated continuity TT&E program and operational capability.

Facilities and Information Technology

Following an emergency event, organizational facilities and infrastructure may be partially damaged or destroyed by the event. Physical damage to facilities and IT requires damage assessments and facility procurement procedures for addressing facility issues.

Reconstitution Processes for Facilities and Information Technology include the following:

  • Damage Assessment
  • Facility Recovery Plan
  • Information Technology Recovery Plan
  • Consolidated Facility and Information Technology Recovery Implementation Plan

Transition to Normal Operations

When reconstitution is nearly completed and the rebuilt or new CMS facility is ready to be occupied, reconstitution personnel must verify that:

  • All IT systems, communications, essential records, and other required capabilities are tested, available, and fully operational.

DR Capability Considerations

DR Parameters

The DR parameters can assist in determining specific requirements in an event of an identified disaster event. The DR parameters help determine required capabilities for DR. The key DR parameters are:

  • Recovery Point Objective (RPO) – The amount of time, prior to a disruption or system outage, to which mission/business process data can be recovered (given the most recent backup copy of the data) after an outage.
  • Recovery Time Objective (RTO) – The maximum amount of time that a system resource can remain unavailable before there is an unacceptable impact on other system resources, supported mission/business processes, and the MTD.
  • Work Recovery Time (WRT) – The maximum amount of time needed to verify the system and/or data integrity.
  • Maximum Tolerable Downtime (MTD) – The total amount of time the system owner/authorizing official is willing to accept for a mission/business process outage or disruption.

Disaster Recovery Parameters shows a timeline for DR parameters:

Disaster Recovery Parameters (page 13)

The DR parameters help CMS business and system owners build the DR requirements into CMS life cycle (e.g., TLC). Per the CIO Memorandum Recovery Time Objective Requirements, July 22, 2020, these recovery parameters should be determined within the Business Impact Analysis (BIA).

The activities for DR requirements at different phases of CMS life cycle:

  • Initiate
    • Identify a service requirement and the need for an IT system
    • Conduct a Business Impact Analysis (System BIA) to determine the operational parameters including RTO, RPO, MTD and WRT
    • Enter all system attributes in the appropriate systems of record including CFACTS
  • Develop
    • Complete a system design including DR parameters following the SDLC (TLC)
    • Update the System BIA to reflect design trade-offs
    • Document how to meet the DR parameters in the DR Plan and ISCP
    • Test to verify the DR plan and ISCP
  • Operate
    • Conduct ongoing testing, training & exercises (TT&E) of DR and ICSP
    • Participate in COOP exercises
    • Complete role based training
    • Document lessons learned, including changes to the System BIA
  • Retire
    • Transition any required functionality to existing systems
    • Update DR plans
    • Follow the system Disposition Plan to retire the system

DR Tiers

The concept of tier-based disaster recovery was developed by users of IBM mainframes as a rule of thumb for planning. CMS no longer uses this. The CMS Information System Contingency Plan (ISCP) guidance is based on NIST SP 800-34, Contingency Planning Guide for Federal Information Systems and NIST SP 800-53, Security and Privacy Controls for Information Systems and Organizations. There are no industry standards for the number of DR tiers. DR tiers are designed to meet the needs of the business and therefore vary by organization.

The DR tiers shown in Table - CMS Disaster Recovery Tiers serves as a tool organized by system RTO requirements. The choice of which tier a system belongs to is influenced by the following business system characteristics:

  • Whether it supports one or more MEFs
  • Whether it has identified financial and operational impacts
  • Its dependencies on other systems
Table - CMS Disaster Recovery Tiers
TierRTOApproachData Loss Recovery
1< 1 dayHighly automated takeover using component mirroringZero or near-zero data loss; fast and reliable
21–5 daysHighly automated takeover using replicationMay lose a few hours to a day of data
36–29 daysA combination of highly automated takeover using replication and manual processesMay lose several hours to days of data
430-plus daysNo backup hardware, data stored off-site on tapeData recovery may not be possible; best case may lose days/weeks of data

 

Recovery Facilities

Facility recovery involves the restoration or replacement of damaged facilities or the use of alternate facilities. This includes communications capabilities such as voice and network. There are various types of backup sites. For on-premises data centers, the options are; cold, warm, and hot. For cloud-based datacenters the options are; multi-site, warm-standby, and pilot light. Each type varies in cost depending on the effort required for implementation.

Table - Facility Recovery Options for Disaster Recovery
Site TypeDefinition
On-Premises Cold SiteTypically, facilities with adequate space and infrastructure (electric power, telecommunications connections, and environmental controls) to support information system recovery activities. This is the least expensive type of backup site and takes the longest amount of time to configure and restore to normal operations.
On-Premises Warm SiteCompromise between a hot site and a cold site. While the physical building, connectivity, and hardware may be available, the data may still need to be copied over to the system.
On-Premises Hot Site

A complete duplication of the original site, with full computer systems as well as near-complete backups of user data.

A hot site requires a mirroring architecture with real-time synchronization between the two sites. This facilitates business system relocation with minimal loss to normal operations. It requires no additional labor during an event as the architecture is fully automated.

 

Table - Cloud Recovery Options for Disaster Recovery
Site TypeDefinition
Multi-Site CloudA complete duplication of original virtual network environments (compute, data, networks and services) in an alternate region. Both sites are active and available for use. Traffic is directed through a load balancing mechanism via DNS. Multi-site cloud requires a mirroring architecture with real-time synchronization between the two sites. This facilitates business system relocation with minimal loss to normal operations. It requires no additional labor during an event as the architecture is fully automated.
Warm-Standby CloudA duplicate, but minimal version of the original virtual network environments (compute, data, networks and services) in an alternate region. Failover site is in standby mode and only receives data synchronization from live environment. During failover, the load balancing mechanisms redirect traffic and the failover environment is auto-scaled to size based on load requirements.
Pilot Light CloudA duplicate, but minimal version of the original virtual network environments but only data and essential services are live in an alternate region. Failover site is in standby mode and only receives data synchronization from live environment. During failover, a new environment is built out using infrastructure as code (scripting, defined configurations and base machine images). This is then appropriately connected to the synchronized data. The load balancing mechanisms then redirect traffic.

 

System and Data Recovery

The results of Business Impact Analyses (BIAs) can provide insight into whether CMS will incur substantial financial and operational impacts from a disruption of CMS business functions and operations exceeding a certain number of days. Those functions and operations with a critical timeframe must be resumed or recovered to a minimum level of service within a specified timeframe. The goal is to ensure that CMS survives as a viable entity in the event of a disaster. For critical functions, such as MEFs, this may require cloning the normal business and working environment to ensure success.

Based on analysis of the situation after a disaster occurs, the recovery resources requirements can be divided into the three distinct recovery phases as described in Table - Recovery Resource Requirements.

Table - Recovery Resource Requirements
Recovery PhaseResource Requirements
Activation and NotificationFor the first 3 days following the disruption, resource requirements focus on the ability to provide rapid resumption or recovery of time-sensitive business functions.
Recovery & Continuity OperationsIf the displaced business functions cannot return to their normal place of business within 2–3 weeks, temporary office space, furnishings, equipment, telephones, etc. will be acquired, and will be sufficient in size to accommodate all of the employees and operations of the affected functions.
ReconstitutionReconstitution or replacement of the CMS facility will require acquisition of replacements for all damaged furnishing, equipment, etc. It may also involve extensive demolition and construction, depending on the extent of the damage. If the facility cannot be restored, either another facility will be acquired or that location will be demolished, and a new building constructed.

Disaster Recovery Services

This topic provides guidance on various solutions/services for consideration for DR based on requirements for CMS.

DR Alternatives

The goal of the recovery strategy is to provide a recovery capability that balances implementation and ongoing costs against potential impacts and exposures to CMS.

The agency can consider several available alternatives—from taking no action or doing nothing to full system mirroring/high availability. The recovery strategy selected may be a combination of two or more of these alternatives:

  • Do nothing. This option provides no additional funds or resources for providing recovery capability, (i.e., “do nothing”). This would leave CMS unprotected against a major disruption of agency operations. There must be an authorized risk acceptance and approval process in place for this option.
  • Normal restoration/replacement. This option depends on the restoration or replacement of the damaged or destroyed facilities, furnishings, equipment, etc. in the normal course of business.
  • Self-recovery (VDC). CMS can provide its own disaster recovery capability by having at least two Virtual Data Center (VDC) facilities where similar types of functions and operations are performed. Leveraging HA architecture for this solution is an option. This alternative requires that the facilities:
    • Are sufficiently geographically distant from each other
    • Contain enough “spare” room
    • Ensure that the entire baseline infrastructure the affected functions will require is pre-installed or have all hardware for “hook-up” pre-installed
  • Replication (VDC). Maintaining availability of CMS critical applications is a key part of disaster recovery and business continuity planning. CMS could choose to provide its own disaster recovery capability by replicating (also known as mirroring) the critical systems at an off-site third-party location. Again, leveraging HA architecture for this solution is an option.

The recovery alternative selected for CMS systems should be chosen based on operational requirements for fulfilling the approved RTO, cost efficiencies, and MEFs supported.

“Disaster Recovery as a Service” (DRaaS) encompasses combinations of alternatives listed above.

DR Assessment

A DR assessment determines the completeness of a system's DR readiness through periodic testing, training, and exercise (TT&E) activities.

DR Readiness

DR Readiness is the ability of an organization to respond to a continuity activation. Readiness is an aspect of the planning and training activities, but ultimately CMS leadership is responsible for the overall determination of readiness. It must know that systems can perform disaster recovery operations, including essential functions before, during, and after emergencies.

Disaster Recovery Business Rules

BR-DR-1: Annual Review of Disaster Recovery Plans

Disaster recovery plans and their supporting documents must be reviewed and reevaluated on an annual basis or upon a significant change to the operating environment.

U.S. Department of Homeland Security, FEMA Office of National Continuity Programs, Federal Executive Branch Continuity Program Management Requirements, Federal Continuity Directive, August 2024

Rationale:

TT&E requirement under Testing.

BR-DR-2: (Rule Withdrawn after TRA 2024R4): Disaster Recovery Tier Selection

BR-DR-3: All CMS FISMA systems must have a plan for DR

As required by FISMA.

Related CMS ARS Security Controls include: CP-2 Contingency Plan and CP-4 Contingency Plan Testing and Exercises.

Rationale:

DR planning and preparation are essential for resumption of services following a disaster.

BR-DR-4: Required Risk Analysis, System BIA, and ISCP

A Risk Analysis, System Business Impact Analysis (BIA), and ISCP must be documented for all applications/systems for CMS to correctly select the appropriate Disaster Recovery Tier for the application.

Related: CMS Target Life Cycle (TLC) Initiate/Develop phases.

Completion of Risk Assessment, Systems Business Impact Assessment, and Information System Contingency Plan are required activities in preparation of process to receive Authority to Operate

BR-DR-5: (Rule Withdrawn after TRA 2024R4): Number of Disaster Recovery Tiers

BR-DR-6: The BIA is the Primary Determinant of DR Parameters

The system Business Impact Analysis (BIA) provides the basis for the system's Recovery Point Objective (RPO), Recovery Time Objective (RTO), Work Recovery Time (WRT), and Maximum Tolerable Downtime (MTD).

Rationale:

These parameters must be based on business requirements.

 

APPLICATION DEVELOPMENT

Application Development Introduction and Principles

Purpose

The intent of this Application Development Guidelines chapter is to establish the core set of business application development guidelines for Centers for Medicare & Medicaid Services (CMS) software developers and maintainers. CMS is confident that consistent adherence to a standard set of development practices that align with industry best practices should produce greater repeatability of process and higher quality in application delivery.

Given the rapid pace of technology evolution, CMS requires a common baseline for software engineering that is technology agnostic but relies on information technology (IT) industry best practices. The focus of this chapter is to disseminate the practices and business rules deemed most beneficial to the production of high-quality, secure software that is compatible with the CMS IT environment.

Scope

This chapter represents the mandatory business rules (BR) and recommended software engineering best practices (RP) that CMS and CMS / Contractor partners should use in the CMS Processing Environments. The guidance and policies stated in this chapter reflect the CMS agreed-upon, industry, and government best practices to support the most viable approach for CMS that meets legislatively mandated security and privacy requirements as well as current technical standards and specifications.

Artificial Intelligence-assisted coding tools can improve developer productivity and code quality when used appropriately. However, these tools present unique security considerations that must be addressed through proper controls and practices. Guidance on securely integrating AI-assisted coding tools within CMS environments while protecting sensitive information and maintaining code quality is included in Software Coding Business Rules and Recommended Practices.

Concepts and Terminology

Role Definitions for Application Development

Application development is a cross-disciplinary function. The production of quality software solutions requires the cooperation of multiple roles. Table 1. Software Engineering Roles presents the application development / software engineering roles and definitions critical to understanding and applying the guidance of this supplement. These roles and definitions are consistent with the Release Management guidance in this section and commonly used across CMS.

Table 1. Software Engineering Roles
RoleDefinition
Business AnalystThe party eliciting and documenting requirements in collaboration with the application’s stakeholders.
System Developer (or simply Developer)The party producing the initial implementation of a software-based system, including such activities as design, coding, unit testing, integration testing, building, and releasing software.
System Maintainer (or simply Maintainer)The party maintaining an existing implementation of a software-based system, including such activities as producing software patches, tracking defects, producing change packages, and releasing software fixes. The system maintainer is responsible for producing a high-quality product, which includes functional capabilities (such as features) and non-functional requirements (such as performance, scale, and capacity).
System Operator (or simply Operator)The party operating a software-based system, including such activities as hosting applications, monitoring, allocating storage, performing backups, starting jobs, restoring systems, applying patches, upgrading systems, and maintaining inventory records. At CMS, this is typically a CMS Virtual Data Center (VDC) operator.
Hosting ProviderThe party providing the network, computing, and storage resources (physical or virtual) used by the system. Although the role of hosting provider is distinct from the system operator, the same party may perform these roles. At CMS, this is typically a VDC operator, but could also be a CMS-approved Cloud Service Provider (CSP).
Business OwnerThe party or parties for whom the system is developed or maintained.
End UserThe party or parties who will use the system.
Information System Security Officer (ISSO)The party in charge of adherence to security practices for the full life cycle of the system from the perspective of the project.
Software AssuranceThe party responsible for running security analysis software and interpreting the reports to continually improve the security quality of the software. It is a quality assurance role with a focus on security. As defined by contract, the system maintainer, system operator, or others may perform this role.

Note: In this chapter, the system maintainer is responsible for any role not specifically assigned to another.

Business Rules and Recommended Practices

This chapter presents both business rules and recommended practices. Business rules reflect CMS standards and are mandatory; conformance to recommended practices is optional but highly encouraged. In addition, recommended practices are likely to become future business rules at CMS’s discretion and following in CMS TRA Foundation, Architecture Change Request (ACR) process.

The provided rationale for each business rule and recommended practice offers additional insight into the reason for the rule or practice as well as context for interpreting the rule or practice. The rationale does not constitute normative guidance.

Internal and External Quality

Internal Quality is the quality that is apparent to the software engineers working on the system. It reflects the code, test data, specifications, and all other component artifacts of the system. Internal quality comprises the attributes of performance, maintainability, scalability, ability to operate, security, reliability, and resilience as understood by the engineering and operations staff as well as management.

External Quality is the quality that end users experience and includes end-user perception of system performance. User interfaces, reports, email notifications, and other forms of user-to-system communication are typical ways to observe external quality.

This chapter acknowledges the importance of both forms of quality.

Principles

Methodology Independence

The software engineering industry has produced various methodologies ranging from waterfall, spiral, and most recently Agile methods, such as SCRUM and Extreme Programming (XP). The guidance in this chapter takes no position on the merits or shortfalls of any method; however, it does establish standards for engineering discipline and practices that all methods must follow.

Technology Agnostic

The architecture described in this chapter meets the modern definition of service-oriented architecture (SOA), which is CMS’s standard. Any references to products and technologies are as examples only: this chapter is technology and product agnostic. Therefore, unless specifically mandated, adoption of any product or technology is a decision beyond the scope of this document.

The term “commercially available” means both proprietary and Open Source Software (OSS). In addition, software that is custom-written by or for the Federal Government may be available, because of the SHARE IT Act of 2024, without licensing costs (CMS will still incur support costs). Find more details about this in the CMS Open Source Strategy section.

Introducing New Software

Application developers must perform due diligence to select software that is sustainable, supported, and good value for the system’s CMS customer. Due diligence, here, might consist of:

  • Evaluate the use case for this software, to avoid buying more features than will be needed.
  • Check the most recently published System Census (internal link) to see whether another CMS system uses it:
    • Compare the other system’s use case and rationale for features used, in common or distinct.
    • Learn from the other system’s implementation of the software, including security controls.
    • Discover any licensing or acquisition vehicle considerations. Consult the Enterprise Software Licensing website for details about existing Enterprise License Agreements (ELAs).
  • If appropriate, check the SaaS Governance (SaaSG) page’s resources to see whether it’s been approved or reviewed for another CMS system.
  • Perform market research to compare candidate software with alternatives.

System Design Principles for Cloud and Virtualized Environments

This topic emphasizes new design principles that enable the Cloud and other virtualized environments. These principles are essential to support CMS’s transformation of the bulk of its processing environments to more virtualized environments. The following topics describe relevant design approaches, issues, and security guidance applicable to each design principle. Additional information is available in CMS Hybrid Cloud: Cloud Consumption Playbook.

Design for Scale

Software developers must design software that scales to meet business needs.

Related CMS Acceptable Risk Safeguards (ARS) Security Controls include: SA-2 - Allocation of Resources, SC-30, and SC-2 - Separation of System and User Functionality.

Design for Reliability and Resilience

Business requirements for availability determine whether to implement a system using a highly available design. These requirements are documented as part of disaster recovery (DR) planning. In addition, designs should account for rapid recovery in the event of an availability issue.

Related CMS ARS Security Controls include: CP-9 - System Backup, CP-9(8) - Cryptographic Protection.

Design for Loose Coupling of Components

Service-oriented, Application Programming Interface (API)-based architectures encourage loose coupling of components, with benefits that include resilience, scalability, and flexibility.

The counterpart to loose coupling is tight cohesion. Tight cohesion requires that a service do only one thing. Another way to approach this problem is the single responsibility principle (SRP). SRP holds that a service should have only one reason to change. If a service must change for more than one reason, it probably does too many things.

Design for Elasticity

Where architecture scalability addresses the ability to meet business needs, elasticity addresses the necessary automation to quickly respond to demand, automatically, within predefined limits. Architectures should account for the presence of external control and will require some startup time before newly allocated resources are fully available. It is important from a cost perspective to consider elasticity when establishing elasticity controls.

Related CMS ARS Security Controls include: SA-2 - Allocation of Resources.

Design for Rapidly Deploying Environments

Architectures should be rapidly deployable, in automated fashion, to the designated target environments. This ensures a seamless, repeatable process across the entire development and deployment life cycle.

Design for Caching

For static data that does not change often—such as images, video, audio, Portable Document Format (PDF) files, JavaScript (JS), and Cascading Style Sheet (CSS) files, CMS recommends caching mechanisms to keep the data as close to the end user as possible. External-facing content is usually deployed outside a data center using Content Distribution Networks (CDN), such as Akamai (internal link), to be as close as possible to the user. This closeness helps mitigate access latency. It also reduces load on intermediate systems as requests are satisfied further from core systems. However, the caching mechanism must check for changes to avoid fetching expired data. Everything delivered to a user’s browser should specify an expiration mechanism. At a minimum, content should specify an HTTP directive like “Cache-Control: must-revalidate.” The need to advise a user to “Clear your cache” indicates a defect where stale content (script or stylesheet) is cached. Security access control must also be maintained.

Nearly all CMS public-facing systems use Akamai for front-end or edge-level caching. Back-end caching may be unnecessary in the CMS federated data mesh environments, but where used it needs careful engineering to:

  • Ensure that limits on persistence are enforced to avoid stale objects.
  • Protect any personal information, used to enable users to resume work across disconnections or outages, so that only the same user can access it.
  • Limit the persistence of cached personal information to avoid the risk of Freedom of Information Act (FOIA) exposure.

Cloud environments provide database services that specifically address caching.

Design for Dynamic Data Near Processing

Distributed data center architecture introduces the risk of Internet latency and the added cost of bandwidth. Keeping dynamic data close to processing elements is a software best practice to reduce network latency and reinforce cost efficiency. Transferring data into and out of a Cloud architecture requires different design principles than transferring data within the Cloud. In some cases, it may be more efficient to transfer a large volume of data into a Cloud infrastructure to take advantage of parallel processing capabilities. Applications that consume data generated in a data center should be deployed locally to take advantage of lower latencies. Design decisions to deploy applications across data centers raise cautionary performance and cost implications.

This design principle emphasizes the cost of inter-data center (or inter-cloud) processing.

Design for Parallelization

Designing hardware and software to carry out calculations in parallel rather than in sequence leads to more efficient and overall faster processing times for operations. Operations on data from request to storage should take advantage of opportunities for parallel manipulation. Processing data collections can also take advantage of parallelization, with incoming requests distributed across multiple nodes for greater efficiency. The choice of parallelization may dictate the use of radically different technology from small-scale or serial processing.

In addition, developers should consider methods and techniques that limit or avoid resource locking, which is a common side effect of poorly performing parallel architecture.

Design for Security

CMS requires implementation of adequate security to protect all elements in the virtualized application. The security design must protect data-in-transit (please refer to BR-SA-15) as well as data-at-rest (please refer to BR-SA-16) according to very specific rules, as summarized in the following paragraphs. Developers must meet or exceed the policy established in the latest published version of the CMS ARS as well as the business rules of the TRA Network Services section Security Services topic.

Threat Modeling

An important concept in designing applications with security in mind is the use of Threat Modeling (see RP-SS-8). Threat Modeling is a process for capturing, organizing, and analyzing a variety of application and threat information. It enables informed decision-making about application security risks. In addition to producing a model, the process also produces a prioritized list of security improvements to the conception, requirements gathering, design, or implementation of an application.

Threat Modeling works to identify, communicate, and understand threats and mitigations within the context of protecting something of value. Threat Modeling is a structured approach of identifying and prioritizing potential threats to a system, and determining the value that potential mitigations would have in reducing or neutralizing those threats.

The Threat Modeling process is simple to understand and execute for any project team, regardless of existing experience. The use of Threat Modeling:

  • Helps ADO teams to improve the security and compliance of their applications.
  • Provides documents to support and improve compliance in a variety of situations (e.g., internal or external assessments, ATO, impact analyses).
  • Aids penetration testing by providing information about the threats a system could face.

For detailed information on Threat Modeling and how to perform it, see:

AI Code Generation

The integration of AI in the CMS software development lifecycle can assist in a range of tasks, from initial planning to rapid-prototyping, piloting, and post-deployment maintenance. Modern AI development and tools can suggest code snippets based on the context of the codebase, develop and test new features, and even build entire working prototypes and applications.

However, AI-generated data and code can serve as an attack vector, introducing hidden bugs, security vulnerabilities, or performance issues if not carefully reviewed. Models may produce code that looks correct but fails when encountering edge cases or incorporates solutions that do not align with project requirements.

AI services present risks including but not limited to data leaks relating to system, programmatic, or other sensitive information, including credentials or market-moving information. For these reasons, ensuring human oversight, proper security configurations, and robust testing are essential elements to integrating AI into the CMS lifecycle.

AI Principles for Developers

The following principles are represented in the TRA Application Development Business Rules and Recommended Practices section, as well as the TRA Foundation Artificial Intelligence guidance.

AI Code Generation Accountability

  • Human-in-the-Loop
    Application developers and teams that use AI-generated code at CMS are ultimately responsible for that code, regardless of the tools and methodologies used. Applying proper testing and human code reviews is required for all AI-assisted code.
  • Continuous Human Oversight and Validation
    AI tools serve as a powerful assistant for code generation, completion, and documentation. These tools can assist with boilerplate code and suggest snippets based on the context of the codebase. However, changes that impact production systems require human review and approval.
  • The “trust” paradox
    Human validation is essential to counter the “trust paradox”: the AI’s output must never be implicitly trusted. This requires a cultural shift where the developer’s role is not to simply write code but to critically evaluate what has been generated for them.

Tool Selection and Data Privacy Considerations

  • External AI Tools and Data Rights Policy
    External tools or services that train or fine-tune AI models on CMS code, data, or metadata may not be used with non-public CMS information unless the CMS business or system owner has ensured compliance with CMS security and privacy requirements and appropriate CMS Data Use Agreements (DUAs) are in place if required. For inquiries, consult with the CMS Privacy Office.
  • AI Code Generation Tool Selection
    Individuals and application development teams interested in using AI assisted software development must perform due diligence for tool selection and methodologies, ensuring proper configurations and compliance with CMS IT policy and federal mandates. For inquiries, consult with AI Governance at CMS IT Governance.

AI Development Methodologies

AI-powered code generation and development assistance tools have undergone substantial maturation in recent years in industry, and are increasingly changing how software is developed, iterated on and maintained. The pace of innovation in this space continues to accelerate, with new capabilities, methodologies, and best practices emerging regularly. CMS recognizes both the transformative potential of these technologies and the critical importance of implementing them responsibly within our enterprise environment. This guidance will continue to adapt to technological advances while maintaining our commitment to security, privacy, and operational excellence.

“Vibe Coding” Policy Restrictions

Vibe Coding is a style of development that emphasizes speed, creativity, and intuition over rigid specifications. It involves developers diving straight into building, guided by instinct and real-time collaboration. This approach, fueled by AI tools, is effective for rapid ideation and reducing cognitive load.

Risks of “Vibe Coding”: While fast for initial concepts, “vibe coding” can be a chaotic and unpredictable process that doesn’t scale well with current technology. The model may generate code that appears correct, but it can subtly miss the point, ignore key project constraints, or fail to grasp the broader project context.

Acceptable Use of Vibe Coding: When speed and flexibility matter more than polish, “vibe coding” is suitable for mockups, prototypes, proofs of concept, low-risk internal tools and utilities, sandbox experiments, and pilots that are not High-Impact (as defined by M-25-21).

Prohibited Use of Vibe Coding: Do not use vibe coding in the software development lifecycle for initiatives where accuracy, security, privacy compliance, and long-term stability are essential. This includes all High-Impact, business-critical systems, or upper environments that handle sensitive data.

Suitable Approaches to AI-Assisted Development

The AI technology landscape is rapidly changing. Individuals and teams employing AI code generation should assess emerging best practices and methodologies, ensuring compliance with CMS IT policy (including security, privacy, and AI) and other federal mandates.

Here is a non-comprehensive comparative look at some of the methodologies used today:

  1. Formal Specification-Driven Approaches. This approach treats a formal, executable specification as the single source of truth. By treating the specification as the driver, any changes to the spec automatically flow into the code. This reduces technical debt, keeps documentation aligned, and strengthens long-term maintainability.
  2. Intuitive or Agent-Driven Workflows. These workflows rely on tools that act like agents: they break down goals, call tools, and iterate toward solutions. In this setup, developers focus on managing tasks, guiding prompts, and reviewing outputs. While fast and flexible, the ease of use can lead to “vibe coding,” which creates major risks in sensitive or regulated environments.
  3. Recommended Hybrid Approach for the Enterprise. The current most practical model is a hybrid of formal specifications and AI-native workflows. This balances the structure of formal specifications with the speed and adaptability of AI agents, while ensuring human oversight.

See RP-AI-9 for a description of this workflow.

AI in the SDLC — Shifting Left: Integrating Quality and Security

To manage the velocity of AI-assisted code generation, it is critical to shift quality and security left in the development lifecycle. Teams should embed automated tools that identify issues early and have proper code reviews before they reach production.

Recommended Practices:

  • Automated Testing: Utilize unit, integration, and end-to-end tests to validate the functionality of AI-generated code.
  • Linting and Style Checks: Enforce consistent code quality and adherence to project-specific conventions using automated linting tools.
  • Automated Security Scans: Integrate static application security testing (SAST) and software composition analysis (SCA) into the CI/CD pipeline to proactively detect vulnerabilities, malicious packages, or outdated dependencies in generated code.

See details of these in the TRA Application Development Business Rules and Recommended Practices section AI Assisted Coding.

Examples of AI Benefits in the Software Development Lifecycle

With proper planning, human review and oversight, AI technology can provide support to product teams throughout the SDLC for different types of AI efforts at CMS. The following is a non-comprehensive list of examples:

Requirements and Design

  • Requirements Generation: AI can be used to analyze natural language descriptions and transform them into structured requirements, user stories, or initial feature lists.
  • Architectural Guidance: AI can be used to suggest optimal architectural patterns, component structures, or UI/UX mockups based on project requirements, helping teams validate risky assumptions and make informed decisions early in the process.

Testing and Quality Assurance

  • Automated Test Case Generation: AI can analyze codebases to suggest and build unit, integration, and end-to-end test cases, assist in increasing test coverage, and reduce manual efforts.
  • Predictive Bug Detection: By analyzing historical data and code changes, AI can predict which areas of an application are most likely to fail, allowing teams to prioritize testing efforts.

Deployment and Operations

  • CI/CD Optimization: AI can analyze data to predict and prevent deployment failures, recommend adjustments to pipeline configurations, and optimize build times.
  • Monitoring and Maintenance: In production, AI-powered tools can monitor systems in real-time to detect anomalies and predict potential issues, often identifying problems before they impact users.

Documentation and Maintenance

  • Automatic Documentation: AI can analyze a codebase to automatically generate and update technical documentation, including API references, in-code comments, and user guides.
  • Code Refactoring: AI can identify code that is difficult to maintain or inefficient and suggest refactored, optimized solutions, helping to reduce technical debt over time.
  • Crucial Reminder: Regardless of the specific AI technology or the stage of the SDLC it supports, these technologies are tools that can augment human capabilities but should not replace human judgment at CMS. Humans are ultimately responsible for all outcomes and must apply critical oversight to all AI-generated or AI-assisted work.

Digital Service Delivery and Human Centered Design

CMS systems directly impact constituents. As a result, user-centric design emphasizes ease of use and empathy for users who will be interacting with their government. CMS has always been a leader in digital service delivery, converting paper-based processes to fully digital solutions. This transformation helps reduce costs to the taxpayer while simultaneously increasing reach and impact. CMS embraces practices that help deliver on these principles. Here are some resources:

Project Services (Context below)

Software development projects at CMS are responsible for providing the bulk of the services they require. This chapter defines three sets of services in support of software development efforts: cross-project services, minimal required engineering support services, and additional services. The following topics address each set of services.

Cross-Project Services

There are a few services provided for software development that cross projects. The following exceptions are typically cross-project services:

  • Technical review and consultative services provided by the CMS Technical Review Board (TRB)
  • Change control, provided by change control boards (CCB) that coordinate between projects with inter-dependencies
  • Data center support services, provided by the hosting provider (either a VDC or CMS-approved Cloud Service Provider)
  • Database management support services, which include data management support and data modeling
  • Security services, such as Security Control Assessments (SCA) or Adaptive Capability Testing (ACT)
  • Accessibility and Section 508 assessment services

Projects should discuss available cross-project services with their hosting provider because services may vary by provider.

Integrated Engineering Support Services

This chapter and its associated BRs rules and RPs assume the existence industry-common integrated engineering support services for application development. Appendix A provides a more complete description of the following engineering support services addressed in this topic:

  • Defect Tracking
  • Version Control
  • Continuous Integration
  • Build Automation
  • Package Repository
  • Test Automation
  • Package Deployment
  • Accessibility and Section 508 Testing

Context

CMS TRA Services Framework

The CMS TRA Services Framework (see CMS Services Framework) provides a standardized template for implementing systems at CMS. This template or pattern details security and interoperability requirements. The Services Framework architecture enables flexibility while continuing to apply defense-in-depth principles. The services framework applies to CMS data centers and cloud implementations.

The services framework is the basis for the CMS Multi-zone Architecture (see CMS Multi-Zone Architecture) where the framework services, detailed by the functions they performed, are mapped to zones within the architecture. CMS is focused on protecting CMS assets via security challenges and enforcing defense-depth-principles. Zones are a way of grouping and sharing resources based on a shared security posture. Distinct zones generally exist in legacy data centers but are not as prevalent in cloud implementations where virtual resources, such as network, compute, and storage are project based implementations and there is tight integration with other cloud service provider (CSP) provided services.

While CMS does not specify the number of zones, CMS does require that CMS data be at least three independent legitimacy tests away from the open Internet. A test is defined as the challenge, filtering and transformation of the data request.

CMS TRA Multi-Zone Architecture

The CMS TRA specifies a zone architecture that provides defense against security attacks and implements layers of challenges to ensure only authorized access to CMS resources. The services framework details the functions provided by each of the service types, and these services are represented by zones in the CMS TRA multi-zone architecture (see CMS Multi-Zone Architecture).

CMS does not require distinct zones for protecting assets, but it does require that data is protected by at least three security challenges. This is a minimum requirement, as additional security measures may be helpful in reducing the risk of a security breach.

The complete description of the CMS TRA Multi-zone Architecture appears in the CMS TRA  Foundation Business Rules and Network Services section, Security Services topic.

Related CMS ARS Security Controls include: SC-5 - Denial of Service Protection, SC-7 - Boundary Protection, and SC-32 - Non-Mandatory: Information System Partitioning.

Web Services or APIs

The CMS TRA describes a service-oriented architecture, supporting both interactive and batch processing modes. The Web Services topics in this chapter provide a more detailed description of web services.

SOAP web services should be documented with separately controlled Web Services Definition Language (WSDL) files (even if initially auto generated). XML-based REST (Representational State Transfer) services can be documented with WSDL 2.0 (or later) files. Java Script Object Notation (JSON)-based services can be described using JSON Schema.

Related CMS ARS Security Controls include: SC-8 - Transmission Confidentiality and Integrity.

Enterprise Messaging

Use of enterprise messaging is recommended for CMS applications. JSON based enterprise messaging increases security by abstracting object locations, moving data access logic closer to the database and requiring business applications to access data via message-based services rather than by direct SQL statements. This helps reduce vulnerability to SQL-enabled exfiltration of data and SQL injection attacks. Enterprise messaging provides additional resilience for applications because it allows for asynchronous coupling between sender and receiver. When enabled, enterprise messaging can provide guaranteed delivery of messaging where messages are transmitted in store-and-forward manner. Thus, if a receiver is unavailable, messages intended for that receiver are not lost. Instead, they are queued to disk and delivered once the receiver is ready. Of course, this is not optimal if the originator no longer needs the response, as is typical with web applications where a response is required quickly before users lose patience and either resubmit or terminate the sessions. Finally, enterprise messaging provides location transparency and decoupling of the service provider from the service consumer. This enables flexibility in the deployment architecture without application changes.

Business Rules and Recommended Practices

CMS takes a methodology-agnostic approach to software engineering. Irrespective of methodology applied, CMS prescribes some required practices. The following business rules are grouped by software development practice. Every practice is compatible with any development life cycle methodology, including Waterfall, Scrum, and Extreme Programming.

The provided rationale for each BR and RP should aid in understanding and tailoring application development within the CMS Processing Environments. Each practice area presents the BRs—the mandatory guidance—first, followed by the RPs. If the TRB promotes a recommended practice to a business rule, it will be listed as the next business rule in the series.

Table 2. CMS Application Development Practice Areas, Business Rules, and Recommended Practices (internal link) presents the organization of the BRs and RPs for this chapter.

Table 2. CMS Application Development Practice Areas, Business Rules, and Recommended Practices  (internal links)
Practice AreaBusiness RulesRecommended Practices
MethodologyBR-ADM-1, BR-ADM-2N/A
Software ArchitectureBR-SA-1 through BR-SA-10, BR-SA-14 through BR-SA-16RP-SA-11 through RP-SA-13
Software DesignBR-SD-1 through BR-SD-2RP-SD-3 through RP-SD-8
Software CodingBR-SC-1 through BR-SC-2RP-SC-3 through RP-SC-12
Software QualityBR-SQ-1 through BR-SQ-6RP-SQ-7 through RP-SQ-9
Secure SoftwareBR-SS-1 through BR-SS-7RP-SS-8
Engineering DocumentationBR-ED-1, BR-ED-2RP-ED-3
System MaintenanceN/ARP-SM-1 through RP-SM-3
Data and Database ManagementBR-DBM-1 through BR-DBM-3N/A
Software Configuration ManagementBR-SCM-1, BR-SCM-2, BR-SCM-4RP-SCM-3
Defect and Issue TrackingBR-DIT-1, BR-DIT-2RP-DIT-3
Software Build and IntegrationBR-SBI-1 through BR-SBI-3RP-SBI-4, RP-SBI-5
Packaging and DeliveryBR-PD-1 through BR-PD-3RP-PD-4 through RP-PD-6
DeploymentBR-D-1, BR-D-2RP-D-3 through RP-D-6
Release ManagementN/ARP-RM-01

Application Development Methodology

CMS has adopted the following Application Development Methodology (ADM) business rules.

BR-ADM-1: Use of the CMS Life Cycle Is Mandatory

The CMS Target Life Cycle (TLC) is required of all Information Technology projects, whether new or existing.

Rationale:

The TLC is the official life cycle for CMS IT projects, as required by the CIO. It supersedes the XLC.

Related CMS ARS Security Controls include: SA-5 - Information System Documentation.

BR-ADM-2: The Development Methodology and Artifacts Must Be Documented

The CMS TLC does not mandate the use a specific development methodology but requires that the methodology choice and all associated project artifacts to be produced, defined and documented. While the specific artifacts produced and their structure will vary based on the project requirements and the chosen development methodology, they should provide comprehensive, well organized coverage of key topics including:

  • Business Planning
    • Business need, alternatives, development options
    • Program governance
  • Architecture and Design
    • Solution architecture and interface control diagrams. This may be the System Design Document (SDD) as required by CFACTS
    • Relationship between the architecture and associated code components
    • Data archiving and Reporting
    • Test Plans and Reports
  • Software
    • Developed software code, including any configuration files, to support the installation and operations of the information system
  • Operations and Maintenance
    • Operational guide for the solution, including installation, failover and restoration guides

 

 PREFERRED CMS OIT provides resources that support IT Governance. TLC resources include:

 

Rationale:

Both waterfall and Agile are classes of methodologies, not specific methodologies. Documenting the development methodology makes it possible to set expectations for all parties regarding the process of development and the expected artifacts. It is important to know what to expect specifically from a project team to best ensure the project team addresses all required activities and that no activities are inadvertently overlooked.

Since most methodologies can be tailored, it is important to explicitly describe and share the tailoring approach with all relevant parties. For example, a project team using the XP methodology might document its adoption of test-driven development. Project managers could then expect writing unit test code before functional code. A project team using the SCRUM methodology would be expected to create and manage a product backlog.

Software Architecture

This topic addresses the business rules and recommended practices for software architecture (SA) and network architecture. A number of these items appear originally in some form in either the Foundation Network Services sections. They are cross-referenced and discussed here to elaborate on CMS requirements from the application development point of view.

BR-SA-1: Use CMS Shared Services

It is the software developer’s responsibility to research the available CMS Enterprise Shared Services, such as enterprise shared services and common platform services, as stated in the TRA Foundation Principles topic on Reuse.

At the time of publication, CMS Shared Services include:

  • Identity Management System (IDM)
  • Enterprise Portal
  • Master Data Management (MDM

System developers and maintainers must use applicable shared services. The business and technical rationale for any situations which require deviation from the use of Shared Services should be reviewed with the TRB during a Consult or Design session.

Rationale

Use of shared services reduces both data and code / logic redundancy and centralizes functions within the agency in accordance with the Federal IT Shared Services Strategy.

Shared services also help improve security because typical security issues have already been addressed and the services tested and used by others.

BR-SA-2: Integrate with the CMS Identity Management Services

All CMS Enterprise applications must use a CMS-approved identity management system. Any exception constitutes a security, operational, and architectural risk that should be reviewed by the TRB as part of a Consult or Design session.

Rationale:

The IDM shared service provides a single source of identity management within the CMS Enterprise. This shared service reduces user effort in switching between applications. The typical applications include the IDM Lightweight Directory Access Protocol (LDAP), IBM Resource Access Control Facility (RACF), or CMS Active Directory (AD). For specific implementation guidance please contact the TRB.

BR-SA-3: No Custom Application Code Is Permitted in the Presentation Zone

The CMS Presentation Zone, which houses edge services, supports static content for CMS applications. Under no circumstances will application services in the Presentation Zone write to persistent storage in the Presentation Zone (a) any data submitted by a client or (b) sensitive information passed to the client.

Commercial Off-the-Shelf (COTS) software packages are exempt from this restriction.

Please note that this rule does not include any code needed to support the configuration of edge resources (e.g. API gateway).

Rationale:

Static content such as Hypertext Markup Language (HTML) and JavaScript are allowed because they run in the browser, but files such as PHP: Hypertext Preprocessor (PHP) files and Java Server Pages (JSP) files are not permitted because they execute on servers in the Presentation Zone.

A Presentation Zone represents a so-called “Demilitarized Zone” (DMZ) and is Internet facing. Due to its role, the Presentation Zone represents a “less trusted” zone than the Application or Data Zones.

BR-SA-4: Use CMS-Validated Mediation and Data Access Services to Access Data in the Data Zone

Applications must use Data Access Services, integrating mediation principles, rather than access databases directly. The CMS standard is to design and implement data access services to abstract data sources. Mediation principles, implemented within Data Access Services, obfuscate the access requirements to the data source.

Thus, direct access from applications to databases, such as using Java Database Connectivity (JDBC) or Open Database Connectivity (ODBC), is not permitted.

If there is a need to deviate from this rule, the business owner, system owner, ISSO, and application developer are strongly encouraged to discuss the requirements and proposed architecture with the CMS TRB, including compensating controls and any analysis of alternatives. The CMS TRB can provide guidance, but the final decision on any deviation from this rule rests with the business owner, who must accept any associated risk.

Rationale:

There are several advantages to using this two-part system between the Application and Data Zones:

  1. Decoupling applications from data stores (location transparency).
  2. Potential scalability advantage, allowing for database “sharding.” (Sharding is horizontal partitioning of an application or database. Typically, horizontal partitioning of database instances will reduce the number of rows in any one instance. Each instance has the same schema but (potentially) different rows.
  3. Ability to perform maintenance on the data stores during outage windows, provided applications are using asynchronous queues and can tolerate a delayed response.
  4. Additional security provided by using a previously tested and trusted service rather than direct database access.
  5. Additional security because database access credentials are present only in the Data Zone.
  6. Additional security because data requests and responses can be validated and inspected.

Disadvantages include:

  1. Few COTS products include messaging support as an option instead of JDBC or ODBC, for example.
  2. Additional cost and complexity from creating data services and obtaining the mediation layer.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations.

BR-SA-5: No Long-Term, Persistent Sensitive Application Data Storage in the Presentation or Application Zones

Personally Identifiable Information (PII), Protected Health Information (PHI), or other sensitive data may not be stored in the Application Zone. Other non-sensitive data, such as reference data or caches, may be stored in the Application Zone to improve efficiency.

Public data may be stored indefinitely in the Application and Presentation Zones to provide better performance. In-memory caches are permitted in the Application Zone as a performance enhancement. These caches must be configured to purge expired data. CMS permits storing sensitive, PII, or PHI data temporarily in the Application Zone, but no longer than a maximum storage time of six (6) hours from the time the whole file or record is received for processing or for transfer to the Data Zone. Temporary files and cached data must be removed from the Application Zone once the transfer or processing is confirmed and successful.

This business rule also applies to message queues.

Rationale:

In keeping with the defense-in-depth strategy of the CMS TRA Multi-Zone Architecture, production and sensitive data should be persisted in the Data Zone. The TRB established the time limit for this rule.

BR-SA-6: Network Communications Must Meet the TRA Rules for Encryption

Please refer to TRA Network Services section Security Services chapter for the latest rules.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations.

BR-SA-7: Substantive Changes to the Architecture, Products, or Technology of an Existing Application Must Be Documented and Reviewed by the CMS TRB

Substantive changes to an existing application must be reviewed by the TRB to ensure compliance with CMS architectural standards and CMS and federal security standards. A Security Impact Analysis (SIA) must be conducted whenever the architecture, products, or technology of an existing system undergo substantive changes.

Rationale:

A substantive change may introduce security vulnerabilities and should be discussed with CMS ISPG and the TRB before committing to a course of action.

Related CMS ARS Security Controls include: RA-3 - Risk Assessment and CM-4 - Impact Analysis.

BR-SA-8: Logging Must Be Configurable and Use Common Platform Standards

Application logging capability must be configurable, allowing system operators to reconfigure the system for file-based logging, database logging, or network (UNIX® Syslog)-based logging.

Programming Language-specific guidance:

  • Java programs must use a logging framework such as Log4J, Apache Commons Logging, and the Java Logging API
  • .NET programs must use Log4Net compatible logging

Applications written in other programming languages should attempt to use Log4J-compatible output logs and configuration files if available.

The preferred logging format is the Common Log Format (CLF), although it is a good practice to coordinate this with the hosting provider and system operator. Note: If CLF is used, the ISO 8601 date requirement is not required because that format uses a different representation.

Rationale:

Configurability allows CMS flexibility in deployment.

Use of de facto and common file formats reduces the processing burden for system operations and security.

Related CMS ARS Security Controls include: AU-2 - Event Logging, AU-3 - Content of Audit Records, AU-5 - Response to Audit Logging Processing Failures, AU-7 - Audit Record Reduction and Report Generation, AU-8 - Time Stamps, AU-9 - Protection of Audit Information, AU-10 - Non-Repudiation (High), and AU-11 - Audit Record Retention.

BR-SA-9: Systems Must Define Metrics for IT Health Monitoring

It is critical to provide metrics relevant to the health of the overall system environment, whether cloud or data center based, to measure whether elements are operating as expected and to enable proactive response to emerging problems. Note that the metrics utilized in monitoring for IT performance are similar to health monitoring. Additional information regarding performance monitoring is available in the Infrastructure Services section Application Performance Monitoring chapter for further guidance.

Rationale:

To assess the IT health of an application, it is essential to evaluate various metrics. For data center environments, this typically includes metrics relating to CPU, network, storage, and memory usage as well as application and database service statistics. For cloud environments, while some of those same metrics may apply to virtualized systems, the distributed nature of cloud architectures and the use of cloud native services requires considering different metrics. These could include requests per minute, response duration, server/node availability, average compute and storage costs, latency, and others. Equally important is adhering to industry de facto and CMS IT standards for instrumentation and gathering such metrics using CMS’s monitoring infrastructure.

BR-SA-10: Applications in CMS Data Centers May Not Use Some Native Email Protocols

CMS prohibits the use of Messaging Application Programming Interface (MAPI) and Internet Mail Access Protocol (IMAP) protocols by business applications within a CMS data center or cloud enclave. Simple Mail Transport Protocol (SMTP) may be used only to connect to an Enterprise Email as a Service relay.

 PREFERRED

CMS recommends the use of CMS Enterprise Email as a Service for outbound mail. CMS Hybrid Cloud maintains SMTP relays that provide TLS-encrypted, authenticated mail connections from all CMS environments. Use of these is required by September 1, 2024.

To send email, applications must use message queuing or a web service to send a message through CMS internal SMTP relay servers. These relay servers send the outbound message via the CMS email services infrastructure (the Microsoft Exchange Web Services (EWS) protocol is permitted because it is Web Service-based).

CMS applications may directly receive email messages only if the email messages were addressed to a “cms.gov” or “hhs.gov” domain. Inbound email must be scanned for malware and checked for appropriate format prior to ingestion by downstream applications.

If there is a need to deviate from this rule, the business owner, system owner, ISSO, and application developer are strongly encouraged to discuss the requirements and proposed architecture with the CMS TRB, including compensating controls and any analysis of alternatives. The CMS TRB can provide guidance, but the final decision on any deviation from this rule rests with the business owner, who must accept any associated risk.

Rationale:

Official email from CMS must always have a “.gov” sending address. E-mail with “.gov” sending addresses may exit from a data center only via the hosting contractor’s designated, security-hardened email proxy servers, which must then forward all mail to an HHS trusted email server.

Use of SMTP in production data centers simplifies exfiltration of data by malicious agents or software. Use of SMTP, proxied by message-oriented middleware (message queues), renders this kind of exploit more difficult. It also provides a single point for auditing outbound SMTP traffic.

Projects that intend to use uncommon protocols must receive the latest guidance. Therefore, they must inform the TRB and ISPG team about such protocols during project design consultations.

Related CMS ARS Security Controls include: SC-8 - Transmission Confidentiality and Integrity.

Related National Institute of Standards and Technology (NIST) Special Publication (SP): SP 800-45 Revision 2, Guidelines on Electronic Mail Security.

RP-SA-11: Servers Should Include Instrumentation for Application Performance Monitoring

To provide useful information to business owners in their own terms, applications should provide application performance instrumentation. This instrumentation is responsible for providing metrics in business terms, specific to the given application monitored. Specifically, Java applications must leverage the Java Management Extensions (JMX) standard APIs, and .NET applications must leverage the Microsoft Windows standard APIs.

Rationale:

This business rule ensures that applications report information that is of interest at the business level; otherwise, monitoring might only report on such available generic infrastructure metrics that may not be vital to business owners. For example, applications can provide counts of users served, number of simultaneous logins, and other data. These are often the same metrics used in the formulation of Service Level Agreements (SLA).

CMS recommends that operational monitoring tools aggregate performance monitoring information rather than application-specific tools.

RP-SA-12: Minimize Manual File Copying by Using Integrating File Transfer Automation

Minimize manual copying, transfer, and extraction of data files even if it is less frequent (e.g., annual file transfer).

Rationale:

Eliminating manual steps improves quality by making processes more repeatable and improves security by limiting human access to critical systems and information while ensuring that secure practices are consistently applied. CMS provides an enterprise file transfer (EFT) shared service that can be used for this purpose.

RP-SA-13: Consider Data Services in the Data Zone to Improve Performance of Database-Intensive Services

Services that require a lot of data manipulation or direct access to databases should be developed as data services in the Data Zone.

Services in the Data Zone should abstract all data sources and repositories being accessed, and responses should only include data needed as part of the response. For more information please see CMS Services Framework (CMS Services Framework) and the CMS multi-zone architecture (CMS Multi-Zone Architecture).

Rationale:

Data services implemented in the Data Zone have potentially higher performance because they are “closer” to the data and can therefore benefit from reduced data transfer latency and fewer network hops to access databases. Because their responses only contain necessary information, data services improve Application Zone service performance by reducing bandwidth needed to send results.

BR-SA-14: Use of Short Message Service/ Multimedia Message Service by CMS Applications

Short Message Service (SMS) and Multimedia Message Service (MMS) must not be used for sensitive information, nor can a SMS or MMS source identity be trusted. When used by applications, SMS and MMS must be validated in the same way as email, ensuring against malware and use of proper input format.

Rationale:

SMS and MMS are not encrypted and the source of messages cannot be verified.

BR-SA-15: Protect Sensitive Information in Transit

All Personally Identifiable Information, Protected Health Information, or other sensitive data entering, exiting, or in transit within the data center (within or across zones) must be encrypted and secured according to the guidance in the CMS ARS. Applications must use Transport Layer Security (TLS) at the highest available level to exchange information securely. When possible, applications should use mutual authentication to ensure the identity of both parties in an information exchange. If encrypting sensitive information is not technically feasible or demonstrably affects the ability to support mission operations, compensatory controls must be implemented as part of a CIO approved risk acceptance plan.

Rationale:

CMS ARS SC-8 requires encryption of any transmitted data containing sensitive information to protect against unauthorized snooping of traffic. This control applies to both internal and external networks and all types of information system components from which information can be transmitted.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, CP-9 - CP9(8) - Cryptographic Protection, System Backup, MP-5 - Media Transport, SC-8 - Transmission Confidentiality and Integrity, SC-12 - Cryptographic Key Establishment and Management, SC-13 - Cryptographic Protection, AC-2 - Account Management, AC-3 - Access Enforcement, AC-5 - Separation of Duties, AC-6 - Least Privilege, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives, SI-7 - Software, Firmware, and Information Integrity, SI-10 - Information Input Validation, and AC-21 - Information Sharing.

BR-SA-16: Protect Sensitive Information at Rest

The application must protect the confidentiality and integrity of all sensitive-information (including all PHI or PII), according to the guidance in the CMS ARS. This includes using encryption that meets or exceeds the FIPS 140-2 encryption standard, utilizing an approved FIPS crypto module. The implemented level of encryption must be aligned to the sensitivity of the information. If encrypting sensitive information is not technically feasible or demonstrably affects the ability to support mission operations, compensatory controls must be implemented as part of a CIO approved risk acceptance plan.

Rationale:

Encryption protects sensitive information from unauthorized access and disclosure.

Related CMS ARS Security Controls include: MP-4 - Media Storage, SC-12 - Cryptographic Key Establishment and Management, SC-13 - Cryptographic Protection, AC-2 - Account Management, AC-3 - Access Enforcement, AC-5 - Separation of Duties, AC-6 - Least Privilege, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives,SI-7 - Software, Firmware, and Information Integrity, SI-10 - Information Input Validation, SC-1 - Policy and Procedures, and SC-28 - Protection of Information at Rest.

Software Design

CMS has identified the following software design (SD) business rules and recommended practices to guide application development.

BR-SD-1: External Configuration Is Mandatory

All configuration settings related to such components as network connections, ports, date, and Domain Name System (DNS) names must be stored outside the application code and not hardcoded into the application.

Rationale:

Inter- or intra-module configuration must be defined external to the application, such as in configuration files or databases, to:

  1. Assure that changes to the configuration can be performed without rebuilding the code.
  2. Allow code testing in different configurations without code modification.
  3. Deploy code in different environments with different hardware or software configurations.
  4. Ensure that hardcoding does not hamper horizontal scalability.

 

BR-SD-2: Web-Based User Interfaces Must Comply with TRA Guidance

The Web-based User Interface Services topics in this chapter establishes the business rules for designing web-based user interfaces, including mobile web interfaces.

Rationale:

Applications requiring web-based interfaces must follow existing guidance. This ensures that users have a consistent user experience across CMS web-based user interfaces.

Related CMS ARS Security Controls include: AC-19 - Access Control for Mobile Devices.

RP-SD-3: Configurations Should Be Validated and Checked on Each System Startup

When possible, the system operator should validate and check configurations at each system restart. CMS recognizes that some technologies (such as Spring dependency injection) make this difficult to enforce.

 PREFERRED

CMS testing tools SonarQube and Snyk (“sneak”) evaluate code against different languages and standards.

 

Rationale:

Validating configurations at startup prevents situations where an invalid setting could cause an abnormal system termination. These problems are avoidable by checking such configurations early in the startup and warning operators of configuration issues.

RP-SD-4: Consider Dependency Injection to Achieve External Configuration

Dependency injection, particularly for Microsoft .NET and Oracle Java-based applications, is recommended practice in industry as a method for easier application configuration and testing. Numerous Open Source and proprietary frameworks are available to make this an effective choice in application design.

Rationale:

Changing dependencies simplifies software testing by allowing reconfiguration without having to rebuild the software.

RP-SD-5: External System Dependencies Should Be Stubbed Out for Development and Testing

CMS recommends developing stubs to substitute for external system dependencies to increase system isolation during testing and decouple project timelines.

Rationale:

Inter-system integration can be difficult to set up because of complexity and availability of resources. Developers can save time by designing a set of stubs that allow development and testing to proceed. This can be as simple as “mock objects” that respond within the application in predefined ways or as sophisticated as stub servers that respond to inter-process service requests. See “xUnit Test Patterns: Refactoring Test Code” by Gerard Meszaros, 2007.

Every significant system has external dependencies. Within CMS, external dependencies are addressed via the Interface Control Documents (ICD) that describe the interfaces between a service consumer and provider. Given the ICD, it is possible to build a stub service as a substitute for the full system.

This practice can help with:

  1. Concurrent releases where the existing test system may not be as current as the stub.
  2. System availability in test, where a fully functional test system might not be available. By substituting the stub, development and testing can proceed.
  3. Faster testing cycles (and therefore performance) because a stub may be faster than the fully functional service.
  4. Simulation because the stub can be made to respond artificially slower or with specific data, thus allowing testing to occur in an environment of artificial scarcity.

Note: Integration testing typically would not use such stubs because they are contrary to the purpose of integration testing.

RP-SD-6: Timestamps Logged by the System Must Be in UTC or GMT and Should Be Expressed in ISO-8601 Format

Unless there is an overriding issue, developers shall use the ISO standard 8601 format for representing date and time stamps in all logs. The CMS ARS requires audit records that can be mapped to Universal Time Code (UTC) or Greenwich Mean Time (GMT) and accurate within thirty (30) seconds.

Rationale:

This practice makes dates easier to parse and reduces burden on log-scanning software.

Related CMS ARS Security Controls include: AU-8 - Time Stamps.

RP-SD-7: Software Should Be Designed Based on SOA Principles

Service should be autonomous and provide a complete unit of work. Multiple services can be assembled or composed into one service if need be.

The Web Services SOA Service Design Principles topic in this chapter provides additional information on how to structure and develop services for CMS using SOA principles.

Rationale:

Use of SOA principles encourages software and data reuse, consistent with CMS IT strategic objectives.

RP-SD-8: Consider Non-Blocking Service Implementations to Improve Performance and Scalability

The use of non-blocking service implementation technologies increases the scalability of services. Developers should consider the use of such designs, particularly if the services are simple and require high scalability. A response should always be provided back to the requestor.

Rationale:

Event-driven, non-blocking services reduce or eliminate the use of synchronization operations (such as semaphores and mutexes) in application code. As a result, they do not require the use of threads and have correspondingly lower memory utilization and higher scalability. CMS recommends using threads only when true CPU concurrency is needed. See “Why threads are a bad idea (for most purposes)”, John Ousterhout, Stanford U.

Software Coding

CMS does not mandate the use of specific programming languages. Irrespective of programming language, certain software coding (SC) best practices apply.

BR-SC-1: Inventory all Open Source Software Licenses

Every piece of open source software incorporated into the production release must be inventoried and the specific license documented. As new software is added or old software retired, the inventory must be kept up to date. The Open Source Software section also mandates this business rule.

Rationale:

Determining the currency and status of software licenses can be costly and time consuming. In addition, keeping OSS patched is important and requires knowledge of what OSS is in use.

The inventory list can be as simple as a text file that is stored along with the source code in the Version Control System (VCS) repository.

Related CMS ARS Security Controls include: SA-6.

BR-SC-2: All Custom-Written Source Code for a Project Must Conform to an Identified Coding Standard

A project must adopt a coding standard, document it, and adhere to it, as validated during code review or inspection.

CMS does not supply a specific coding standard. Industry offers many available options. Table 3 presents some commonly accepted coding standards that a project may adopt or adapt. The table below shows commonly adopted coding standards. For other languages, the Google Style Guide project provides a good starting point.

Table 3. Commonly Adopted Coding Standards
LanguageSuggested Standard
C / C++Ellemtel Standard
JavaGoogle Java Style Guide
JavaScriptGoogle TypeScript Style Guide
COBOLA.J. Marston’s COBOL Coding Standard
PythonGoogle Python Style Guide

It is recommended that projects adopt a tool for source code formatting and establish the coding standards in that tool to maintain consistency.

 PREFERRED - CMS strongly recommends the integration of SonarQube and Snyk with development environments to help enforce these standards and ensure code quality.

 

Rationale:

Having and applying a consistent standard makes both maintenance and code review easier.

Externally produced libraries and other forms of reusable capabilities (e.g., Open Source Software) do not have to meet these coding standards because they typically adhere to their own coding standards and are not produced specifically for a given project at CMS.

Note: Non-compliance with coding standards is a Defense Information Systems Agency (DISA) Applications Security and Development Security Technical Implementation Guide (STIG) Category II vulnerability.

BR-SC-3: All Custom-Written Code for a Project Must Be Shareable

The SHARE IT Act of 2024 requires that, with certain exceptions, any custom-written software that is created or modified after June 23, 2025, must be made available to other federal agencies.

Rationale:

This follows previous guidance from the Office of Management and Budget (OMB) to enable reuse. Technical guidance for reporting and sharing is produced by the Open Source Program Office (OSPO). Find more information in the Open Source Software section.

 

RP-SC-3: Do Not Intermingle Code in Different Programming Languages in the Same File

Many programming systems allow commingling of two or more programming languages in the same source file. This can lead, however, to difficulty in static analysis and issues in mixing content with presentation. As a result, CMS recommends against intermingling code in the same source file.

For example, CMS recommends separation of the following programming languages:

  • HTML from JavaScript
  • Java Server Pages from JavaScript
  • Java from SQL
  • HTML from Cascading Style Sheet

In cases where total separation is not possible, such as JSP, CMS advises adherence to the standard by striving for minimal inclusion of Java code in JSP files. A reference from the primary language to the secondary language is an example of minimal inclusion. Table 4 provides an example of file-type separation.

Table 4. File Type Separation Example
Correct UsageIncorrect Usage

<html>
    <body>
        <!--this is a reference to javascript -->
        <script src="/js/myscript.js"></script>
    </body>
</html>
                                

<html>
    <body>
        <script >
         // this is inline JavaScript
        alert("This is inline JS");
        </script>
    </body>
</html>                        

Embedded SQL (or ESQL) is one notable but also relatively rare exception. Like JSP, ESQL intentionally combines a host language, such as COBOL, C, or Java, with SQL to generate a new source file consisting only of the host language and database library calls to implement the ESQL logic.

Rationale:

To facilitate use of static analysis tools, profilers, and other source code scanning tools, code must not be intermingled. This also helps enforce other good practices like separating presentation from business logic.

RP-SC-2: Capture Code Metrics and Defect Tracking Metrics for Quality Improvement Purposes

Regular capture of such code metrics as size and complexity in the automated build process, as well as capture of defect quality metrics from the defect tracking system, allows for trending analysis and eventual quality improvement practices.

RP-SC-5: When Using Flat Files for Data Transfer, Include Helpful Metadata in the File

It is often helpful to embed such metadata as schema version, record count, version of the application that created the file, or other such information, directly into flat files prior to data transfer. This allows the file receiver to perform basic validation before using the file. If the metadata does not match expectations, the system operator should be notified and the notification logged.

One-time file transfers are exempted from this recommendation.

Rationale:

Flat file formats (schemas) often change over the lifetime of the file. By embedding metadata to the file, it becomes easier to determine what is needed to read and process the file correctly.

Rather than developing their own methods, developers should consider adopting the SemVer (semantic versioning) V2.0.0 proposal, which was designed to address these kinds of issues.

CMS grants an exemption for one-time file transfers because of the potential implementation cost.

RP-SC-6: When Using Flat Files for Data Transfer, Consider Including a Machine-Readable Schema

Sufficient metadata information should accompany each flat file to allow reading and validating the file. The typical limitation on “sufficient” metadata is the name of each field in the file, but may also include data types and lengths as well as field separators, if used.

CMS does not prescribe the machine-readable schema, which can take many forms. Table 5. Common Flat File Transfer Formats presents some common flat file transfer formats and the recommendations for their use.

Table 5. Common Flat File Transfer Formats
FormatRecommendation
CSVThe first row of the file should contain a comma-separated list of field names.
JSONEach Java Script Object Notation field has a field name and a value.
XMLThe file should be accompanied by an Extensible Markup Language Schema Definition. Alternatively, the data file may reference an XML Schema Definition (XSD) using the standard XML mechanisms. Some projects may choose to use Schematron to perform cross-schema validation.

Rationale:

Application of this recommendation makes it easier to read the file and effectively includes the ICD definitions with every file. It also makes the system more robust to changes in file formats because an application can detect and choose to stop gracefully or continue. For example, if a file were in an unexpected format, a graceful stop would provide a clear message to that effect to operators. Likewise, it could continue by ignoring unexpected data (if appropriate) and logging a warning. Either alternative is preferable to data corruption or system failure due to unexpected data. Irrespective of the representation, it is easier to exchange and maintain self-documenting data formats.

RP-SC-7: Use Decimal Math Types for Financial Calculations

Unless developers take extreme care when using floating point or integer math data types, it is generally preferable to use decimal math data types, such as BigDecimal in Java.

Rationale:

This practice reduces the likelihood of unintended round-off errors in financial calculations.

RP-SC-8: Consider Synthetic Transactions

Synthetic transactions are transactions to verify correct integration of a newly installed system with dependent services. These transactions also monitor “heartbeats” as well as throughput in production systems.

Synthetic transactions present at least two known perils: (1) to prevent misuse, it is crucial to design these transactions securely, and (2) they can skew monitoring statistics if issued in sufficient quantities.

CMS recommends disabling synthetic transactions during installation (default mode). During production, synthetic transactions must be conducted securely and should be configured to execute with a frequency that minimizes application load while still providing meaningful business reporting of application functionality.

The use of synthetic transactions in production must be approved by the application’s business owner.

Rationale:

Using synthetic transactions for “smoke testing” a new deployment can verify that everything is in working order because these transactions will fully exercise the system.

Artificial-Intelligence-Assisted Coding

For additional guidance on AI-Assisted Coding, see the TRA Application Development section AI Code Generation as well as Foundation Artificial Intelligence Guidance.

RP-SC-9: AI-Generated Code Should Undergo Human Review

All code generated by AI tools should be reviewed by at least two persons, including the person who requested the code generation. This review should verify:

  • Security controls and best practices are properly implemented
  • No sensitive information or internal implementation details are exposed
  • Generated code meets CMS coding standards and quality requirements
  • Input validation and error handling follow security best practices
  • All generated code dependencies are verified and approved
  • Code meets CMS accessibility and compliance requirements

Rationale:

AI-generated code may contain security vulnerabilities, expose sensitive information, or implement patterns that do not align with CMS standards. Multiple reviewers help catch potential issues that one reviewer might miss. Human review provides essential validation of security controls and code quality. When code includes references to external packages or dependencies, reviewers can verify their legitimacy and security posture.

RP-SC-10: AI-Generated Code Should Protect Sensitive Information

Teams should not share sensitive files and/or directories as context or prompts for AI coding assistants. This includes:

  • Credentials, secrets, or security tokens
  • Internal system architecture details
  • Business logic implementing fraud detection
  • Security control implementations

Rationale:

AI coding assistants may retain or use submitted information in ways that could expose sensitive data. Threat modeling has identified specific risks around prompt injection attacks and potential data leakage through context provided to the AI system (see A Practical Understanding of Threat Modeling for AI Systems). Even if prompts are filtered by the AI service provider, CMS must maintain control over what information is shared. This prevents accidental disclosure of protected information and maintains the confidentiality of CMS’s security implementations and architectural details. This aligns with CMS ARS controls for protecting sensitive information and system security information.

RP-SC-11: Configure Content Exclusions for AI-Generated Code

Teams should configure repository-level and organization-level content exclusions to prevent AI tools from accessing sensitive files and directories. Teams should also maintain documentation of which files and directories are excluded and the rationale for each exclusion.

Rationale:

Content exclusions provide a proactive mechanism to prevent AI tools from accessing and potentially exposing sensitive information. Even with provider-side filtering, organizations need to implement their own controls around what content is accessible to AI tools. This creates an additional layer of protection beyond relying solely on developer discretion or provider-side controls.

RP-SC-12: Implement Code Review Safeguards for AI-Generated Code

Teams should:

  • Enable branch protection rules requiring pull request reviews
  • Configure required status checks for security scans
  • Enforce signed commits
  • Limit merge capabilities to authorized team members
  • Label pull requests containing AI-generated code with an “AI-assisted” tag
  • Document the specific AI tool(s) used for code generation and indicate AI usage in pull request descriptions

Rationale:

These safeguards ensure proper review processes are followed and create transparency around the use of AI-assisted coding tools. Highlighted risks around malicious code being introduced through AI suggestions are areas of concern. Branch protection rules and required reviews help mitigate these risks. Labeling AI-generated code ensures reviewers can apply appropriate scrutiny. This is particularly important as AI-generated code may require different review considerations than human-written code.

Software Quality

The CMS Testing Framework is the source of all definitions of testing and test procedures such as unit tests, integration tests, and smoke tests.

CMS has adopted the following SQ assurance business rules and recommended practices for the CMS Processing Environments.

Testing custom software applications may require approaches such as static analysis, dynamic analysis, binary analysis, or a hybrid of the three approaches. Developers can employ these analysis approaches in a variety of tools (e.g., web-based application scanners, static analysis tools, and binary analyzers) and in source code reviews.

BR-SQ-1: All Custom-Written Software Must Have Associated Automated Unit Tests

Writing automated unit tests is an industry-accepted best practice, regardless of programming language used.

Rationale:

Automated unit testing frameworks are now available in most if not all commonly used programming languages. Consequently, software developers can express unit tests directly as programs or modules as appropriate. Such unit tests become an executable specification of a requirement that can be verified by inspection and by discussion with business subject matter experts (SME). The use of automated unit tests also encourages writing software that is modular and easy to test. This increases cohesion and decreases coupling, which are good goals for software and systems.

BR-SQ-2: Run Automated Unit Tests during Full Builds

CMS requires the execution of all automated unit tests during full builds of the entire source base. A full build does not use results from prior builds. It starts from a clean build area and results in deployable packages, test results, and in-line documentation.

Rationale:

Running tests during full builds is mandatory. Testing during full builds is typically done before a full release of the code; it is the last time a full set of unit tests can be run before release. This is a minimum requirement. Projects are encouraged to run unit tests for all builds because unit testing shortens the feedback loop between changes.

BR-SQ-3: Automated Unit Tests Must Use a Commercially Available Unit Testing Framework or Test Runner

Automated unit tests must use a commercially available unit testing framework or test runner that can be run from the command line and that can be executed from within a continuous integration server.

Rationale:

Industry has settled on two basic frameworks for unit testing:

There are xUnit frameworks available for nearly every programming language. Consequently, there is little reason to develop and maintain an xUnit framework specific to a project or system. If the programming language (such as COBOL) makes it difficult to use an xUnit-style framework, CMS will consider a proposed alternative, such as TAP. Choosing a protocol such as TAP makes it possible to provide easy-to-read and easily understood reports on test results.

BR-SQ-4: All CMS User Interfaces Must Meet Section 508 Accessibility Requirements

This business rule applies to web-based user interfaces but is not exclusive to web uses. The Web-based User Interfaces topic prescribes the applicable business rules.

Rationale:

CMS and HHS enforce compliance with Section 508 of the Rehabilitation Act.

BR-SQ-5: Manual Code and Design Reviews Are Mandatory

Someone other than the original author (or change author) must manually inspect all code and designs, and document the review results. Note: These code and design reviews are distinct from any TRB Consult or Design review sessions.

In addition to performing, recording, and documenting code and design reviews, the design team must provide notice of code reviews to CMS and must invite, at the business owner’s discretion, CMS auditors to attend the reviews.

Rationale:

Manual code and design reviews are very effective tools for raising software quality and identifying security weaknesses. Inviting CMS auditors supports the review process by providing the government an opportunity to assess the effectiveness of the reviews. Business owners have the discretion of foregoing such audits because of time commitment and potential scheduling conflicts.

CMS recommends considering automation in the form of code and design review tools.

BR-SQ-6: De-Identification of Production Data Is Required in Non-Production Environments

CMS requires using de-identification (see NIST SP 800-188, “De-Identifying Government Datasets: Techniques and Governance,” September 2023 and NIST SP 800-122 “Guide to Protecting the Confidentiality of Personally Identifiable Information (PII)”, April 2010) and data masking tools when the quality activities in non-production (sometimes called “lower environments” in CMS vernacular) use test data that originated as production data.

The Health Insurance Portability and Accountability Act (HIPAA) of 1996 requires de-identifying medical data. HIPAA specifies the list of Protected Health Information, 45 C.F.R. § 160.103.

Federal Tax Information (FTI) data is similarly subject to de-identification.

Exception:

CMS allows an exception to this rule if the non-production environments are configured and controlled to the same stringent security standard as the production environment and the non-production environment has received an Authorization to Operate (ATO).

Discussion:

Do not assume that fields can be de-identified without considering the context of the information because this could lead to re-identification, for example, by combining information with other sources. A structured analysis is recommended to determine the privacy risk associated with the original data, the requirements for the de-identified data to be useful, and the transformations that will be used to mitigate the privacy risk while preserving the necessary utility.

Rationale:

Although it is often helpful to test with realistic data, the federal policies related to the uses and disclosures of PII and PHI make it necessary to de-identify such information before use outside the production environment. As data moves from more trusted to less trusted environments (in reverse of the normal flow), it becomes necessary to perform tasks like de-identification.

RP-SQ-7: Code Coverage Analysis Is Highly Encouraged During Unit Testing

CMS recommends minimally seventy-five (75) percent code (statement) coverage.

Rationale:

Code coverage analysis gives an indication of how well unit tests are exercising the code in response to test data. Higher levels of code coverage give correspondingly higher confidence that the code is operating as intended. Note: One hundred percent statement coverage does not mean that every code path has been tested. Complete code path coverage testing is not computationally feasible for large programs.

RP-SQ-8: Use Static Analysis Tools During Build to Catch Common Coding Errors

CMS recommends using static analysis tools to check the bulk of files created by programmers. Good tool candidates include code quality analysis, coding standards, security vulnerability, and Section 508 analyses.

Rationale:

Ordinary compilers do not always enforce good practices nor do they discourage poor practices. Static analyzers can be configured to do both. In addition, their speed and consistency help diminish the burden of manual code inspection.

Static analyzers can also point out problem areas. For example, a tool producing the McCabe Cyclomatic Complexity number can identify areas of increased complexity relative to the overall complexity of the application. Higher complexity typically means greater risk of defects. Because these tools can quickly analyze a large amount of source code, they can help prioritize software quality improvement efforts. For more information about static code analysis, Please refer to references from CISA, OWASP and Wikipedia.

The DISA Applications Security and Development STIG recommends static security analysis in conjunction with manual review of the code.

Specific guidance about performing static code analysis within the CMS Cloud environment and integrating static code analysis into the DevOps CI/CD pipeline are available on the CMS Cloud DevOps website, including Snyk and SonarQube.

RP-SQ-9: Developers Assist Testers in Generating Test Data

Developers have intimate knowledge of the internals of their system; accordingly, they can assist in generating test data, providing detailed schemas of valid data, or providing data generators to the testing team.

Rationale:

Testing teams often do not have the necessary programming skills to perform automated test data generation. Developers can assist by providing needed skills while testers select what and how to test.

Secure Software Practices

The SANS Institute and The MITRE Corporation publish a list of Top 25 Most Dangerous Software Errors. As stated on the Common Weakness Enumeration (CWE™) website, “The CWE site contains data on more than 800 programming errors, design errors, and architecture errors that can lead to exploitable vulnerabilities.” CMS has established the following business rules for secure software (SS) practices.

Related CMS ARS Security Controls (for the following business rules in this topic) include: SA-8 - Security and Privacy Engineering Principles and SI-2 -Flaw Remediation .

BR-SS-1: All Software on CMS Production Servers Must Have Recorded Provenance

All software installed on production servers must come from known and documented sources, and must be auditable. The only allowable sources for software on a CMS production system are CMS-controlled media, build server, or repositories managed by the system operator.

Rationale:

Software installed from unknown sources constitutes a project and security risk. The use of a controlled build pipeline and controlled repositories reduces or eliminates this risk by recording every piece of software installed. The appearance of unexpected code could indicate malware infection.

Related CMS ARS Security Controls include: CM-2 - Baseline Configuration and CM-3 - Configuration Change Control.

BR-SS-2: Use NIST SP 800-132-Specified Password-Based Key Derivation Functions (PBKDFs)

When passwords must be stored in a CMS database, they must be stored only as a salted cryptographic digest (sometimes called a hash), computed by a security function approved for applicability with Federal Information Processing Standards Publication (FIPS PUB) 140-2, and following key derivation techniques specified in NIST SP 800-132. The current recommended algorithm is PBKDF2, specified in IETF RFC8018 This is typically implemented as part of secrets management, such as AWS Secrets Manager. See Keys and Secrets Management for more information.

A Password-Based Key Derivation Function (PBKDF) that uses the SHA-256 (or better) hashing algorithm and a randomly generated 128-bit salt is recommended for CMS information systems. An iteration count of at least 1,000 is also recommended.

Rationale:

Passwords should almost never appear on CMS systems. They should be managed using a CMS-approved identity management solution. This business rule applies in those cases where password management is local to the application.

A salted cryptographic hash using a large number of iterations is the current best practice for storing passwords. This approach increases computational cost associated with each derivation and helps thwart dictionary or brute force attacks.

Passwords encrypted in this manner are not to be decrypted by the CMS system, but rather are compared for equality in their encrypted form.

Related CMS ARS Security Controls include: IA-5 - Authenticator Management.

Related: NIST Special Publication SP 800-132, Recommendation for Password-based Key Derivation Part 1: Storage Applications. Password-Based Cryptography Specification Version 2.1 (PBKDF2), specified in IETF RFC8018, the OWASP Password Cheat Sheet.

BR-SS-3: SQL Code Must Use Binding Variables

Use of SQL binding variables reduces an application’s exposure to SQL injection attacks. This technique is available in most major brands of relational databases although the exact syntax may vary by product.

Rationale:

All modern implementations of SQL provide the capability to define binding variables to pass data safely back and forth between SQL statements and program code.

Building Dynamic SQL code via string manipulation may introduce security vulnerabilities, and especially to SQL injection. Simply using stored procedures does not offer sufficient protection. SQL statements that include user-provided data must use binding variables except when prohibited by the language. Only a few SQL statements disallow binding variables (such as DROP TABLE). In these cases, it is important to avoid user-provided data or use a robust input validation system to avoid SQL injection vulnerabilities.

Related CMS ARS Security Controls include: AC-19 - Access Control for Mobile Devices.

BR-SS-4: Check for Common Security Vulnerabilities

Projects must check for common security vulnerabilities in their code using a combination of testing and analysis tools.

CMS relies on The MITRE Corporation Common Weakness Enumeration (CWE™) and the CWE™ Top 25 list of Most Dangerous Software Weaknesses to define the types of vulnerabilities to check for. This includes the following:

  • Cross-site scripting (XSS)
  • SQL injection
  • JavaScript injection

The party responsible for software assurance must scan all custom-written software for vulnerabilities using the CMS ISPG-approved scanning software.

Rationale:

Penetration testing alone is insufficient to comply with this business rule; rather, a combination of manual code review, static and dynamic testing, and penetration testing should be leveraged.

Use of a Web Application Firewall (WAF) or other tool does not obviate testing for security vulnerabilities.

CMS ARS requires that the source code be free of known vulnerabilities; system developers must test for weaknesses throughout the development process.

Related CMS ARS Security Controls include: RA-5 - Vulnerability Monitoring and Scanning.

BR-SS-5: Use Static Analysis Tools to Catch Common Security Weaknesses

This business rule supports BR-SS-4. Static analysis tools must be used to check the bulk of files created by programmers for known security weaknesses and vulnerabilities.

Rationale:

Although manual inspection can find security weaknesses, static analysis tools, such as HP Fortify, University of Maryland FindBugs, or Open Source Splint, can help detect security problems and reduce the burden of manual inspection. CISA provides a list of Free Cybersecurity Services and Tools. The DISA Applications Security and Development STIG recommends static security analysis in conjunction with manual review of the code.

 PREFERRED

CMS recommends the use of SonarQube for static code analysis.

RP-SS-6: Use Profiling to Perform Dynamic Code Analysis

CMS recommends that project teams perform Dynamic Code Analysis (also called profiling) to investigate the application’s behavior at runtime (unlike static analysis, which only uses source code).

Rationale:

Dynamic analysis provides a runtime picture of the application’s behavior, including memory consumption, file and database access, and network usage. This information is helpful in determining application performance and security characteristics, grounded in observation of behavior. Dynamic analysis is also helpful in identifying security vulnerabilities. Projects should consider the HHS AppScan service for this purpose.

BR-SS-7: Error Handling Must Not Reveal Information That Could Lead to an Exploit

In production operation, systems must not produce error messages that reveal information that could be used to maliciously compromise or otherwise exploit the system. Examples include providing internal ID numbers, database metadata, and other such messages.

Rationale:

Some web-based systems have default, development-mode configurations that reveal information about the processing to help developers easily identify problems. Unfortunately, these settings also provide information valuable to attackers. As stated in the CMS ARS Security Control SI-11, “organizations [must] carefully consider the structure / content of error messages.”

Related CMS ARS Security Controls include: SI-11 - Error Handling.

RP-SS-8: Perform Threat Modeling During the Design Phase to Identify Potential System Threats

CMS recommends that project teams perform Threat Modeling as early as possible in the System Development Life Cycle (SDLC), with updates as needed throughout the SDLC. Ideally during the Design / Requirements phases, or if following an Agile / DevOps approach, during each Sprint planning meeting. This practice promotes early identification and remediation of vulnerabilities, as well as the continuous monitoring of effects from internal or external changes.

Rationale:

It is insufficient to reactively respond to discovered software security issues and vulnerabilities. Threat modeling allows ADO Teams to identify security risks and vulnerabilities early in the system development life cycle (SDLC). By analyzing potential threats during the design phase, an ADO Team can address them proactively, reducing the chances of security issues cropping up later. For existing products and projects, a threat model can help validate decisions made around secure design, as well as inform improvements to the secure design of system. Threat Modeling is addressed further in Application Development Principles.

Engineering Documentation

The primary purpose of engineering documentation is to record the information used during software development. It is not end-user documentation. CMS has established the following business rules and recommended practices for engineering documentation (ED).

BR-ED-1: Custom-Written Software Must Include Inline Documentation for Public APIs

A public API (or Web Service) is a constant or variable method, procedure, or function that is accessible from another module or system. Public APIs must be documented. In addition, distributed processing APIs, such as remote procedure calls, REST APIs, Simple Object Access Protocol (SOAP) calls, or other APIs, must be documented.

The software build procedure must generate human-readable documentation that is regularly published to a project internal site or folder for reference by developers and maintainers.

Rationale:

The documentation of public APIs must go beyond providing function signatures because function signatures alone do not fully specify operational assumptions. For example, any occurrence of a change to system state (such as updating a database record) is not specified in the function signature. It is therefore important to provide documentation that is trustworthy and accurate. Inline documentation is an industry best practice that has proven itself valuable in documenting such APIs.

Inline documentation is consistent with the use of such tools as JavaDocs, doxygen, and robodoc. Other similar tools are available for all popular programming languages.

BR-ED-2: The CMS TLC Phase Review Artifacts Must Be Produced

Regardless of software development methodology employed, the CMS required project artifacts must be produced to support TLC phase reviews as well as any required TRB design consultations. In addition, business and program teams must be able to provide the documents/artifacts that support their system within 2 business days, upon request, to fulfill review or audit requests from outside agencies such as OIG and GAO.

Related CMS ARS Security Controls include: SA-3 - System Development Life Cycle.

Rationale:

Typically, CMS systems have long operational lifespans. It therefore is necessary to have sufficient, accurate documentation to support the initial deployment as well as the full operations and maintenance life cycle. Refer to the CMS TLC website for additional information regarding the artifacts expected within each phase.

RP-ED-3: Engineering Documentation Should Be Versioned Along with Source Code in the Same Repository

All engineering documentation such as design documents and diagrams should be versioned along with source code in the same repository. Such documentation must be stored in source (modifiable) form, not just distributable form (such as Adobe PDF).

Rationale:

This versioning approach makes it possible to keep both kinds of information (source code and documentation) up to date and baselined together. Modifiable documentation is necessary to support continuous maintenance.

Note: Some version control systems are incapable of storing binary data. Merging is a difficult activity for complex binary data even if the version control system can store binary data. In these cases, it may be preferable to store such media assets in a media asset management system or web content management system.

System Maintenance

Design for maintenance recognizes that successful systems have long operational lives. As a result, CMS advocates building capabilities for diagnosis, debugging, and health monitoring into systems. CMS has adopted the following recommended practices for system maintenance (SM).

RP-SM-1: Consider Building Self-Diagnosis Capability into Systems

The ability to diagnose system problems rapidly and accurately can be very helpful. CMS recommends that developers consider building in diagnosis tools to perform pre- and post-run validation of the system.

Rationale:

Writing a simple script for scanning logs or configuration for certain kinds of easily correctable errors can be a timesaver. Another helpful heuristic is counting errors and halting processing when a specific threshold is exceeded.

RP-SM-2: Consider Designing Maintenance Capability into Systems

Developers should consider the following system maintenance capability requirements adapted from NASA Johnson Space Center’s Man-System Integration Standards, Volume I, Section 12: Design for Maintainability (internal link):

  • Physical access, visual access, removal, replacement, and modularity requirements
  • Fault detection and isolation requirements
  • Test point design
  • Maintenance data management system

Developers should consider the following factors:

  • Non-interference of preventive maintenance
  • Flexible, preventive maintenance schedule
  • Reduce training requirement for system operations
  • Reduce skill requirements for system operations
  • Reduce time spent on preventive and corrective maintenance
  • Increase maintenance capabilities (especially corrective) during mission
  • Decrease probability of damage to part, module, product, or data itself

Maintenance capability techniques include:

  • Mistake proofing (to ensure that a part or module can only be installed correctly)
  • Self-diagnostic indicators, gauges, annunciators, pop-up dialogs, and dashboard indicators
  • No or minimal adjustment (self-adjustment as well)
  • Tracking metrics over time to identify problem areas

Lack of access to production servers by developers can hamper efforts to determine root cause and perform troubleshooting. CMS recommends that software designs include such troubleshooting capabilities as logs, variable dumps, execution traces, or other techniques.

Note: Debugging data may include PHI or PII and must be secured.

Data and Database Management

Data and database management (DBM) are critical parts of CMS systems. Flexibility in data and data management can improve the efficiency of quality assurance activities by allowing for more and varied access to alternate data sources. For example, quickly switching between databases can make it possible to prepare and use test data.

Software developers must adhere to the following business rules to ensure CMS systems are more robust in the face of change. In addition, CMS requires compliance with specific Data and Database Management Standards (please refer to BR -DBM-3).

BR-DBM-1: Systems Must Meet Federal Record Management Requirements

CMS mandates compliance with National Archives and Records Administration (NARA) requirements for federal government records management.

Rationale:

Federal Record Management Requirements are complex and may have design and operational impact. For more information, please refer to National Archives and Records Administration (NARA).

BR-DBM-2: Systems Must Meet Federal Government FOIA Requirements

CMS directs compliance with the Freedom of Information Act (FOIA).

Rationale:

To meet FOIA requirements while reducing the burden on the agency, it may be necessary to consider data tagging or other techniques to flag data in databases for eligibility or ineligibility for FOIA. For specific FOIA requirements, please refer to FOIA.

BR-DBM-3: Systems Must Meet CMS Data and Database Management Standards

CMS has published the following TRA chapters that establish the Agency’s guidance and standards on data and database management:

Data Architecture

For such information as data design patterns, standard terms, naming and definition standards, modeling tools and resources, and to reference the CMS Data Reference Model (DRM), a data taxonomy that describes data and subject areas fundamental to achieving CMS’s mission, please refer to the ​​OIT Division of Enterprise Architecture (DEA) Data Architecture & Engineering Services site.

The documents on the DA page are kept up to date through continuous improvements and new guidelines are added. Feedback on DA publications can be shared through the DA mailbox:

 

Standards and Guidelines Documents

The Data Architecture team develops various standards and guideline documents around data design, architecture and technology which include, but are not limited to the following topics: Data Naming, Data Definitions, Data Domains, Data Assets, Data Dictionary, Business Glossary, Data Catalog, and data modeling tools. DA’s standards and guidelines are created in part by following industry standards like ISO 1179, Data Catalog Vocabulary (DCAT), and Dublin Core.

CMS Standard Terms

The DA team maintains a growing list of data terms that are used by CMS project teams, the CMS Standard Terms List (STL). The list is made available for download and kept up to date on the DA page. Data names in Logical Data Models should be composed of one or more terms in the STL and adhere to its naming conventions. Projects that use this set of terminology for their data names also provide a common understanding to the rest of the CMS data community making it easier for the enterprise as a whole to function seamlessly. Project teams may request a change or addition to the list by filling out the Standard Term Request Form, available on the DA page, and submitting it to the DA Mailbox.

DA Consultations and Data Model Reviews

To check if a database design meets CMS standards, the project team should contact the DA team within the Division of Enterprise Architecture (DEA) at the DA Mailbox to validate their data artifacts including logical data models, data dictionaries, or other artifacts. Project teams may also set up consultations with the DA team for general data design and architecture advice as well as potential solution engagements to explore new data management technologies. To submit a request for consultation, please reference instructions on the DA page.

Rationale.

CMS’s standards aim to promote better interoperability among applications by ensuring data is clearly, simply and consistently described (named and defined) for consumers.

Database Administration

For information on the roles and responsibilities of a DBA at CMS and for guidelines on commonly used database platforms such as SQL Server, Oracle, and DB2 refer to the following publication from the Database administration:

Rationale.

CMS keeps separate roles and responsibilities for Central DBAs and Local DBAs. The Central DBA will have final approval for all database objects running on all database servers. The Local DBA will refer to the day-to-day operational support person responsible for activities necessary to implement and maintain the database for a project.

Software Configuration Management

CMS requires software configuration management (SCM) in accordance with CMS Risk Management Handbook: Configuration Management (CM), which provides additional details.

Related CMS ARS Security Controls include: SA-10 - Developer Configuration Management, CM-1 - Configuration Management Policy and Procedures, and CM-2 - Baseline Configuration.

BR-SCM-1: All Source Code Must Be Checked in to Version Control

The source code is a valuable asset of the system and must be maintained via version control. CMS does not specify a single enterprise repository; instead, each project is chartered to maintain its own repository. The SHARE IT Act of 2024 requires the sharing of custom-developed code.

Source code includes all build scripts, test scripts, test data generators, packaging specifications, Open Source Software, Infrastructure as Code (IAC) configurations, and any other files used as dependencies for the build and packing process.

 PREFERRED

CMS strongly recommends the use of its enterprise GitHub repository services.

 

Rationale:

Version control is necessary from the standpoints of asset management and security.

Open source Software comes without a manufacturer’s warranty. Thus, it is necessary to maintain the original source under version control. This facilitates testing and maintenance activities that often require source code to fully diagnose and correct issues.

Related CMS ARS Security Controls include: SA-10 - Developer Configuration Management and CM-2 - Baseline Configuration.

BR-SCM-2: All Code Must Be Baselined Prior to Release into Implementation, Validation, and ATO(ed) Production Environments

All version control systems have mechanisms for baselining a release of software. Whether this is called baselining, labeling, or tagging, the effect is that a set of files, each at their own revision level, are identified as belonging to a named baseline. CMS does not specify any naming convention for baselines.

The baseline manifest must contain a summary of the errors, change requests (CR), new features, and other changes that differ from the prior release.

Rationale:

Identifying a baseline is a key step in configuration management. It is necessary to know what release of code was used to build software in test and production. Baselining is accomplished differently in every version control tool, but the effect is the same—to label or tag a specific level of code.

Related CMS ARS Security Controls include: SA-10 - Developer Configuration Management, CM-2 - Baseline Configuration, and CM-3 - Configuration Change Control.

RP-SCM-3: Apply Database-Oriented Configuration Management Practices

Use database-oriented configuration management practices to ensure that changes to the database schemas are synchronized with the schemas expected by program source code.

Rationale:

Without database-oriented configuration management practices, it is easy for database schemas to get out of sync with the source code intended to manipulate them. Databases evolve by applying changes to an existing database, but deployed code evolves by installing new code each time.

BR-SCM-4: Configurations Must Be Checked in to Version Control

Configurations stored in textual form (except for encryption keys and passwords) must be checked in to version control to facilitate system configuration management. This includes any Infrastructure as Code (IaC) implementations, such as Terraform, AWS CloudFormation.or Kion (Cloudtamer), as well as cloud rules and policies.

Rationale:

Controlling the versions of configurations makes auditing possible and reduces the risk of unexpected configurations.

Related CMS ARS Security Controls include: CM-6 - Configuration Settings and CM-6(1) - Automated Management, Application, and Verification.

Defect and Issue Tracking

To maintain quality, it is essential that projects track defects and change requests in a defect tracking system. The following business rules and recommended practices govern defect and issue tracking (DIT) on CMS projects.

BR-DIT-1: All CMS Software Development Projects Must Use a Defect Tracking System

Each software system at CMS must use a defect tracking system to track and manage defects. System maintainers must select either a COTS package or, in the alternative, program-specific custom solutions. COTS or custom solutions are acceptable if the full defect and issue history (along with all attachments) can be easily exported and transferred to subsequent contractors.

 PREFERRED

CMS strongly recommends the use of the Enterprise Jira Bug Workflow (internal link).

 

Rationale:

The defect history is essential for understanding the evolution of a software system.

Related CMS ARS Security Controls include: SI-2 - Flaw Remediation.

BR-DIT-2: A Defined Defect Classification Standard Is Mandatory

Projects must define and document their classification scheme for defects.

Rationale:

Without a common standard, it is not possible to gather the necessary metrics to understand the evolution of the software.

The following four classification standards are available:

In addition, a software system maintainer (organization) may propose an alternative classification standard.

RP-DIT-3: Defects Should Be Correlated to Baselines

Baselines, sometimes called commits or tags, should be tracked in defect reports. To fix a defect, it is necessary to track what was changed.

Rationale:

Baselines help release managers record the change history and understand in business terms what is changing.

Software Build and Integration

Software build and integration (SBI) is the process of converting source code into the target representation. In this set of disciplines, code produced by many developers is integrated into a single build and the output is a set of packages for installation. CMS has established the following business rules and recommended practices for SBI.

BR-SBI-1: All Builds Must Occur in Controlled Environments

Build servers must be isolated from external sources of source code such as developers’ computers, with two exceptions: Package or Library servers, and authoritative version control systems.

In addition, CMS requires identification and inventory of the software tools used in generation of code, such as compilers, case tools, etc. At a minimum, this should be documented within the project artifacts. In some environments, this information can also be recorded in configuration files stored under version control.

Build control files, such as Makefiles, Apache Maven POM (Project Object Model), or Apache ANT build.xml files, must be obtained from version control prior to building.

Rationale:

All software on a production server must be auditable. Inclusion of code, configuration, or other uncontrolled data in a build constitutes an avoidable security vulnerability.

Related CMS ARS Security Controls include: SI-7 - Software, Firmware, and Information Integrity.

BR-SBI-2: All Production-Deployed Custom Code Must Be Built and Installed from Version-Controlled Source Code

All custom-written sources must come from a version control system.

Rationale:

All software on a production server must be auditable.

Installing or modifying code in production without first checking it in to a version control tool constitutes a project risk and security vulnerability.

BR-SBI-3: Production Builds Must Have Zero Compile Errors

Production builds must not have known compile-time errors (excluding warnings or informational notices). Compliance with this business rule does not permit suppressing errors.

Rationale:

Compilation errors constitute a project technical risk. Many times, the compiler will not produce an output. In other cases, the compiler will make a best guess. Either outcome represents an avoidable technical risk.

CMS strongly recommends that production code have zero compile-time warnings, and does not suppress warnings through compiler options and flags. Because different compilers have different levels of tolerance to errors, it is not possible to issue a single business rule on the subject.

RP-SBI-4: Use Explicit Library and Build Dependency Management

Packages typically have dependencies on other software packages or libraries to conduct builds. All software library or package dependencies should be documented. The documentation may be a human-readable document, or preferably, a machine parse-able specification. The documentation should be checked into version control.

A corollary to this rule is that production builds must use explicit dependencies and not automatically upgrade to the latest versions of packages unless the project has planned for a full regression test. A system checkout report should be produced to ensure that the inventory of packages is as expected and to flag any combinations of packages that have known (declared) incompatibilities. This is, of course, package system dependent.

Rationale:

Library and build dependency management is an essential part of configuration management for builds. Modern build tools such as Apache Ivy (part of Apache ANT), Apache Maven, and Gradle all use package dependency management, which makes it possible to trace the exact versions used to build a baseline release of code.

RP-SBI-5: Consider Instituting Continuous Integration

Continuous integration is the practice of continuously checking out source code, building it to ensure that all parts integrate, and then running automated unit and integration tests.

Rationale:

Continuous integration provides an insight into product quality, allowing early identification of issues (particularly subsystem / module integration issues) in the development cycle. Continuous Integration also provides an easy place to introduce quality improvement activities, such as automated regression tests or static analysis.

Packaging and Delivery

Packaging is the process of bundling related object files into an archive suitable for installation into a system. A target release may constitute one or more target packages, each of which may be installed on a potentially different system. The target release will provide the recommended order of installing these packages on the target systems and should also include a copy of release notes.

A source release is the set of source code along with a list of commercial and custom software required to build the software. Source releases should also include a copy of release notes.

Delivery places the packages in a repository for later deployment. The package catalog on each server tracks the software currently deployed on that server. This is the normal operating mode for distributed systems. Mainframes may follow different conventions for managing the installed software assets inventory.

CMS has established the following business rules and recommended practices for packaging and delivery (PD).

BR-PD-1: Software Must Be Packaged for Deployment

A package includes a manifest listing of the following:

  1. All object files, configuration files, and other files necessary for operation along with installation and de-installation instructions either in machine-readable or human-readable form. Machine-readable form is preferred.
  2. Complete inventory of any third-party components with version numbers.
  3. The rest of the package is the software itself along with instructions for installing, updating, or removing the software.

Packaged software must be installable via a single command, in a standard package formats, for each supported target platform and operating system.

Rationale:

Packaging software allows system operators to install packages, which also update operating system or language system catalogs for easier configuration management. This helps ensure that systems are operating with the expected software, which reduces both operational and security risks.

Single step installation and removal is key to reducing otherwise error-prone installation procedures.

Related CMS ARS Security Controls include: SI-3 - Malicious Code Protection.

BR-PD-2: Software Target Packaging Must Be in Either the Operating System or Language Platform Native Form

Operating System (OS) native installation packaging includes the following examples as shown in Table 6. Packaging System by Platform (Illustrative not Normative).

Table 6. Packaging System by Platform (Illustrative not Normative)
Operating SystemMechanism
Red Hat Enterprise Linux or CentOSRPM, which is installed with YUM
SUSE Enterprise LinuxRPM, which is installed with Zypper
Debian LinuxDEB, which is installed with APT
WindowsMSI, which is installed manually or via Microsoft System Center Configuration Manager or equivalent
Oracle / Sun SolarisSun Packaging system
IBM Mainframe platform (z/OS)At CMS, Endevor is used to perform installation using package control.

Installation specifications may include installation and de-installation code. Mobile platforms may have different installers based on manufacturer.

Some language platforms include packaging standards. These are acceptable as a packaging mechanism for CMS. Table 7. Language-Specific Packaging Standards presents recognized language-specific platform packaging standards.

Table 7. Language-Specific Packaging Standards
Programming LanguagePackaging Format
RubyGems
JavaJAR, WAR, EAR
PythonPIP
Node.JSNode Package Manager (NPM)
Microsoft .NETNuGet Packages
PerlPerl Libraries (CPAN)

The packaging formats in Table 7 can produce an inventory on demand of the software installed on a system.

Note: Syntactic units that are part of certain programming languages are sometimes called packages (such as Ada, Java, or PL/SQL). These do not constitute “packaging” based on this definition. They provide modularity, but do not specify a binary release package format.

BR-PD-3: Database Changes Must Include Back-Out Scripts

When packaging database scripts will change the Data Definition Language (DDL) or include Data Modification Language (DML), back-out scripts must be provided in the event the script must be rolled back and the database state restored.

Rationale:

The capability to restore the database to the pre-release state is an essential risk mitigation technique.

RP-PD-4: The Package Manifest Should Include a List of All Defects Corrected in the Release

Packages should include a list of all defects corrected. There should be a defect number for each defect and a one-line summary or abstract. Any known and uncorrected defects related to a package should be identified in the same manner.

Rationale:

This list of defects, which helps operations understand the impact of changes, can be generated automatically from data in the defect tracking system and the version control system.

RP-PD-5: Changes Applied to Databases Should Be Recorded in the Database Itself

It is often necessary to apply changes to a database that alter the schema or data. To enable changes in idempotent fashion, it is helpful to record which changes have been applied. This record prevents double application of a change. It also facilitates auditing changes applied to a database, which is useful during operations. Please refer to The Agile Data (AD) Method for other strategies.

RP-PD-6: Support A/B Testing of User Interfaces

Providing feature flags or other kinds of runtime configuration supports A/B testing and allows different groups of users to experience slightly different versions of a web site. This is particularly useful for large-volume web sites where the user experience is crucial to the success of the program.

Rationale:

A/B testing is a technique that allows CMS to offer a better user experience by scientifically testing user reactions to two or more different variants of the web site.

If A/B testing is valuable for an application, certain design changes can be applied to support A/B testing and should be considered at the outset.

Deployment

Deployment is the activity of installing or upgrading an environment with installation packages from a trusted repository. Deployment may require taking a server down or placing it into maintenance mode, when it may be either unavailable or can only proceed with limited availability (for example, read-only mode). CMS has established the following business rules and recommended practices for deployment (D).

BR-D-1: Developers Do Not Have Unsupervised Administrative Access to Production Servers

Developers must not have unsupervised administrative access to production servers. This requirement is enforceable, for example, by having separate operations staff to access production servers or by pairing developers with operations or management staff to supervise access to production servers.

Rationale:

For reasons of separation of duties (CMS ARS Security Control AC-5) and Least Privilege (CMS ARS Security Control AC-6), developers typically do not have access to production systems.

Related CMS ARS Security Controls include: AC-2 - Account Management, AC-5 - Separation of Duties, AC-6 - Least Privilege, and CM-5 - Access Restrictions for Change.

BR-D-2: All Installation and Back-Out Scripts Must Have Been Tested in Lower Environments Prior to Use in Production

All installation, upgrade, removal, and back-out scripts must have been tested (as with any custom software) in lower environments (development, validation, or integration) prior to use in production.

Rationale:

Any software run in production must have been tested in lower environments in accordance with the Release Management guidance in this volume. Otherwise, the probability is great that the back-out scripts will not work when used in production.

RP-D-3: Support Rolling Deployment

One recommended practice is rolling deployments. In a rolling deployment, some portion of the web site is removed from the load balancer, brought off line, updated, brought online, and reintroduced to the load balancer to repeat the process on a different portion of the web site until the entire site has been migrated to the new code. This technique can be combined with feature flags, a configuration setting that activates a new feature across all new sites or a portion as needed.

Rationale:

A rolling deployment allows the performance of system upgrades with reduced or no outage because both the old and new system are available for use during the transition to the new system.

RP-D-4: Use Feature Flags to Gradually Introduce New Features to Users

Feature flags are runtime flags that allow a business capability to be turned on or off at runtime in production. They allow for decoupling the time of deployment from the time of use.

Rationale:

Feature flags provide controls for the business to decide which features to expose to users. Flags enable features, as well as easy removal of a change that may cause issues. With sufficient intelligence, a feature flag can also enable A/B testing, allowing exposure of different user groups to different capabilities to gauge the relative difference in adoption or sentiment by the user groups.

RP-D-5: Deployment Should Integrate with Monitoring to Coordinate Outages

Because deployment causes some systems to be taken offline, deployments may trigger alarms in the monitoring system. This is avoidable by informing the monitoring system that the deployment action is part of a deliberate, planned outage.

Rationale:

It is preferable to coordinate any planned outages, such as deployment or maintenance activities that could take part of or an entire system offline. Integration with monitoring reduces coordination errors and is more efficient than human coordination techniques, such as telephone calls.

RP-D-6: Support Rollback of Package Installation

Every package designed for installation into production should be capable of rollback. Rollback can entail database changes, such as schema changes.

Rationale:

It is sometimes necessary to back-out changes to the production system. If a pac kage changes a schema, it should be possible to undo the change either by following a manual procedure or executing a prepared script. Another alternative used in highly virtualized environments is to provide the capability to revert the entire virtual machine to a prior incarnation; however, this would not account for database schema changes.

RP-D-7: Support Automated Startup, Shutdown, and Maintenance Mode Entry / Exit

Application developers should provide automated startup and shutdown scripts for a delivered application. These scripts should be developed in collaboration with the appropriate monitoring team and be invoked by the tools the data center provides for scheduling services, such as Tivoli Work Scheduler.

In addition, developers should provide scripts for entering and exiting a maintenance mode, if appropriate.

Rationale:

Starting and stopping applications should be as straightforward as possible.

Maintenance modes offer a mechanism to allow limited access to application functionality during updates to capability, data backup, or some other limitation that prevents full system access. There should be a visual indication of maintenance mode as well as a flag for non-visual interfaces (if necessary).

Release Management

The purpose of the Release Management (RM) process is to govern and manage the release of software baselines throughout CMS environments in a reliable, efficient way. Given that industry is innovating rapidly in this area and there are many ways to conduct release management, this topic covers just a few recommended practices.

RP-RM-1: Establish and Follow Organizational Standards for Deployment of Custom Software

Use the Agency or organizationally approved release management service for deployment of custom software into each of these environments in accordance with the approved RM process for deployments.

Rationale:

For example, in IBM mainframe environments, the Computer Associates’ Endeavor system is used to perform deployments. Deviating from such a standard increases support and training costs, limiting the ability to exchange resources and know-how between systems. Such deviation also increases the risk that different deployment methods might interfere with one another.

RP-RM-2: (Retired after TRA 2018R1): Contractors Must Deliver Certain Configuration Items

RP-RM-3: (Retired after TRA 2018R1): Minimum Acceptance Test Criteria

Common Engineering Support Services

Defect Tracking Services

A defect tracking system is a web-accessible, multi-user repository that is the source of record for software defects identified by any of the project stakeholders. Defects are entered, classified, worked, fixed, retested, and closed following the software developer’s defect correction process.

Version Control Services

The Centers for Medicaid & Medicaid Services requires designation of one version control repository as the official repository (source of record) used by the application system for software build and test. The contents of the Version Control System (VCS) repository represent the complete change history of all source code in the system.

At a minimum, the VCS must store and retrieve textual artifacts and establish named baselines for artifacts (also known as configuration items).

Note: CMS ARS Security Control SA-10 (“Developer Configuration Management”) requires configuration management practices.

Continuous Integration Services

The Continuous Integration services orchestrate the flow of source artifacts into and the flow of packages out of the build automation services. Builds are triggered on VCS check-in or on a schedule. Unit tests and integration tests may also be performed if the builds are successful. Results are posted on a Continuous Integration dashboard.

Build Automation Services

The build automation services generate a set of output packages (including documentation and other output products) based on source code, configuration files library packages, and other input sources. The output packages should be stored in the package repository.

Package Repository Services

A package is an installable component that contains instructions for how to install and uninstall the component. It should also contain a list of its dependent packages.

A release is effectively a versioned set of packages. A package repository is a service that contains trusted packages. These packages may have been acquired by contract, obtained via Open Source Software, or produced on behalf of CMS. A package repository contains both source packages and binary packages. Source packages can be used to build new binary packages. Binary packages are suitable for the system operator’s deployment into a production environment.

A package specification is a document or report that describes the contents of the package (sometimes called the manifest) as well as instructions for installation and removal. Typically, packages include mechanisms for dependency identification. Thus, if package A depends on packages B and C, then this is recorded in package A’s specification.

Often, a package builder uses package specifications to construct a package from source artifacts. A package installer takes one or more packages and installs them properly within the target system environment.

Inspection Repository Services

CMS recommends reviews and inspections as very effective practices for finding and preventing defects. See Software Inspection, Tom Gilb and Dorothy Graham, Addison-Wesley, 1993. and Software Inspection v Software Testing: How do they differ?

The inspection repository contains the artifacts of the inspection and the outcomes. Each inspection outcome of a code and design review should be coded with an indication of the number of issues or risks identified in the review to support useful metrics on the results.

Test Automation Services

Developer test automation services are essential for modern software development. Typically, by integrating test automation into the build automation services, tests run automatically after each successful build. Developer testing should include both white and black box testing. Unit testing alone is not sufficient.

Package Deployment Services

The package deployment services are charged with deploying trusted packages from the package repository and deploying them into the appropriate CMS environments (development, test, integration, or production), given appropriate authority. These services record these events in an audit trail. CMS does not specify the use of push- or pull-based mechanisms for package deployment.

Accessibility and Section 508 Test Services

Testing software systems for compliance with Section 508 accessibility requirements is mandatory. While automated testing is helpful, achieving all the requirements does require expert interpretation. An accessibility test service with software tools and experts knowledgeable in the field can help greatly. It is highly recommended that accessibility be built into the design of applications because this will increase the likelihood of successful Section 508 testing.

Performance Test Automation Services

These services can be used to performance test, stress, and soak test applications to determine their behavior under high and sustained load. Soak testing is useful for finding memory or resource leaks that can bring down a system even when performance is within expected parameters.

Security Assurance Services

While static analyzers can be very helpful at finding security vulnerabilities in code, it is important to supplement this analysis with dynamic analysis using tool suites designed to find vulnerabilities in the running system. Such a test bed can subject a system to systematic probing and attacks while still under development to better address security issues early in the software development process. Similarly, using data loss prevention techniques can detect causes of data leakage.

Note: CMS ARS Security Control CM-4(1) requires using a separate test environment before implementing software in an operational (production) environment to look “for security impacts due to flaws, weaknesses, incompatibility, or intentional malice.”

Web Services and Web APIs

Introduction

Background

A Service-Oriented Architecture enables a sustainable interoperability model between systems within the enterprise as well as for external systems requiring services to interact with CMS enterprise systems. To maximize the interoperability benefits with the appropriate security controls, CMS has established a set of enterprise standards and options for implementing SOA within the CMS enterprise.

CMS has determined that Web Services (WS) is the preferred implementation technology for SOA. For purposes of this architecture, Web Services includes both SOAP and REST implementations. Both are based on open Internet technologies and have diverse technical implementations in the marketplace.

CMS envisions SOA and Web Services as key enabling technologies for providing E-Government services over the Internet as well as within the enterprise.

The Web Services characterizations within this chapter align with CMS’s strategic priorities to develop an enterprise SOA for the CMS Processing Environments:

  • Promoting Consumer Centric Design
  • Promoting Standardization Across CMS
  • Promoting Developer User Interface / Experience

Purpose

The intent of this Web Services chapter is to delineate the Web Services architectural standards and relevant constraints for implementing those standards within the CMS TRA-defined environment, and for Web Services transactions between the CMS enterprise and external CMS partners. This chapter also addresses the integration of Web Service components within the CMS TRA-defined environment.

Scope

The concepts, strategies, and guidelines discussed in this chapter align with concepts, semantics, latest industry standards, and recommended practices defined by the World Wide Web Consortium (W3C), including but not limited to, Web Services architecture (available at: http://www.w3.org/TR/ws-arch/), Web Services policy (available at: http://www.w3.org/TR/ws-policy), XML (available at: http://www.w3.org/TR/xml11/), and other standards bodies as specified here.

SOAP, in general, uses the constellation of W3 WS-* and Organization for the Advancement of Structured Information Standards (OASIS) WS-I standards. SOAP supports information exchange via XML-formatted messages and metadata.

In contrast, REST uses “resources” as the central organizing concept. It is common for REST APIs to provide responses in several forms, such as XML and JSON, as well as other formats like GIF or MPEG as representations of those resources. REST relies heavily on the infrastructure of the World Wide Web, including Universal Resource Identifier (URI), HyperText Transport Protocol/Secure (HTTP/S), Really Simple Syndication (RSS), Atom, and other standards. It is possible to specify REST services using Web Services Description Language (WSDL) 2.0.

This chapter provides guidance for designing scalable, secure, and interoperable Web Services implementations by information systems but not the specifics of such system designs or supporting infrastructure. It supplies a framework or road map for defining Web Services-based interfaces between systems in the current and future enterprise. Although this chapter is not a substitute for any set of specialized business requirements, it offers a starting point for systems architects and developers to build interoperable solutions that address specific, granular requirements.

Changes to and Deviations from TRA Guidance

The TRB approves any changes to the CMS TRAsections and topics, following the process documented in CMS TRA Foundation section, Architecture Change Request Process.

All project requests for grant of special considerations or deviation from CMS TRA guidance must be provided in writing to the TRB. The TRB will respond to all requests in writing.

Requests to the TRB must include the following information:

  • Description of the requested deviation
  • Reason for deviation(s)
  • Full text of the TRA business rule or guidance in the TRA chapter for which there is a request for deviation
  • Description of other alternatives to the deviation(s) that have been considered
  • Enumeration of risks to CMS as well as other affected systems and stakeholders
  • A plan of action with dated milestones for remediating the deviation(s) and complying with the TRA.

Related TRA Chapters and Guidance

The following CMS TRA chapters may provide additional helpful information.

  • The CMS TRA Infrastructure Services section File Transfer topic may help when considering how to integrate batch-processing paradigms into the CMS SOA
  • The CMS TRA Application Development chapter – provides guidance for developing Web Services and data access services and their arrangement in the CMS TRA Multi-Zone Architecture.
  • The Network Services section Security topic – provides instruction on how to configure security applications to permit SOA network communication as well as permissible communication patterns between and among TRA zones and services.

A mailing list entitled, “CIO Resource Library Communications”, is available to notify subscribers when new or revised IT-related directives, policies, technical standards, TRA guidance / chapters, or guidelines are available. To subscribe to the new list, please select the following URL and enter your email address: https://public.govdelivery.com/accounts/USCMS/subscriber/new?topic_id=USCMS_12066

Concepts

This chapter employs commonly used CMS and Internet terminology, as described below.

SOAP

SOAP is an XML-based, open standards protocol for integrating applications and sending messages. SOAP Web Services use the XML data language for expressing three different concepts: service description, data, and metadata. Each concept has a distinct XML-based language as follows:

  1. Web Services Description Language describes a SOAP Web Service. The WSDL describes two parts—the messages (requests and replies) as well as service communication information.
  2. SOAP is the on-the-wire XML stream for a message. It contains an envelope with a header and body. The header includes addressing and security that describe how to process the body. The body consists of the bulk of the message.
  3. XML Schema Description (XSD) is the definition language to describe the structure of the XML schemas used as well as the legal values for each element. Complex schemas may employ an additional language called Schematron to express cross-schema validation (see also ISO/IEC 19757-3:2025Information technology – Document Schema Definition Language (DSDL) – Part 3: Rule-based validation using Schematron). The table SOAP-Specific Web Concepts and Definitions presents the applicable SOAP-specific Web concepts and definitions used in this chapter.
SOAP-Specific Web Concepts and Definitions
Web ConceptDefinition
ResourcesAn informational resource is any piece of data or content that a URI can uniquely identify and that the Web can express in a textual or binary representation and transmit across a network.
Uniform Resource Identifier (URI)Unlike REST, SOAP does not rely on the URI of a web service. SOAP is effectively transport independent.
TransportSOAP is normally invoked over HTTP, but it is also possible to use Enterprise Messaging.
Data RepresentationsSOAP messages are XML documents, both for requests and responses.
AttachmentsSOAP messages may have attachments that do not have to be XML based.
OperationsSOAP defines operations in the WSDL file and allows for an unlimited number.

SOAP messages are transport independent. Applications transmit SOAP-formatted messages using web standard protocols such as HTTP and Advanced Message Queuing Protocol (AMQP) as well as proprietary protocols such as IBM MQ Series or Microsoft Message Queuing (MSMQ). The Web Services standards allow SOAP messages to contain text or binary attachments.

REST

REST is an architectural pattern based on the underpinnings of the World Wide Web. REST works with a variety of web protocols, data formats, Multipurpose Internet Mail Extension (MIME) types, operations, message sequences, and resources. Resources include web pages visible to users and data objects consumed by software applications. REST uses HTTP or HTTPS to identify and act on information resources. It constrains the interface to a set of well-known, standard operations (e.g., GET, POST, PUT, and DELETE). REST’s aim is to exchange content with resources that maintain the state of information. REST uses URIs to identify resources and HTTP operations to act on those resources. SOAP relies principally on XML, while REST relies on HTTP and JSON.

Ultimately, REST proposes to leverage the current operations of the Web rather than adapting the Web to a new way of working. REST-Specific Web Concepts and Definitions presents the applicable REST-specific Web concepts and definitions used in this chapter.

REST-Specific Web Concepts and Definitions
Web ConceptDefinition
ResourcesAn informational resource is any piece of data or content that a URI can uniquely identify and that the Web can express in a textual or binary representation and transmit across a network.
Uniform Resource Identifier

A URI identifies a REST resource.

Unlike SOAP Web Services that typically use a single URI endpoint, REST uses as many URIs as necessary to describe each of the resources used.

Data Representations

A REST request message specifies the representation, or Multipurpose Internet Extension (MIME) type 5 (“application/” per RFC 2046), of the response message. REST specifies the response MIME type by:

  • Using the HTTP Request Header Field “Accept”
  • Using a URI query parameter to specify the MIME type
  • Using the filename extension of the requested resource
  • Using the default (html or text)

The response also contains a MIME type to designate the information type.

OperationsREST uses the operations defined in the HTTP protocol (RFC 2616), namely GET, PUT, POST, DELETE, HEAD, OPTIONS, TRACE, and CONNECT, but it is not limited to those operations. Extensions should be carefully considered because this may limit reuse and hamper compatibility.
StatelessnessREST uses HTTP request and response messages. The resources hold their own state, and the REST requests cause changes in the state of those resources. The requestor (whether client or another service) holds the state of the conversation; this explains the large scalability of REST. HTTP cookies are often used to maintain session state.
HypermediaHypermedia is the fundamental organizing structure of the Web. Hypermedia can have a role in Web Services design and can enable both dynamic discovery and state management. There is a form of REST called “Hypermedia As The Engine Of Application State” (HATEOAS).

REST supports various other patterns:

  • Asynchronous JavaScript and XML (AJAX) calls. AJAX allows a web page to invoke a REST web service directly from the browser. This allows for responsive webpages as well as Single Page Applications (SPA).
  • Really Simple Syndication feeds. RSS is an XML-based polling publish-subscribe method that allows a user to subscribe to a file and get an update when the file changes. Blogs and other online publications use RSS extensively.
  • Atom feeds. Atom feeds (RFC 4287) are another XML-based publish-subscribe method similar to RSS, accessed via the Atom Publishing Protocol (AtomPub).

The REST pattern may use alternate operations (other than HTTP GET, PUT, POST, or DELETE) and alternate MIME types other than XML and JSON, although this is uncommon. REST messages can transmit textual or binary attachments via hypertext or by sending a binary file as the REST message body.

SOA and Web Service Roles

There are three principal roles in a SOA—service providers, service consumers, and intermediaries:

  • Service Provider. An organization that creates, deploys, maintains, and operates a service for use by others. The provider is responsible for implementing and operating the web service as a sustainable capability over its lifetime of use.
  • Service Consumer. An organization that consumes (i.e., invokes, executes, or calls) web services. A web service consumer could be another web service, a business application, or a system component that is not a service.
  • Service Intermediary. An organization that acts as both a service consumer and provider.

Service Interfaces

As defined by the OASIS Reference Architecture Foundation for Service Oriented Architecture (SOA) and the DOJ Global Reference Architecture (GRA), the service interface is “the means for interacting with a service. It includes the specific protocols, commands, and information exchange by which actions are initiated [on the service].” Best practice calls for describing a service interface in an open standard, machine-interpretable format (i.e., a format whose contents are capable of automated processing by a computer), such as WSDL.

The service interface is the technical specification for the data exchanges (inputs, outputs, and error conditions), including specific description of data types, formats, constraints, assumptions, and expectations. The software code for an interface typically only describes the data exchanges but not the assumptions and expectations. It is important to document these as well because they may not appear expressed in the software. For example, a service that appends data to a database may expect rejection of any duplicates; however, there is no way to express this in the machine-readable software interface. It is important to document such assumptions in the human-readable text. The CMS Interface Control Document records this interface.

The core of a service contract will comprise the service description that expresses its technical interface as well as the assumptions, preconditions, and post-conditions that define its service contract. In addition to normal operation, it must define how service consumers will recognize service failure. For example, for SOAP Services, it would define the SOAP Faults as well as use of HTTP status codes (if appropriate).

In practice, a service interface is, in effect, an agreement (a service contract) between consumers and providers committing each to a course of action. The service contract is not a legal document but rather a technical specification that defines what the service will and will not do. The service contract gives service consumers a clear understanding of each service’s definition, limitations, and constraints.

Message Structure

Regardless of implementation, the basic structure of a web service is message oriented and involves an envelope, a header, and a body. Optionally, there may be attachments.

The envelope represents the message as a whole. In the physical world, there is data on the outside of the envelope and data on the inside. In the web services world, the data on the outside (which refers to source and destination address) is called the header. The data on the inside is the body or content.

The header provides routing information, timestamps, and other data about the message. The body is either the request or the response data.

The body can often be encrypted or digitally signed independently from the header. The header identifies whether to encrypt the body (or a portion of the body).

The HTTP headers, depending on SOAP version, play less of a role in SOAP Web Service execution than in REST services. REST Web Services use the HTTP header extensively and place the message body in the content portion of the HTTP request or response.

Transport

In a Web Services architecture, the transport functions as the network layer that conveys Web Service messages between service providers and consumers. HTTP (and HTTP/S) are the most commonly used transport layers for Web Services. CMS makes no transport or protocol restrictions other than adherence to CMS security standards. As a result, it is possible to use alternative protocols such as AMQP (ISO/IEC 19464), proprietary protocols such as IBM WebSphere MQ, or even Internet Inter-ORB Protocol (IIOP) to convey web services. CMS recommends using standard, openly available protocols (like HTTP and AMQP) for greater interoperability without cost to either consumer or provider whenever possible.

Service Repository

A service repository is a platform-independent directory that catalogs all services, providing an authoritative source for service descriptions, classification, and categorization information. It may contain such information as service level agreements, version descriptions, capacity, integrity, and security constraints. The repository organizes the services for easy discovery.

Repositories support both web service providers and consumers by providing an authoritative source for service information. Publishers post and manage configuration of metadata artifacts to support information sharing and to ensure their reuse across the boundaries of participating stakeholders. Consumers can search for services and retrieve essential information to use Web Services. The service repository should play an important role in governance and configuration management for Web Services. Some repositories support workflow, allowing submission of Web Services for approval and eventual publication through a documented process. This supports the governance process by offering service management automation.

Enterprise Service Bus

The Enterprise Service Bus (ESB) is an integration infrastructure for implementing independent sharing of data and business processes among connected Web Services.

ESB-connected consumers interact with various service providers to complete business processes. The ESB routes SOAP requests to multiple service providers. Components Connected with an ESB depicts Web Services and other components as they connect via an ESB.

Components Connected with an ESB (page 26)

Enterprise Messaging

Enterprise message (not e-mail) is a mechanism for sending and receiving messages, including but not limited to XML-based messages, either with request-reply or publish and subscribe (“pub-sub”) semantics between applications. In addition, enterprise messaging provides delivery options such as message queuing, fire and forget, and guaranteed delivery. Depending on the specific product implementation, enterprise messaging can offer disconnected operation, persistent messaging, internal acknowledgment of message receipt, and additional transactional recovery and message redelivery options.

Service-Oriented Architecture Overview

The SOA approach influences the solution architecture of CMS IT systems, from the conception of projects and throughout the system development life cycle (the CMS TLC). SOA simplifies development efforts by providing reusable portions of functionality and separating application architectures between the consumers of functionality and the providers of functionality.

Advantages and Disadvantages

A SOA is beneficial because it distributes computing rather than data. SOA systems provide functionality, called services, which they make available to other systems. The architecture for delivering these services is service-oriented because the service is the atomic building block, not the system. A “service” thus implements a unit of business or technical capability, provides an interface so it can be used (or consumed), and contains a description that defines its purpose and other pertinent metadata. Rather than copying data or functionality, a SOA uses network technology to access these services in a distributed fashion. The SOA offers a collective of interdependent services that function as a system to provide business value, unlike traditional architecture that supports a collection of systems that only operate on data and functionality within their own sphere of control.

The principal disadvantage of a SOA is that it distributes functionality and data sources. For some applications, this can degrade performance and reduce security. Architects should consider how to mitigate these issues when choosing an appropriate application and service architecture.

Need for SOA at CMS

Five sets of challenges frame the need for SOA at CMS:

  • Flexibility
  • Control
  • Mobile support and customer engagement
  • Data publication
  • Integration and modernization

CMS values flexibility and the capability of adapting quickly to change. At the same time, the Agency also must control the proliferation of data sources within the enterprise. CMS is committed to mobile platforms, which require a supporting SOA. Finally, enterprise IT modernization must allow for graceful retirement or replacement of systems without disruption to the rest of the IT environment. The following topics address these challenges, each warrant deeper discussion. “Success Factors for the SOA Solutions” presents the advantages of SOA as a compelling solution.

Flexibility and Adaptation to Change

CMS is committed to providing a flexible and adaptable environment to accommodate needs and changing requirements. Flexibility means reusing current capabilities in new ways to meet new goals. Adaptation means evolving those capabilities to address new goals. In CMS’s fast-changing environment, technology solutions that require a completely new implementation to meet business needs are less satisfactory because they take too long to execute and introduce greater risk.

Controlling Data Source Proliferation

The data source landscape of CMS’s current environment presents key problems for managing data:

  1. Business applications make their own copies of the relevant CMS data to fulfill their unique business requirements. Consequently, there are copies of the same information in multiple locations, and each copy has its own managing organization.
  2. Copies are altered to meet unique business requirements.

Over time, the continuing proliferation of data derivatives and data sources complicates the determination of the authoritative instance of the data.

The resulting data landscape has copies of data, each different from the other, and with unclear derivation histories. This landscape makes reuse problematic and encourages further deviation. CMS understands there is a tendency to produce point solutions when an organization’s development decisions only consider the needs of a single project or system. In distributed systems, these are point-to-point solutions. If unchecked, this approach spawns a self-perpetuating cycle of data proliferation: multiple copies of information abound, often scattered across independently managed instances, with limited integration and little opportunity or incentive for reuse. The filtering and augmentation criteria used to alter local copies is often undocumented, which forces projects to return to the source and make copies.

Enabling and Encouraging Data Publication

Data publication via APIs ensure that external consumers operate with an authoritative data source, rather than working from copies (as mentioned earlier). The availability of data via APIs means that mashups and other forms of data integration can be applied to derive new forms of data.

Any decision to publish data for external consumers must also balance their need for data with the need for privacy in accordance with CMS and HIPAA privacy guidelines.

Supporting Mobile Platforms for Customer Engagement

For many Americans, mobile platforms are their primary means of Internet communication, either by economic necessity or by choice. Mobile platforms, such as mobile phones, tablets, and other portable devices, interact with the digital world using web services. Making CMS services available to these mobile devices fulfills the vision of an Enterprise as a Service (EaaS) that delivers services to citizens.

REST-based web services are the de facto standard for mobile application support. For intermittently connected (“offline”) mobile and web applications, CMS recommends considering other forms of more complex middleware such as synchronization or change-set propagation (see Untether: Middleware Components to Support Intermittently Connected Web-Applications), or use of local storage (see Offline Web Applications). These solutions are out of scope for this chapter.

While allowing mobile users (constituents, employees, and contractors) to leverage CMS services securely in the field, CMS must protect the Agency infrastructure and data from malicious intent to ensure effective Agency operations.

Challenges of Modernization and Integration

All systems inevitably become candidates for modernization, replacement, retirement, or consolidation. Rarely do CMS systems operate in isolation; rather, they are likely to depend on other systems and data as they perform tasks as part of a business process or data flow. CMS systems have been developed and maintained for many years, with a variety of different technologies. As technologies and products are eventually retired, it is critical to access system functionality and data until CMS acquires suitable replacements. This introduces a strategic interoperability challenge.

CMS needs architectures that facilitate systems modernization and delivery of new capabilities without disrupting the existing process, data flow, and customer services.

Success Factors for the SOA Solutions

CMS has adopted a SOA methodology to fulfill the Agency’s needs in conducting business and developing IT systems. Web Services will be the principal, but not exclusive, method of implementing SOA in the CMS environment. Messaging technologies can also implement a SOA, although the proprietary protocols may limit reuse and increase costs over freely available Internet standards, such as TCP/IP, HTTP, XML, SAML, and other markup language-based technologies. The following success factors should guide development and implementation of SOA solutions that offer greater value to CMS organizations:

  • Adhere to a Service Contract. Web Services must adhere to a service contract, guaranteeing that service producers will not alter existing Web Services without either providing a migration path or a backwards-compatible solution. A typical solution includes a versioning method.
  • Comply with Open Standards. Open Web Services standards for systems integration (such as the WS-* standards, XML, JSON, HTTP, and URI) have evolved sufficiently to provide the basis for an overall approach to integration across a large, diverse undertaking like the CMS enterprise. Design and investment decisions should favor technologies (products, designs, and approaches) based on open industry standards to avoid vendor lock-in and improve system interoperability for CMS stakeholders.
  • Adopt an Operating Model. Successful SOA implementation includes a model for sustaining and maintaining shared services across organizational boundaries. The relationship is asymmetric—the service consumer (and by extension, the enterprise) stands to gain more than the service producer. A financially sustainable operating model is necessary to ensure that service producers can continue to provide services and service consumers can continue to rely on the availability and operation of those services. The principal issue is equitable compensation for services provided; however, compensation as a business, organizational, and financial challenge in implementing shared services is outside the scope of this chapter. In addition, service consumers and providers need to establish SLAs with appropriate compensation to ensure sustainable service performance while reducing the risk of lower performance and higher operating cost.

Establish Authoritative Sources for Information

CMS’s goal is to use a SOA to establish authoritative sources to manage and expose information using Web Services rather than continue making copies. There must be a single version of the truth in the enterprise. System maintainers and architects should strive to:

  • Avoid duplication of data
  • Avoid duplication of business rules or business processes
  • Expose a single, consistent way to perform a specific business capability

Many vendors now offer support for Web Service integration capabilities, such as transformation, content-based routing, mapping integration, composable capabilities (“mash-ups”), and orchestration. Vendors also support off-the-shelf adapters for enterprise applications.

Enable Reuse by Exposing Existing Capabilities as Services

Exposing existing capabilities as services is an excellent first step toward eventual systems modernization, migration, consolidation, or replacement. It effectively isolates other applications from the specifics of a target system by identifying and formalizing interfaces.

Service architects should consider exposing existing system capabilities as web services, especially when an existing system is a system of record for a specific kind of data. This approach will help migrate existing business processes and information systems into the CMS SOA. Good candidates for web services are finer-grained data exchanges rather than large “batch” data transfers more suitable for managed file transfer and batch processing. For more information on this mode of computation, please refer to CMS TRA – Infrastructure Services, File Transfer.

Exposing existing capabilities as services is also an essential step in supporting mobile platforms to advance the enterprise mission. Mobile platforms use SOA to leverage services to a variety of stakeholders. By providing these services, CMS fosters an ecosystem of service and data consumers while also enabling its own business.

SOA Service Design Principles

The following design principles apply to the creation of new services as well as enhancement and improvement of existing services. (Many of these principles were adapted from SOA Principles of Service Design, by Thomas Erl.) CMS has adopted performance as an additional characteristic from ITIL v3 to cover the relationship of SLAs to Web Services. Consistent with CMS’s longstanding practice, architects should consider security early in the design process.

Many of the following SOA Service Design principles are closely related and complementary:

  • Standardized Service Contract
  • Loose Coupling
  • Abstraction
  • Reusability
  • Design for Evolution
  • Autonomy
  • Statelessness
  • Discoverability
  • Composability
  • Performance
  • Security

Standardized Service Contracts

CMS recommends defining and expressing services through standardized service contracts. CMS provides a template for an ICD that documents the Web Service interface (or service contract). The ICD must contain the explicit service definition as well as implicit definitions such as performance, security, capacity, and assumptions. To ensure recording such information, organizations should consider adopting a service specification template, which extends the ICD’s information and can address the specific nature of Web Services.

Loose Coupling

Loose coupling is a general, cross-cutting design principle maximized by applying all principles described in the following topics. Service developers should reduce dependencies between a service’s implementation and its Web Service consumers in all aspects of service design. Service consumers should depend on the interface and not the implementation.

For example, CMS advises service developers to expose information using a Message Model that is a different schema than the data storage model. This decouples the interface from implementation, allowing for greater flexibility and migration to different implementations. By contrast, coupling service consumers to the data source implementation limits implementation options in the future, such as migration to new technology, or database consolidation.

Abstraction

Abstraction helps make Web Services more reusable and more durable because the service can endure implementation changes without any effect on consumers. CMS expects service developers to abstract away the intrinsic details of the underlying web service implementation. The service consumer benefits when service developers hide such details as the programming language types, location-specific information for database or computer names, and parameters closely tied to COTS products.

Autonomy

CMS desires Web Services that can be independently tracked, monitored, deployed, and modified. Architects should identify dependencies as part of implementation documentation but not as part of the service interface. Hiding inter-service dependencies from service consumers will prevent unnecessary coupling between such internal design decisions. Service consumers should base their use of a service on the service interface, not the details of a service’s implementation.

Reusability

Reusability is a primary consideration in making the business decision to produce a CMS Web Service. Reusability decreases overall cost and increases speed of delivery by employing already tested and working services. Developers should emphasize generality in design services to encourage reuse and avoid too much specificity.

Design for Evolution

Change is a constant feature of services and service needs. The SOA must allow service consumers and providers to evolve the architecture with some degree of independence. For example, services can evolve and exist simultaneously in several versions (with different compatibility requirements). Change management allows consumers and providers to coordinate. Without the capability to operate multiple versions, service changes will affect service consumers, necessitating complex coordination for migrations.

Statelessness

In a stateless service, the service maintains no internal state of its own: all data is passed in (except perhaps reference data), computed on, and results returned. Thus, there is nothing to copy or maintain between multiple instances. Although this is ideal, it is often not practical.

When a service must be stateful, it should represent a business transaction as a complete unit of work. This allows the service to use transactional data stores effectively and ensure that data is in a consistent state. Stateless services are efficiently scalable.

Discoverability

A Web Service should be discoverable, with associated, machine-readable metadata that allows a service consumer to discover and understand the service. In production use, however, services must already be known to the application and not discoverable.

Some Web Services containers will emit the WSDL for the service if the service URI endpoint has a ‘?WSDL’ or ‘/WSDL’ appended. This capability is permitted only within the development and test environments. CMS requires disabling such capabilities on production Web Service endpoints.

Composability

A Web Service should be capable of combination into larger units (sometimes called composite web services or lightweight composition) as well as orchestration into a process (such as via business process automation or WS-BPEL), depending on the complexity of the orchestration. This design principle means that a Web Service forms a component model. At CMS, composition occurs in either the Application or Data Zones because services are not permitted to perform business processing in the Presentation Zone.

Performance

Web Services should perform within their required and documented service levels. Thus, Web Services should have documented performance, capacity, and latency levels, especially for service consumers outside the service developer’s organization.

Separate services create flexibility and reusability for developers, but can increase machine workloads for a given task. Although this may be an acceptable tradeoff for the developer, specific performance-related impacts should be considered, especially in CMS’s high data-volume and high user-workload business functions. Application architects must fit execution cycles for synchronous tasks into user response-time expectations, and batch execution cycles into the time allotted for processing. Examples of such tradeoff considerations include:

  • Splitting service components into separate services versus a larger single service that may have several methods. When server-to-server calls are added, extra workloads must be considered.
  • Although generic data access services may simplify data access for applications, they may create extra data calls and force table joins over large volumes of data. The preferable goal is creating specific data services that serve fewer use cases, while allowing for smaller, more pointed queries in any one call. Other alternatives include staging tables and caching to allow generic data access services to function in a reasonable performance window.

During the design phase of the project, service designers should analyze the known use cases of potential consuming applications for performance considerations.

Security

Web Services must operate within a well-defined security context, with explicit definition of trust expectations and mechanisms for securing both the operation and data affected by the Web Service.

The service designer must account for all three dimensions of security (confidentiality, availability, and integrity) in the design.

Business Rules

CMS has adopted the following business rules governing the use of Web Services (WS) within the CMS Processing Environments. To comply with TRA standards, a SOAP or REST Web Service must adhere to the following business rules. Where appropriate, related CMS ARS controls are identified to provide supporting guidance for associated business rules.

BR-WS-1: Describe All Services

CMS requires documentation of Web Services, regardless of implementation technology, in a system’s ICD (or equivalent). This includes a description of the service’s functional and non-functional characteristics, as follows:

  • Functional characteristics
    • Preconditions
    • Post-conditions
    • Assumptions
    • Constraints
    • Inputs and Outputs
  • Non-functional characteristics
    • Performance characteristics
    • Capacity characteristics
    • Transactional characteristics
    • Security characteristics
    • Privacy implications, including PII, PHI, etc.

Service developers should consider submitting descriptions of their high-level services to the CMS Enterprise Architects to be documented in the HHS Enterprise Architecture Repository (HEAR). Through this activity, CMS documents the relationship between shared technology services, business services, and organizations.

In addition to these universal documentation requirements, all Web Services must satisfy the following SOAP- and REST-specific requirements.

Related CMS ARS Security Controls include: , PT-2 - Authority to Process Personally Identifiable Information, PT-3 - Personally Identifiable Information Processing Purposes, RA-3 - Risk Assessment, RA-8 -Privacy Impact Assessments, SA-1 - Policy and Procedures, SA-4 - Acquisition Process, SA-9 - External System Services, CA-2 - Control Assessments, PM-21 - Accounting of Disclosures and SI-18 - Personally Identifiable Information Quality Operations.

SOAP Rules

A SOAP Web Service must expose its service contract using WSDL, and must document its service contract. The service contract must describe the following items:

  1. Actions
  2. Data (including media types)
  3. Messages
  4. Interfaces

The documentation may be in the developer’s own template. Examples appear at: Amazon SQS template snippets and at 2012-11-05/QueueService.wsdl (describing the WSDL and human-readable documentation for the Simple Queuing Service).

REST Rules

A REST Web Service must document its service contract by describing the following items:

  1. The resources
  2. The operations that each resource supports
  3. The representations (content types, which include MIME types) of the resource
  4. The input and output parameters in the URI that reference each resource
  5. The HTTP headers associated with the operations

The webpage may be in the developer’s own template. An example appears at Amazon S3 template snippets for the Simple Storage Service (S3).

Rationale:

CMS Web Services must have a human-readable specification for Web Service consumers and providers. It is not sufficient to provide the web service schemas. The specification must explain the pre-conditions, assumptions, and post-conditions of using a given web service. This information helps service consumers understand the functional and non-functional capabilities and limitations of each service.

The use of machine-readable specifications makes it possible to validate web services and share them more easily between service consumers and service providers.

In addition, it is highly recommended that implementations follow de facto industry standards, such as Swagger, Web Application Description Language (WADL), RESTful API Modeling Language (RAML), or JSON Schema to provide specification of REST Web Services.

BR-WS-2: Services Must Use Standard Invocations

SOAP Rule

A SOAP Web Service must use SOAP formatting in accordance with the WS-* standards. If CMS adopts a set of standard SOAP header elements, these must be included in all SOAP Web Services.

REST Rule

A REST Web Service must comply with HTTP 1.1 (RFC 9112) or later and use compliant headers and HTTP method names.

Rationale:

A Web Service must be invoked and respond in a standard way to achieve the broadest application and encourage reuse.

BR-WS-3: Services Must Validate Input and Outputs

It is mandatory to validate input parameters for security, bug trapping, and catching illegal user inputs. CMS requires both syntactic and semantic validation. Semantic validation must include type, format, references, value, context, and applicability.

SOAP Rules

A Web Service must use XML Schema Definition to validate syntactically and structurally any XML data received.

REST Rules

The REST Rules regarding XML are the same as the SOAP Rules. For JSON, the service developer is responsible for both syntactic and semantic validation.

Rationale:

Validation of inputs and outputs contributes to better security, privacy, testing, bug detection, and spam prevention. “Improper enforcement of message or data structure” is a common weakness according to the Common Weakness Enumeration.

Recommendation:

Architects should use XML-aware network elements (such as IBM DataPower or F5) to validate XML and JSON objects rather than developing custom validation code.

Related CMS ARS Security Controls include: SI-10 - Information Input Validation.

BR-WS-4: Use TRB-Approved Data Zone Mediation and Data Access Services to Access Data in the Data Zone

This rule is identical to BR-SA-4 in the Application Development Guidelines chapter and is referenced because of its importance.

BR-WS-5: Messages Must Include Timestamp and Originator

SOAP Rule

A SOAP message must contain a current timestamp using ISO 8601 [Universal Time Code (UTC)] and identify the sender.

REST Rule

A REST message must contain a current timestamp using ISO 8601 (UTC) and identify the sender.

Rationale:

For usage logging, debugging, and security, each service invocation and response needs an origination timestamp and originator (identifier). Using such timestamps makes it easier for CMS security to perform event correlation because other CMS logs employ UTC.

Related CMS ARS Security Controls include: AU-7 - Audit Record Reduction and Report Generation.

BR-WS-6: Services Must Use Open Data Formats

SOAP Rule

A SOAP message must implement Unicode Transformation Format (UTF) 8-bit (UTF-8) and must not contain embedded binary objects in the body. If binary objects are small enough, they must be encoded, for example, by Base 64 encoding. Binary objects that are too large to send efficiently in a SOAP message must be sent using a protocol such as Direct Internet Message Encapsulation (DIME) or Message Transmission Optimization Mechanism (MTOM).

REST Rule

A REST message body may contain UTF-8 encoded text, XML, JSON, YAML, multimedia, or other open textual or binary formats (such as HL7, FHIR, XBRL, ACORD.)

Rationale:

Avoid proprietary data formats because they limit reuse and may introduce unintended vendor lock-in.

BR-WS-7: Web Services Must Follow CMS Encryption Policy

CMS TRA Network Services, Security Services Business Rules based on the CMS ARS policy for determining when network communications must, should, or should not be encrypted. These rules apply equally to Web Services traffic and content, regardless of underlying implementation technology. In addition, all cryptography must be performed using validatedFIPS 140-3 security modules.

Related CMS ARS Security Controls include: AC-3 - Access Enforcement, SC-8 - Transmission Confidentiality and Integrity, and SC-13 - Cryptographic Protection.

Rationale:

Adherence to CMS TRA guidance ensures consistency across CMS Web Services.

BR-WS-8: SOAP-Based Services Must Comply with WS-* Standards

A SOAP Web Service must comply with WS-* standards, and thus, should be able to accept WS-* payloads and extensions for security, reliability, and transactions as required.

Rationale:

Following industry standards is the best way to improve interoperability.

BR-WS-9: Inter-Zone Web Services Must Transverse a Mediated Service

Access to a Web Service from outside of the same zone must traverse a service that employs mediation principles.

Active mediation is required between a service consumer and a service provider. That is, this component actively validates requests and does not simply function as a pass-through. In practice, mediation is often implemented using an enterprise messaging service, ESB, or XML appliance. Alternative implementations may be proposed.

Rationale:

A single entry point has several advantages. First, it decouples the service consumer from the service provider. The service consumer does not know how many instances of the service exist behind the mediation service, thus allowing the service provider to scale the service with multiple service provider instances. Second, it requires only one firewall rule to allow entry. This reduces network administration burden. Finally, the entry point is a natural location for monitoring and control.

The mediated service could be the sole occupant of the Presentation zone for an application if it performs active mediation, challenging requests for validation and access control.

BR-WS-10: Web Services Must Be Version Numbered

All CMS Web Services, regardless of implementation technology, must have a published version number that can be (1) checked by client applications, (2) explicitly requested by client applications, and (3) distinguishable from older or newer versions of the same Web Service.

Rationale:

A change to the semantics without change to the interface specification still means altering the version of the Web Service to reflect the change to semantics. In other words, the version does not solely apply to the interface but to the service as a whole.

Following a versioning scheme decouples release cycles for service consumers and service providers. As the number of service consumers increases, the difficulty in coordinating release cycles between producers and consumers increases dramatically. Please refer to the “Web Services Versioning” approach.

BR-WS-11: Use Certificate-Based Mutual Authentication for Machine-to-Machine Web Services

CMS requires cryptographic certificate-based mutual authentication to perform machine-to-machine communications.

This rule is relaxed if both consumer and provider are communicating over the CMSNet / Core Virtual Routing Forwarding (VRF) network, in CMS data centers under current CMS Authorization To Operate, and the message would not traverse the Internet, even over a Virtual Private Network.

Rationale:

This rule meets the requirements set forth in the CMS Risk Management Handbook, Volume III, Standard 3.1 CMS Authentication Standards.

Using mutual certificates ensures end-point authentication and eliminates the need to store usernames and passwords in the clear on servers communicating with CMS.

The rationale for relaxing the requirement is a risk-based decision for messaging between trusted data centers over a trusted network.

BR-WS-12: Messages Must Pass through All Intermediate Zones

A web service message must pass through each intermediate zone on the way to and from its destination. A zone is defined by the partition created by virtual or physical firewalls where security challenges have been implemented. The CMS TRA requires that at least three separate security challenges be implemented to protect CMS data.

Rationale:

This preserves the integrity of the CMS TRA Multi-Zone Architecture’s Defense-in-Depth strategy. All requests and responses pass through each of the intermediate zones, affording opportunity for CMS security infrastructure to inspect and analyze messages.

Thus, if a message or transaction originates in Internet and will be processed in the Data Zone of a CMS Processing Environment, it must traverse the Presentation, and Application Zones before reaching the Data Zone. No shortcuts are permitted. Likewise, the response must traverse each intermediate zone before reaching the Internet. In the cloud environment, where the zones may not be so clearly defined, the responses must pass through the same zones as the request.

BR-WS-13: CMS Public APIs Must Be Published

All CMS Publicly accessible APIs must be documented on the CMS Developer tools website.

Rationale:

This makes finding CMS APIs easier for external developers.

Web Services Best Practices

CMS recommends that service developers adhere to the following recommended practices to ensure the most effective implementation of Web Services.

Implementation Independent

CMS prefers technology-independent implementations. Implementations should follow these recommended practices, regardless of chosen implementation technology.

Consider Using Enterprise Messaging and ESB

For Java-based applications, Java Message Service (JMS) is the preferred mechanism for writing programs that use enterprise messaging because it is the Java Enterprise Edition (EE) standard and has strong vendor support. Because it is an interface specification, application code becomes more portable. Given that JMS does not standardize the implementation protocol, it is necessary to use the same manufacturer’s JMS implementation library on both sides of the channel. Unfortunately, this reduces JMS’s platform independence and flexibility. There are industry initiatives in place to address this, such as AMQP.

For other platforms, such as native IBM Mainframe z/OS or Microsoft.NET, product-specific APIs may be used.

For more information about enterprise messaging, please consult Application Development, Enterprise Messaging.

Use a Transaction Identifier

By providing a transaction identifier (ID) in the web service request, it becomes easier to track a transaction through its life cycle. Transaction IDs are typically whole numbers. Occasionally, it is helpful to use a compound identifier consisting of additional metadata such as a source system or source server.

Use a Session Identifier to Group-Related Transactions

It is easier to track group transactions, perform metrics, and debug complex transactions when web service requests include a session identifier. A session ID is a grouping mechanism for reporting but is not the same as a session cookie. To prevent vulnerability to session hijacking attacks and session prediction attacks, select a secure means for using a session ID to manage sessions.

Typically, session identifiers are whole numbers. It may be helpful to use a compound identifier consisting of additional metadata such as source system or source server.

Consider a “Heartbeat” Service

A “heartbeat” service — or active monitoring (see The Anatomy of APM — Foundational Elements to a Successful Strategy) — checks that the service delivery platform operates correctly without having to process a synthetic” transaction. This is a lightweight service, intended to be called frequently without impacting the capacity of the system.

Consider a Synthetic Transaction

Synthetic transactions verify correct integration of a newly installed system with dependent services. These transactions also monitor “heartbeats” as well as throughput in production systems.

Synthetic transactions present at least two known perils: (1) misuse because of security vulnerability, and (2) potential to skew monitoring statistics if issued in sufficient quantities. It is crucial to achieve a secure design for these transactions.

CMS recommends disabling synthetic transactions during installation (default mode). During production, synthetic transactions must be secure and minimize application load while still providing meaningful business reporting of application functionality.

Application business owners must approve the use of synthetic transactions in production.

Consider for a Mechanism for Production Connectivity Validation

Production interface validation allows external parties to validate connectivity to the service. External parties could connect to the CMS service and submit a small number of specially identified transactions to prove correctly established connectivity and that CMS is receiving transactions. Depending on the implementation, such a mode allows external parties to ascertain (1) acceptance of the HTTP/S certificates and (2) whether the XML or JSON payloads are minimally valid. CMS recommends limiting the number of these transactions. The goal is to demonstrate that the production systems can communicate.

One possible mechanism for this purpose is using a “flag” to identify the transaction for special handling. This mechanism is neither a test capability nor a way to provide testing capabilities to external parties. Testing capabilities should be provided with a separate, designated test system that meets CMS rules for production data management (including but not limited to handling PII, PHI, and sensitive data).

Avoid Programming Language-Dependent Serialization

Many programming languages (such as Java and Python) permit serialization operations; however, these are typically language specific. Services requiring language-specific serialization are not reusable by other programming languages. The realization of any short-term productivity gain is therefore lost over the long term because lock-in limits future options. CMS recommends using the JSON or XML serialization formats that have broad adoption by industry. Accordingly, CMS discourages the use of Java Remote Method Invocation (RMI).

SOAP Best Practices

CMS recommends the following best practices for SOAP Web Services.

Use the Least Number of WS-* Standards Necessary to Achieve Your Goals

Although many of the WS-* standards have value, they can affect portability and stability, especially WS-* standards that have not achieved broad market adoption. Choosing a workable subset is a practical solution.

Use Document / Literal WSDL Operations

New WSDL SOAP bindings should employ “Document/Literal” as opposed to Remote Procedure Call (RPC) style (known formally as “RPC/Encoding”). RPC services use the SOAP RPC conventions and SOAP encoding if required. A developer may use any appropriate reference mechanism that meets WS-* standards.

SOAP API Design Best Practices

The following short summary of API design practices, which is neither exhaustive nor exclusive, help make SOAP APIs more reusable, secure, and stable:

  1. Choose platform independence.
  2. Use SOAP Faults, but do not send stack traces back as part of error and exception handling.
  3. Use ISO 8601 (UTC) date and time formats.
  4. Use SAML assertions.
  5. Support Single Sign-On (SSO) when possible,
  6. For large data sets, allow for pagination of results.
  7. Use UTF-8 encoding.
  8. Consider using API Keys.

REST API Design Best Practices

There are numerous sources for REST API Design Best Practices.

A good set of REST best practices appears in the book REST API Design Rulebook. Practical implementation details appear in the book RESTful Web Services Cookbook. Integrity of REST (relative to the original definition) is not as important as security, simplicity, performance, and interoperability; for example, it is permissible to use REST in a RPC manner. The following summary of API design practices, which is neither exhaustive nor exclusive (and inspired by White House, the GSA 18F group, and Microsoft Guidelines) help make REST APIs more reusable, secure, and stable:

  1. Resource Identification by a unique, persistent URI.
  2. API endpoints identify nouns, not verbs.
  3. Platform independence.
  4. Use HTTP error codes but be cautious about revealing internal structures. Do not send stack traces back as part of error and exception handling.
  5. Discovery via description of URI interfaces.
  6. ISO 8601 (UTC) date and time formats.
  7. JSON web tokens (or equivalent).
  8. Support Single Sign-On.
  9. Use of HTTP standard verbs GET, DELETE, HEAD, and PUT only for idempotent operations (defined as one for which the side effects of N>0 identical requests is the same as for a single request). Beware of such as incrementing numbers and changing dates that are not idempotent
  10. For large data sets, allow for pagination of results.
  11. Use UTF-8 encoding.
  12. Consider using API Keys.
  13. Put version numbers in the URL (e.g., api/v2/resource/{id}.json).
  14. API endpoints use HTTP Accept Headers to determine response media type.
  15. Use Cross Origin Resource Sharing (CORS) for security (not JSONP).
  16. Provide mechanism to throttle API.
  17. Consider using a JSON validation scheme such as the one proposed by json-schema.org.

Service Deployment in the Multi-Zone Architecture

The CMS TRA Foundation, Multi-Zone Architecture section provides definitions of the Zones. It places restrictions on the location and type of Web Services deployed within CMS Processing Environments. Both the CMS TRA Multi-Zone Architecture and the Virtual Data Center concept encourage communications within like zones, regardless of data center. While communication may be possible to the cloud environments from CMS data centers, the cloud boundaries for ‘like’ zones is not always clear. Zones in the cloud may be a mixture of CMS and cloud service provider services, therefore the implementor must verify the security posture when communicating between zones. For the specifics of multi-data center access, please refer to CMS TRA - Network Services, Wide Area Network Services, which provides additional business rules regarding security and networking.

Please note that the multi-zone architecture, most notably in the cloud environment, may not necessarily specify the number of zones. The zones defined in the multi-zone architecture are built upon the services framework and the functions these services perform. The presentation zone supports edge services while the application and data zone support applications and data services correspondingly.

Allowable Zones

The Presentation Zones may host Web Services only if:

  • The services provide access to static data via HTTP GET requests where the http server performs no processing (such as template expansion or reference data) other than delivering the requested data.

All requests originating outside the Presentation Zone must be validated and inspected at the Presentation Zone. The Application Zone is the recommended zone for most Web Services, and in particular, for hosting such services as:

  • Business Rules Services
  • Portlet Services
  • Business Logic Services
  • Business Process Automation Services
  • Enterprise Application Integration Services
  • Email Routing Services
  • Aggregation Services
  • Orchestration Services

The Data Zone is recommended for hosting such services as:

  • Data Access Services
  • Data Transformation Services
  • Data Replication Services
  • Job Control Services
  • Mainframe-based Services

Management Zone services are strictly restricted to support for operations and information security. Ordinary applications may not host their services in the Management Zone .

The following services should be hosted in the Management Zone :

  • Logging Services
  • Control Services
  • Other Security or Infrastructure Support Services

Data Access Services

In accord with the CMS TRA Multi-Zone Architecture, any applications that require information from CMS data sources must not communicate directly with the data sources (e.g., databases or external web services) themselves. Instead, they must use Data Access Services and mediation principles to access data in databases. By decoupling applications from the database implementations, CMS can deploy new technologies and upgrade old ones without undue effect on service consumers.

Within the TRA Application Development section, Business Rules and Recommended Practices topic, BR-SA-4, Use TRB-approved Data Zone Mediation and Data Access Services to Request Data from the Data Zone covers the use and mediation of data access services.

Data Access Services store, access, and update data from CMS internal and external data sources. They are a layer of data access logic between the distributed data stores and the service consumers who access them. Data is accessible through data services as an abstraction, which helps to hide the details of the raw data. A Data Web Service delivers data using a Message Model that includes the metadata and content. Service consumers leverage this Message Model to access the required data, and the Data Web Service handles the collection and distribution of the data to the shared data consumer.

Data Access Services provide loose coupling between the service consumer and the underlying physical data stores (e.g., databases). These services offer the capability to improve data quality, integrate data, apply business / data rules, and transform data from varied sources. They help integrate, aggregate, and manage data sources. Data Access Services also provide authoritative interfaces to enterprise information.

To best ensure the integrity and authority of CMS enterprise data, Data Access Services should be the sole mechanism by which an application web service or business application can interact with the CMS enterprise data. The orchestrated use of Data Access Services can better ensure effective database interaction in business application performance while also avoiding data contention issues. As shown in Data Access Services in the Data Zone, these services are deployed into the Data Zone.

Data Access Services in the Data Zone (page 27)

Data Web Services must maintain the integrity of data sources by managing transactions autonomously. These services must not rely on or expect service consumers to maintain databases or other data sources in a consistent state on behalf of the consumed web service (e.g., passing in data sources as parameters). It may also be necessary for Data Web Services to orchestrate data transactions across multiple database tables to ensure data source consistency. Data transactions must avoid deadlock by design and ensure performance levels remain within service levels.

Web Services Technology Overview

This chapter identifies the set of design guidelines that should affect information system design choices for implementing Web Services. These guidelines address only the architectural aspects of systems, independent of any individual system’s business requirements. This topic establishes descriptive attributes for the architectural support of Web Services implementations to assure common characterization of what constitutes a Web Service.

Web Services Versioning

By design, the implementation of services can change with little impact to the service consumers (downtime, rework, compatibility, etc.). Changing the interfaces—the way one invokes those services—requires changes to the service consumers as well. CMS recommends that architects and designers focus on the impact of each change to service interfaces. Changing an interface cannot be done unilaterally. Instead, adherence to a governance process will facilitate negotiation to resolve differences between Web Service providers and consumers.

Web Services versioning provides a mechanism to allow service consumers to continue using previous versions of a web service while working on their transition plans.

Version Implementation

CMS employs the following principal methods of versioning Web Services:

  • Endpoint (URL) versioning
  • Message versioning (either using JSON or XML versioning)
  • XSD versioning

In URL versioning, the major version number of a web service endpoint is appended to the URL. In Message versioning, the message header contains a version number field. This can be a version in the XML or JSON payload or it can be a XML XSD version. In XSD versioning, the XSD for the WSDL contains the version as part of the schema definition.

URL versioning is generally the simplest to explain and use.

Versioning Scheme

The SemVer 2.0.0 (Semantic Versioning) specification provides a mechanism for naming versions of products following de facto industry standards. The scheme is summarized as follows:

Given a version number MAJOR.MINOR.PATCH, increment the:

  1. MAJOR version when you make incompatible API changes,
  2. MINOR version when you add functionality in a backwards-compatible manner, and
  3. PATCH version when you make backwards-compatible bug fixes.

Additional labels for pre-release and build metadata are available as extensions to the MAJOR.MINOR.PATCH format.

For REST services, the version is reflected in the URI or as a field in the payload. XML payloads may use XML XSD versioning. Most often, the URI would reflect the version.

For SOAP services, the URI or WSDL reflect the version.

Dependencies on Standards

SOAP Web Service Standards

Successfully implementing solutions using this chapter depends on having integrators address multiple standards, appropriate elements, and other CMS TRA chapters.

The extensible nature of Web Services’ architecture has led to exponential growth in Web Services standards and specifications as standards bodies and companies attempt to leverage this technology. New specifications allow new functionality for Web Services, and these specifications are subsequently submitted to standardization bodies, such as the W3C or OASIS. Standardization is necessary to ensure that different implementations of Web Services engines engage one another seamlessly and thus fulfill the interoperability objective of Web Services.

Although this recent activity has created numerous standards and specifications that address many aspects of Web Services, only a narrow set are relevant to this supplement, as shown in SOAP-Specific Web Concepts and Definitions. For a more detailed discussion, please refer to the following sources:

SOAP Web Service Standards by Category
CategoryStandardsNotes
Security
  • WS-Security
  • WS-Policy

Focus on Security Assertion Markup Language (SAML) and the related federated security model:

Processes
  • WS-Business Process Execution Language (BPEL)
Focus on support of distributed processes, where the participant initiating the process maintains a controlling position throughout the lifetime of the process. Future support should improve monitoring a process’s progress based on rules.
Data Formats
  • WS-SOAP
  • XML
SOAP payloads are XML documents.

 

REST Web Service Standards

For REST, the standards are different and there are fewer of them. REST relies on the infrastructure of the World Wide Web and the HTTP protocol for standards. REST relies on media types to identify payload (data) standards. Each Internet Assigned Numbers Authority (IANA) registered media type has a template that defines its use and format. REST Web Service Standards presents the REST Web Service Standards for Security, Transactions, and Data Format.

REST Web Service Standards
CategoryStandardsNotes
Security
  • XML Digital Signature
  • HTTP/S using Transport Layer Security
There is no current standard for message-level encryption for JSON messages. XML has greater choices for this.
Transactions
  • HTTP
HTTP does not support transactions beyond the idempotency rules for HTTP request methods (GET, PUT, POST, HEAD, DELETE, etc.).
Data Format
  • XML
  • JSON
  • Other media

REST is most often used with JSON and sometimes with XML.

HTTP/S does include compression options.

Other media are identified using media types in the HTTP request and response headers.

Web Services Security

SOAP- and REST-based services have very different security infrastructures. SOAP security models follow W3, OASIS, and other XML-based standards. REST security follows the technologies of the Web with no reliance on XML or OASIS standards. This topic describes CMS’s adopted SOAP and REST standards for securing Web Services.

HTTP/S for External SOAP and REST Web Services

The most frequently used (and only common method between both SOAP- and REST-based web services) is HTTP/S. HTTP/S is HTTP protocol over Transport Layer Security connections. TLS protects the connection between client and server against risk of interception by encrypting the data and authenticating endpoints. TLS is capable of server-only or mutual authentication where both the client and server are authenticated cryptographically.

There are a few exceptions: please refer to the Security Services section for specific details.

Protect PII, PHI, and Sensitive Information

The CMS ARS contains a broad set of CMS security controls based on NIST requirements. CMS mandates compliance with the following security specifications:

Protect Sensitive Information in Transit

  • All Personally Identifiable Information, Protected Health Information, or other sensitive data entering, exiting, or in transit within the data center (within or across zones) must be encrypted and secured according to the guidance in the CMS ARS. Applications must use Transport Layer Security (TLS) at the highest available level to exchange information securely. When possible, applications should use mutual authentication to ensure the identity of both parties in an information exchange. If encrypting sensitive information is not technically feasible or demonstrably affects the ability to support mission operations, compensatory controls must be implemented as part of a CIO approved risk acceptance plan.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations, CP-9 - System Backup, CP-9(8) - Cryptographic Protection, PT-7 - Specific Categories of Personally Identifiable Information, MP-5 - Media Transport, SC-7(24) - Personally Identifiable Information, SC-8 - Transmission Confidentiality and Integrity, SC-12 - Cryptographic Key Establishment and Management, SC-13 - Cryptographic Protection, AC-2 - Account Management, AC-3 - Access Enforcement, AC-5 - Separation of Duties, AC-6 - Least Privilege, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives, SI-7 - Software, Firmware, and Information Integrity, SI-10 - Information Input Validation, and AC-21 - Information Sharing.

Protect Sensitive Information at Rest

  • The application must protect the confidentiality and integrity of all sensitive data (including all PHI or PII).
  • The application must protect the confidentiality and integrity of all sensitive information (including all PHI or PII) according to the guidance in the CMS ARS. This includes using encryption that meets or exceeds the FIPS 140-2 encryption standard. The implemented level of encryption must be aligned to the sensitivity of the information. If encrypting sensitive information is not technically feasible or demonstrably affects the ability to support mission operations, compensatory controls must be implemented as part of a CIO approved risk acceptance plan.

Related CMS ARS Security Controls include: MP-4 - Media Storage, SC-12 - Cryptographic Key Establishment and Management, SC-13 - Cryptographic Protection, AC-2 - Account Management, AC-3 - Access Enforcement, AC-5 - Separation of Duties, AC-6 - Least Privilege, SI-4 - System Monitoring, SI-5 - Security Alerts, Advisories, and Directives, SI-7 - Software, Firmware, and Information Integrity, SI-10 - Information Input Validation, SC-1 - Policy and Procedure, and SC-28 - Protection of Information at Rest.

SOAP-Specific Web Services Security

An essential reason to use Web Services is to enable software integration that crosses boundaries and leads to organizational integration. When legacy systems are involved, integration requires one application to extend its architecture to communicate with another application with which it did not previously communicate. Integration may also mean reusing existing services for new purposes.

As CMS’s Web Services evolve, interoperability should encompass communication standards that allow distinct systems to cooperate—an especially important capability when those communications must be secure.

Implement XML Filtering

CMS requires implementation of XML filtering of Web Service transactions before they propagate into the enterprise. Toward that end, XML documents must be:

  • Well-formed (syntactically correct according to the XML 1.0 specification)
  • Schema conformant to the referenced schemas
  • Semantically correct (contains valid data)

Typically, XSLT-based rule sets, executed using dedicated XML gateways, perform XML filtering.

All requests that fail to pass validation, must be logged for further analysis, and prevented from entering the CMS Processing Environment.

Anticipate and Protect against XML Denial-of-Service Attacks

CMS seeks to enable access to resources while simultaneously using XML filtering rules to control entry into the CMS Processing Environment. By providing a hardened intermediary, the XML security gateway protects the internal web service resources from denial-of-service attacks. Using an XML security gateway as a proxy, network managers can configure simple settings on message size, frequency, and connection duration.

Consider Digitally Signing Messages

Message-level security is the key building block for end-to-end security. It provides additional security beyond that supplied by transport, and offers message integrity, message authentication, non-repudiation, and confidentiality. Message integrity with signatures ensures detection of message tampering.

WS-Security specifies mechanisms for message integrity. WS-Security elements describe how to represent and associate a cryptographic signature with specific parts of a SOAP message. This approach allows arbitrary, well-formed fragments of the message to have separate signatures.

Web Service messages that contain sensitive data should be signed by the sender and include creation of a secure audit trail by logging each message with a signature that can be verified post transaction. Because each log entry is signed, its contents cannot be modified or altered, and the receiver gains non-repudiation protection. One may mitigate the processing-intensive nature of signing and verifying every incoming and outgoing message by using a hardware appliance rather than a software-based solution.

Encrypt Message Fields

WS-Security specifies mechanisms for encryption of SOAP messages. Encryption of message fields is rarely done in practice; however, in some instances, encrypting message fields should be considered when XML messages undergo some number of anticipated transformations prior to use.

XML Encryption (and decryption) requires fully parsing the XML transaction and then, for select message section(s), performing a set of processing-intensive XML and cryptographic encryption (decryption) operations. The advantage of using field-level encryption is that the message as a whole may be transformed without requiring re-encryption if transformations do not affect the encrypted fields. This may offer significant performance benefits. The principal disadvantage of XML encryption is its additional complexity and potential performance impact.

XML Encryption and XML digital signatures are resource intensive. Deploying both can significantly affect the performance of high-transaction applications. CMS advises mitigating this approach by using hardware (an appliance, for example) rather than a software-based solution.

Encryption of message fields must follow encryption guidance detailed in the Security Services section.

Time-Stamp Messages

CMS recommends using a Network Time Protocol (NTP) facility to synchronize network nodes to a single, authoritative time-source reference within the CMS enterprise and with CMS stakeholders. An NTP facility augments non-repudiation capabilities when used with XML Digital Signatures.

Consider SOAP Web Services Security

Web Services Security (WSS) is an extension to SOAP to enable declarative security. WSS identifies agreed-upon specifications and policies for message integrity, authentication, confidentiality, and non-repudiation. Because WSS supports message-level security, it can provide end-to-end security that is not possible by using TLS. TLS can only support point-to-point security.

The general Web Services Security model supports such security modes as identity-based authorization, access control lists, and capabilities-based authorization. This model for security tokens allows use of existing technologies, including:

  • X.509 public key certificates (recommended practice at CMS)
  • SAML Assertions (recommended practice at CMS)
  • Kerberos shared-secret tickets
  • Password digests or hashes (generally discouraged at CMS)

Various approaches may be used to create security tokens. A Web Service may use a security token from local information, or a security token may be retrieved from specialized services such as an X.509 certificate authority or a Kerberos domain controller. The preferred methodology is using SAML, an XML-based framework for exchanging security information expressed in the form of assertions about subjects, where a subject is an entity (either human or computer) that has an identity in some security domain.

REST-Specific Web Service Security

REST leverages the security infrastructure of the Web, which means relying on HTTPS and TLS for transport-level security. Typically, message-level encryption is not used. If required, the service provider must document the steps to encrypt and decrypt messages using available technologies. At the time of this publication, there are no interoperable standards for JSON messages.

CMS has adopted the National Security Agency (NSA) Guidelines for Implementation of REST, including:

  1. Use PKI and HTTPS (HTTP with TLS).
  2. Sensitive data should never be cached or put into a visible URI.
  3. Use an authorization service.
  4. Code as if protecting the application.

Like SOAP Web Services, REST-based services must validate the syntax and semantics of every service request.

Other security-related best practices (some of which are repeated from REST API Best Practices ) include:

  1. Use HTTP error codes but be cautious about revealing internal structures. Do not send stack traces back as part of error and exception handling. (please refer to CMS ARS Security Control SI-11 – “Error Handling”)
  2. Use JSON web tokens (or equivalent).
  3. Support Single Sign-On.
  4. Consider using API Keys.
  5. API endpoints use HTTP Accept Headers to determine response media type.
  6. Use CORS for security, not JSON with Padding.
  7. Provide mechanism to throttle API.

Operational (and Future) Considerations of Web Service Management

Operational management policies related to CMS Web Services should encompass, at a minimum, the following information:

  • Clearly defined terminologies and target metrics
  • Methodologies for measuring the metrics to ensure consistent results
  • Service-level prioritization to ensure that the most important services run first

Service management is critical to maintaining a structured life cycle and assuring the desired reliability and continuity to Web Services consumers. The propensity to inject multiple layers of abstraction creates one recurring complexity. The greater the number of layers, the less likely that all layers work together when combined and the harder to validate behaviors of other layers that may affect the one tested. Effectively implemented service management will balance the perceived needs for abstraction against its risks.

Change and Release Management

Any proposed changes to the Web Service must be compatible with the CMS Change Control Process described in the Configuration Management section of this chapter. Web Service deployment must be compatible with the Release Management guidance.

Monitoring

CMS requires the use of mediation services to cluster available Web Services. Service developers and operators can use mediation service monitoring to collect service usage statistics. A mediation service usually contains a logging mechanism for producing usage statistics. Development teams should have access to this information to identify performance bottlenecks, monitor run-time performance, and drive the decision to scale out a service. Service monitoring can also establish a security baseline and support transaction monitoring.

Quality of Service (QoS) is a target measurement of availability and throughput (performance), which the Web Service provider should monitor and manage. Validation at run-time ensures that the structure and content of a message payload comply with defined service interfaces. Security profiles enforce authentication and authorization when accessing Web Services. Rules for compliance may be maintained in the repository, while policies can be assigned to endpoints defined in an ESB. Invoking an endpoint or triggering an event can then invoke the rules for governance.

Web Service providers are encouraged to report metrics on service usage. Tracking usage will identify the most popular services within the evolving Web Services enterprise, which is essential to managing costs.

API Management

CMS recognizes that support for API and mobile applications at scale will require consideration of API management platforms as Web Services interfaces are exposed to support such applications. These management platforms allow for:

  • Developer, Application, and user registration
  • API Key Management
  • Traffic Management and throttling of application requests (by application, by user, or other criteria)
  • Monitoring across APIs (including performance and SLA conformance)
  • Security scanning and filtering
  • Publishing interface descriptions to Internet-based service consumers
  • Enforcement of naming and protocol conventions
  • API life-cycle management

CMS data poses additional challenges with respect to data use agreements, data sharing rules, intellectual property, and API rules of engagement. While these are not purely technical considerations, they need to be part of an API management implementation.

In order to encourage adoptions of CMS APIs, care must be taken to foster an open environment and community among API users. Thus, a big part of API success is developer advocacy and community management. API management platforms should also include support for developer productivity, such as API documentation, Software Development Kits (SDK), mock / testing APIs, code generators, sample code, and tutorials. As part of community engagement, it is also important to publish API roadmaps. Note that many of these engagement strategies also apply to Open Source software development.

API Key Management must encompass the full life-cycle management of digital keys in a manner that ensures the confidentiality, integrity, and availability (CIA) of the keys, from registration to creation, assignment, expiration, revocation, and destruction, as well as assuring that keys are assigned uniquely (per CMS ARS Security Control IA-02) so that CMS API management tools can reliably control access to CMS APIs. API keys are not the only way to identify the source of an API call, but they are among the most specific.

System Maintainers must maintain security (CIA) over API keys and secure them as credentials. API keys must have expiration dates and be rotated regularly. When personnel are terminated, API keys under their control must be revoked (please refer to CMS ARS Security Control PS-04). Different promotion environments (such as development, test, or production) must have different API keys.

Future Considerations

Enterprise Web Service Registry

CMS requires that service developers write service descriptions with appropriate content and clarity to ensure that consumers—and especially consumers outside the developer’s domain—can easily understand the service description. CMS will expect that service developers provide a full explanation of all service inputs, outputs, and schemas. The developer should take care to optimize the description for search mechanisms by including all relevant keywords that accurately describe the purpose and functionality of the service.

Support for Mobile Reference Architecture

Mobile platforms will require additional capabilities to publish, manage, and protect CMS’s Web Services from denial-of-service attacks as well as granular security capabilities. If CMS possesses appropriate security capabilities, it can throttle the use of services according to various criteria.

In addition, CMS will consider whether to issue guidance on the use of mobile authentication frameworks such as OAuth (open standard to authorization).

Web-Based UI Services

Introduction

Purpose

This Web-based UI Services chapter defines the CMS’s enterprise-wide initiative to provide a consistent, accessible, and productive user interface (UI) for all users of CMS websites, portals, and other web-based presentation mechanisms.

Scope

The scope of this chapter is limited to Web-based user interfaces made available to CMS employees, contractors, constituents, and/or other business partners. This includes user interfaces intended to be accessed via personal computers, smart or web-enabled phones, network terminals, mobile devices, tablets, and accessibility systems such as Braille or voice browsers.

The chapter does not include:

  • IBM 3270 terminal user interfaces
  • Microsoft Windows® desktop-based user interfaces
  • Specialized and/or embedded hardware user interfaces
  • Mobile applications (such as iPhone® / Android® “apps”) This refers to locally executed, compiled applications for mobile devices. Mobile-friendly web pages must still meet the business rules presented in this CMS TRA chapter.

Other Web Design Guidance Resources

CMS websites must comply with the guidance resources for web design and user experience published by CMS and HHS at the following web sites:

These resources provide guidance beyond the scope of this document and include topics such as Responsive Design and CMS Branding.

Impact to COTS

Web-based COTS software packages must meet the business rules specified in this chapter. All software packages must be acquired and customized with these rules in mind.

Business Goals

This chapter aims to satisfy the following business goals:

  • Goal 1: Create fully accessible web portals that are designed to empower consumers by enabling them to find relevant information quickly.
  • Goal 2: Inspire consumers’ confidence and trust in health information technology by ensuring the privacy and security of content and services accessed through portals.
  • Goal 3: Simplify access to health information technology for all users.
  • Goal 4: Promote collaboration and information exchange between CMS components that will minimize redundant efforts and result in superior application integration in portals.
  • Goal 5: Modernize the Agency’s websites using newer, CMS-approved technologies that improve their look and feel, marketability, adoption, and ease of use.

Expected Results

Implementing the user-centered design guidelines in this chapter will lead to the following expected results:

  • CMS contractors will be able to design, develop, and maintain a consistent look and feel and branding of web portals.
  • CMS contractors will be able to deploy easy-to-navigate web portals after completing appropriate testing.
  • Web portals and health IT systems created by CMS contractors will see increased marketability and adoption rates.
  • Users’ and visitors’ needs will benefit from centralized portals that minimize redundant logons and data entries.

Expected Benefits

In addition to the expected results, following these guidelines may also produce benefits that can be measured using the American Customer Satisfaction Index or the Citizen Satisfaction Measurement from CFI Group for Government:

  • By carefully designing for the needs of a diverse user community from the beginning, CMS can increase users’ satisfaction by reducing disruptions from unanticipated upgrades.
  • By keeping in mind that stakeholders include consumers, providers, beneficiaries, contractors, and CMS staff, developers can design systems to satisfy stakeholders’ needs and reduce stakeholders’ learning and training time.
  • By designing attractive and intuitive systems, CMS will enhance its staff’s productivity.

Business Rules

CMS has adopted the following business rules for web-based user interfaces within the CMS environments.

User Experience

This topic presents business rules for building web-based user interfaces (UX).

BR-UX-1: Ensure Usability and Accessibility

The Agency must ensure that all CMS websites are accessible to stakeholders both with and without disabilities. To achieve this, website developers must be aware of and observe the following regulations that require testing to ensure accessibility compliance.

Rationale:

  • Section 508 of the Rehabilitation Act is a federal policy that websites must comply with. Those developing websites for federal agencies, including HHS and CMS, must ensure that all employees and members of the public with disabilities are able to access these websites with ease. Additionally, consult both the CMS Section 508 and HHS versions of 508 guidance.
  • Section 504 of the Rehabilitation Act (also known as Reasonable Accommodation) is defined as “…making existing facilities used by employees readily accessible to and usable by individuals with disabilities…” For additional details, consult HHS’s guidance.
  • According to the Plain Writing Act of 2010, government agencies are now required by law to design websites using plain language so that the public can easily understand the communication. Guidelines are also available from the Federal Plain Language website.
  • According to the Government Paperwork Elimination Act, websites must use electronic forms, electronic filing, and electronic signatures to conduct official business with the public.
  • Testing must be conducted to ensure that websites comply with Section 508 and are accessible. HHS approves the Accenture Digital Diagnostics compliance monitor and Adobe Acrobat Professional 8.0 or higher for testing compliance. The World Wide Web Consortium also provides a list of tools for testing compliance.

Please refer to U.S. Web Design System , Digital.gov, and the following for more details.

BR-UX-2: Collect Feedback

CMS must enable stakeholders and the public to provide feedback about their customer experiences.

Rationale:

  • Online surveys are one mechanism for collecting feedback (e.g., Customer Satisfaction Survey available at: http://www.opm.gov/surveys/services/ Customer.asp). By emailing surveys, pop-up, or web-based surveys, CMS can collect feedback to improve training, simplify processes, and establish metrics. Another alternative is to conduct regular usability testing to collect customer feedback. As digital.gov/topics/usability indicates, a formal usability laboratory is not always necessary and testing could be performed remotely.
  • Implementing Executive Order 13571, “Streamlining Service Delivery and Improving Customer Service” (April 2011), and Office of Management and Budget (OMB) Memorandum M-11-24 requires agencies to set service standards and use customer feedback to improve the customer experience.  Refer to the American Customer Satisfaction Index for seeking feedback (http://www.theacsi.org/).

BR-UX-3: Provide Multilingual Capability

CMS must provide multilingual capability when foreign language content (i.e., non-English language content) is available. HHS prescribes that links to foreign language materials must be presented as a link written in that language (i.e., En Español, not In Spanish). Another option is to have a clickable En Español option. Thus, when the user clicks on it, the website is made available in another language, as in the following example from HHS:

  • Healthy Heart
  • Also available en Español, Français, Tiếng Việt

Rationale:

BR-UX-4: Ensure Privacy and Security

CMS must ensure its stakeholders’ privacy and security on all Agency websites. This includes complying with the Health Insurance Portability and Accountability Act of 1996 (HIPAA). Although CMS allows the collection of user information, no Personally Identifiable Information can be collected from simply browsing CMS websites. CMS websites must also comply with the CMS Information Security Acceptable Risk Safeguards (ARS).

Specific related ARS controls include PT-2 - Authority to Process Personally Identifiable Information, PT-3 - Personally Identifiable Information Processing Purposes, PT-4 - Consent, PT-5 - Privacy Notice, and PT-6 - System of Records Notice.

Rationale:

Navigation

This topic focuses on low-level implementation details for building usable and accessible web-based user interfaces.

BR-UI-5: Skip Navigation

CMS sites must allow users to skip repetitive navigational links (i.e., skip navigation) when they use screen readers and have keyboard access. This allows users to jump directly into a site’s main content. For implementation details, follow Digital.gov’s Accessibility for front-end developers for CSS and HTML implementations, as well as WebAIM.org’sSkip Navigation Links.

Rationale:

Please refer to https://www.hhs.gov/web/policies-and-standards/index.html, as well as GSA’s Section 508 ICT Testing Baseline for Web and Web-based Intranet and Internet Information and Applications (1194.22) for additional details.

BR-UI-6: Provide a Home Page Link

CMS sites must provide a home page link in text format on every page associated with a specific website. HHS prescribes that the home page link must be placed at the top of the main (left) navigation panel. The use of “home” or “[Site] Home” as primary text is the standard. If graphics are used as a link to the home page, an accompanying text link must be present on the page. If the site has a unique logo or header, those elements should be linked to the home page as a supplement to the text link.

Rationale:

Please refer to https://www.hhs.gov/web/policies-and-standards/index.html for implementation details.

BR-UI-7: No Frames

CMS sites must not use frames. According to HHS, frames are meta-documents that enable the display of multiple documents within a single browser window, but frames could lead to accessibility and navigation issues for screen reader users.

Rationale:

HHS discusses a set of exceptions for using frames. Please refer to https://www.hhs.gov/web/policies-and-standards/index.html for additional information.

BR-UI-8: Keyboard and Mouse

To ensure accessible navigation, CMS websites must provide keyboard actions as alternatives to mouse or gesture events for interacting with the website.

Rationale:

A keyboard event handler should be provided that executes the same function as the mouse event handler, as shown in Device Handler Correspondences. Please refer to W3C Client-side Scripting Techniques SCR2: Using redundant keyboard and mouse event handlers and SCR20: Using both keyboard and other device-specific functions for implementation details using JavaScript and HTML.

Device Handler Correspondences
Mouse EventKeyboard Event
mousedownkeydown
mouseupkeyup
clickkeypress
mouseoverfocus
mouseoutblur
dblclickN/A
mousemoveN/A

Consult Apple’s Safari Web Content Guide for mapping multi-touch gestures to equivalent mouse events or Windows User Experience Interaction Guidelines for mapping gestures to equivalent mouse and keyboard events.

Privacy and Information Security

This topic focuses on building web-based user interfaces that incorporate privacy protection and security measures.

BR-UI-9: Phishing and Redirection Prevention

Links must be provided to the CMS-recommended State Fraud and Abuse Contact Report and HHS’s Office of the Inspector General’s Report Fraud Form.

Rationale:

CMS websites must minimize phishing attacks, which often involve improper redirection and unnecessary exposure of personal email addresses on websites. Please refer to Common Weakness Enumeration (CWE™) CWE-601: URL Redirection to Untrusted Site for implementation details to prevent redirecting users. 

BR-UI-10: Cross-Site Request Forgery Prevention

CMS websites must use effective methods to protect against Cross-Site Request Forgery (CSRF).

OWASP suggests these methods for more effectively preventing CSRF attacks:

  • A challenge-response, such as Completely Automated Public Turing Test to Tell Computers and Humans Apart (CAPTCHA), re-authentication through password, and one-time tokens are possible defenses against CSRF. As W3C discusses, CAPTCHA is not 508-compatible and could create inaccessibility. W3C suggests other approaches (http://www.w3.org/TR/turingtest/#solutions) such as logic puzzles or auditory CAPTCHA to replace traditional CAPTCHA.
  • Warn users that they should not allow their browsers to save usernames or passwords. Also, do not enable CMS websites to “remember” users’ login information.
  • Warn users that they should not use the same browser to access sensitive applications and to browse the Internet freely (tabbed browsing).
  • Advise users that the use of plugins such as No-Script makes POST based CSRF vulnerabilities difficult to exploit.

According to OWASP, the following four methods are neither appropriate nor effective for preventing CSRF:

  • Session identifiers, like cookies, are simply used by applications to associate requests with specific sessions. The session identifier does not verify that the user intended to submit the request.
  • POST requests are often considered to limit CSRF attacks because attackers cannot construct a malicious link, and therefore, a CSRF attack cannot be executed. Yet there are methods where an attacker can trick a victim into submitting a forged POST request, such as a simple form hosted in an attacker’s website with hidden values.
  • Multi-step transactions are not adequate to prevent CSRF. If an attacker can predict or deduce each step of the completed transaction, then CSRF is possible.
  • Rewriting the session ID in a URL might be a useful CSRF prevention technique because the attacker cannot guess the victim’s session ID; however, the user’s credentials could be exposed over the URL.

Rationale:

CSRF attacks make a target system perform a function (funds transfer, form submission, etc.) via the target’s browser without knowledge of the target user, at least until the unauthorized function has been committed.

BR-UI-12: Authentication

CMS websites may allow anonymous access as follows:

The organization permits actions to be performed without identification and authentication consistent with CMS mission and business functions and with documentation and supporting rationale in the security plan for the system. (ARS AC-14)

Otherwise, a CMS website must authenticate all users according to the policies established in the CMS ARS, which may include HSPD-12 compliance.

Rationale:

Authentication is the process of verifying that individuals are who they claim to be. Authentication is commonly performed by submitting a user name or ID and one or more items of private information that only a given user should know.

To prevent authentication attacks, observe and practice the following:

  • Implement authentication management practices identified in CyberGeek Password Requirements for password complexity and other factors
  • Utilize multi-factor authentication
  • Serve the initial login page via secured connection, e.g., secure socket layer (SSL)
  • Implement account lock-out to prevent brute-force attack

BR-UI-13: Session Management

CMS websites must implement secure sessions.

Rationale:

Session management is a process by which a server remembers how to react to subsequent requests throughout a transaction. A session identifier maintains sessions on the server that can be passed backward and forward between the client and server when transmitting and receiving requests.

To enhance session management, observe and practice the following server-side defenses:

  • Session ID names used by the most common web application development frameworks can be easily identified for potential exploit. This includes PHPSESSID (PHP), JSESSIONID (J2EE), CFID & CFTOKEN (ColdFusion), ASP.NET_SessionId (ASP .NET). Instead of using one of these, change the default session ID name of the web development framework to a generic name, such as “id.”
  • Consider installing a web policy agent that a web server can call to make policy decisions when a client requests access to a protected resource.
  • Use a session ID length of at least 128 bits (16 bytes) to prevent brute force attacks.
  • Session ID must be unpredictable (random enough) to prevent guessing attacks.
  • Use latest web development frameworks, such as J2EE, ASP .NET, PHP, for session management instead of building a home-made one from scratch, because these frameworks are used worldwide on multiple Web environments and have been tested by the web application security and development communities.
  • The most common session ID generation mechanism in use today is a strict one in which web applications will only accept session ID values that have been previously generated by the web application.
  • Set expiration timeouts for every session. Insufficient session expiration by the web application increases the chance for attackers to reuse a valid session ID and hijack the associated session. The session expiration timeout values must be set according to the purpose and nature of the web application and also balance security and usability so that users can comfortably complete the operations within the web application.
  • Web applications must provide a visible and easily accessible logout (logoff, exit, or close session) button that is available on the web application header or menu. The button should be reachable from every web application resource and page, to ensure that users can manually close the session at any time.

To further enhance session management protection, observe and practice the following client-side defenses:

  • Enforce login timeout on the login page and notify users that the maximum amount of time to log in has passed. This allows the renewal of the session ID for authentication and avoids scenarios where a previously used (or manually set) session ID is reused.
  • When users choose to close a web browser tab or window, close the current session for users before closing the web browser. This helps users close a session if they forget to click a logout button.

BR-UI-14: Protecting PHI

CMS websites must secure Protected Health Information.

Mitigations may include limiting what PHI elements are displayed to a user, or displaying only partial elements such as a partial Social Security Number (SSN), a partial email address, or partial telephone number.

Rationale:

The HHS HIPAA Privacy Rule, prescribes that 18 elements are considered PHI (The HIPPA Privacy Rule). These elements must be removed or generalized.

For a summary, please consult Summary of the HIPAA Privacy Rule.

BR-UI-15: User Identifiers

CMS websites must meet the CMS ARS standard for user identifiers. If permitted, use easy to remember and recognizable user identification for identifying and tracking user identity.

Rationale:

When users have the freedom to choose data that are meaningful to them as user identification, that identifying information can be recalled more easily. Consider allowing users to use their first and last initials. Another possibility is to allow email addresses as user identification.

For additional details, please refer to Chapter 4 of HHS’s HIPAA Security Series, Security Standards: Technical Safeguards.

BR-UI-16: Input Validation

CMS websites must validate all user input at a minimum on the server side, although validating at both the client and server side is recommended because it allows better system performance and quicker response for the users.

Rationale:

According to the CWE:

Use an allowlist of acceptable inputs that strictly conform to specifications. Reject any input that does not strictly conform to specifications or transform it into something that does. Do not rely exclusively on looking for malicious or malformed inputs (i.e., do not rely on a denylist). However, denylists can be useful for detecting potential attacks or determining which inputs are so malformed that they should be rejected outright. When performing input validation, consider all potentially relevant properties, including length, type of input, the full range of acceptable values, missing or extra inputs, syntax, consistency across related fields, and conformance to business rules.

For a counter example of improper user input validation and ways to mitigate this issue, please refer to CWE-20: Improper Input Validation and OWASP Improper Data Validation.

For information about other related issues, software errors, or attacks, please refer to CWE/SANS TOP 25 Most Dangerous Software Errors.

BR-UI-17: (Deprecated after TRA 2017R1): Information Security

BR-UI-18: Need to Disclose

A website should disclose only the minimum sensitive information necessary for authorized users to perform the tasks required, as described in the following examples:

  • When users log in, they should not see their entire Social Security Number unless absolutely necessary. This will prevent shoulder surfing, for example.
  • When users check a claim, they should not see more information about that claim than is necessary for the task at hand.
  • When users file a change of address, they do not need to see their prior address in its entirety.

Layout, Style, and Appearance

These rules focus on building web-based user interfaces following best practices to ensure compatibility with the broad range of browsers and devices that users may employ.

BR-UI-19: Color Contrast Ratio

CMS sites must ensure a color contrast ratio of 4.5:1.

To calculate the color contrast ratio, follow this equation:

Rationale:

According to the W3C, a contrast ratio of 3:1 is the minimum level for standard text and vision; however, to satisfy Section 508 accessibility, a 4.5:1 is necessary to account for those who have low visual acuity, congenital or acquired color deficiencies, or the loss of contrast sensitivity due to aging.

BR-UI-20: (Deprecated after TRA 2017R1): Standard Colors

BR-UI-21: (Deprecated after TRA 2017R1): Deprecated Elements

BR-UI-22: (Deprecated after TRA 2017R1): CSS Version

BR-UI-23: (Deprecated after TRA 2017R1): HTML Version

BR-UI-24: (Deprecated after TRA 2017R1): JavaScript

BR-UI-25: Scripts Compatibility

In order to comply with Section 508 standards, CMS websites must be ARIA compliant and ensure that the use of JavaScript does not interfere with the use of screen readers and other assistive technologies.

Rationale:

In order to keep CMS websites accessible by the general public, it is important that they function with most browsers and assistive technologies.

BR-UI-26: (Deprecated after TRA 2017R1): Logo

Other Guiding Principles

This topic presents other subjects not directly addressed in the foregoing UX business rules.

Mobile Interfaces

Mobile Web applications must not store PII local to the mobile device in a persistent manner (although temporary, in-memory use is permissible). This applies, at a minimum, to cached contents, cookies, and HTML5 local data stores.

Best Practices

CMS provides general guidance via a set of websites policies and a handbook. HHS also provides a set of website policies and suggests following the Associated Press Stylebook.

The Managing Content section on Howto.gov also discusses best practices in managing the content of a government website. Table - Summary of Howto.gov Best Practices for Managing Content of Federal Websites highlights and summarizes that discussion. For further information, please refer to digital.gov/resources/checklist-of-requirements-for-federal-digital-services.

Table - Summary of Howto.gov Best Practices for Managing Content of Federal Websites
Best PracticeDescription
Task-drivenUse a task-driven perspective to identify the website’s top so that CMS’s stakeholders can accomplish their work easily and quickly. For a road map to perform task-driven mission identification, please refer to digital.gov.
Current ContentKeeping the website content up to date is critical for providing accurate information. By establishing a content review process, using date stamps, managing links, and following records management requirements, out-of-date information can be minimized.
Employee Information on Public WebsiteFocus public-facing websites on information designed for the public. Use only intranets or extranets to provide information for your employees.
“Contact Us” PageProvide a page called “Contact Us" or “Contact CMS.” The public expects to see this link on every page of your website, usually in the header or footer. This allows the public to ask questions, get additional information, or report problems. If that is not feasible, at least provide a link from your homepage and every major entry point.
“About Us” Page

The page should be placed at the top navigation bar or banner area of the website. At a minimum, include the following pieces of information:

  • Full name of CMS
  • Name of CMS head and other key staff, as appropriate
  • Contact information
  • Basic information about the parent and subsidiary organizations and regional and field offices, as appropriate
  • A description of CMS’s mission, including statutory authority
  • Strategic plan
  • Organizational structure (such as an organizational chart)
“Frequently Asked Questions” Page

People do not always recognize the acronym “FAQ.” “Frequently Asked Questions,” spelled out, is the most common terminology. Consider the following pieces of information as clues to creating a FAQ page:

  • Look at email, phone calls, and letters from the public
  • Review top search terms
  • Talk to the people who answer phones and mail
  • Look at statistics
  • Look at information requested under the Freedom of Information Act
Site MapHave a site map or A-Z subject index to serve as a “table of contents.” This can give visitors a quick and easy way to find what they need.
Forms and PublicationsOffer ease of access to online forms and other publications by identifying most commonly requested and commonly used forms and publications. Also link to federal portals such as Search for government forms on USA.gov and GovInfo.gov to help visitors find other related materials.
Job and Employment InformationJob seekers and curious citizens want to know basic information such as what jobs are available, how to apply, and what it is like to work at CMS. Include these types of information on CMS’s website. Also link to USA Jobs so visitors can find information about jobs available across the federal government.
Grant and Contract InformationWhen appropriate, CMS should provide grants information or contracting on its website and link to federal portals such as grants.gov.
External Linking

Notify visitors when they are leaving CMS websites for any non-federal government websites. Preferred methods for notifying visitors include:

  • Placing an icon next to the link
  • Identifying the destination website in the link text or description itself
  • Inserting an intercepting page that displays the notification, after the user selects the link
  • Displaying all non-federal links in a separate listing from federal links
Cross-agency PortalsLinks to cross-agency websites (portals) can supplement or eliminate the need to create (or recreate) information. These links to other government sites can guide visitors to additional resources they might not otherwise find. This is especially important for federal public websites because many visitors do not know the organizational structure of CMS and may need help finding information and services. Such links might also help visitors get to the most authoritative, current source for information.
USA.gov

Linking to USA.gov is a requirement of Section 204 of the E-Government Act of 2002. Visitors to the CMS website may become frustrated if they are looking for government information and services that are provided by another agency. Linking to USA.gov will help visitors who need information from different agencies or who need a "starting point" to find government information. Such linking also promotes seamless government by allowing visitors to access the vast amount of information from across government without having to know which agency sponsors the information. The following guidance will accomplish this linking:

  • Text links should say “USA.gov”.
  • For Spanish, use “GobiernoUSA.gov”.

Logos and linking instructions can be found at https://www.usa.gov/link-to-us

  • Use the following text for link descriptions and alt text:
    • English: “USA.gov: The U.S. Government's Official Web Portal”
    • Spanish: “GobiernoUSA.gov, el portal official del Gobierno de los Estados Unidos”
Metadata

At a minimum, use these six metadata elements on the homepage and all major entry points:

  • dc.title
  • dc.description
  • dc.creator (the content owner; this should be the name of the organization)
  • dc.date.created (original creation date)
  • dc.date.reviewed
  • dc.language

These six suggested metadata elements are based on internationally recognized Dublin Core standards.

 

Additional Notes:

Please refer to HHS Web Policies and http://www.hhs.gov/web/policies/webstyle.html for HHS’s Dos and Don’ts.

For a list of HHS-approved multimedia tools and formats, please refer to Plug-ins Used by HHS.

Open Source Software

Open Source Introduction, Overview, and Strategy

Introduction

Open Source Software (OSS) is software that is freely licensed to the public to use, copy, modify, and distribute. The source code is openly shared to encourage people to voluntarily improve the design of the software. CMS has been an active consumer of and contributor to OSS on several IT projects. This chapter provides guidelines for the CMS project teams that wish to use and contribute to OSS libraries and packaged OSS for their internal consumption or for development of new and custom software.

Several CMS business units and offices have been actively developing open source projects as part of IT modernization efforts at the agency. CMS has many active open source communities, such as its Developer APIs like Blue Button, the CMS Design System, and hundreds of other repos across many CMS organizations. CMS has embraced Open Source development projects and is looking forward to releasing software to promote its reuse.This policy will help CMS achieve the goals outlined in the OMB directive M-16-21 for Federal agencies, as well as in the 2024 SHARE IT Act.

CMS launched its Open Source Software Policy in 2018 that guides IT Application Development Organizations (ADOs) and Employees that produce software for CMS’s mission-critical programs. The policy is located at github.com/CMSgov/cms-open-source-policy. This CMS policy is a living document, and changes to this policy are handled via issues and pull requests in the CMS GitHub repository.

Scope

The guidance in this chapter is limited to using OSS within the CMS environment and does not address practices that are generally applicable to software engineering efforts. Moreover, the guidance complements and incorporates CMS’s existing policies, standards, and procedures, including those described in other parts of the CMS TRA, such as Cybersecurity Security Services. The TRB manages and approves the use of OSS licenses in accordance with TRA guidance and its prescribed function within the CMS TLC.

Additional references used in this chapter:

Wikipedia, Free and open-source software

Digital.gov, Requirements for achieving efficiency, transparency, and innovation through reusable and open source software

Acquisition.gov, FAR Subpart 27.4 - Rights in Data and Copyrights

OWASP, Vulnerability Scanning Tools

 

Overview

OSS Overview

Open Source Software is computer software whose source code is developed in an open and collaborative way and made available (in a human-readable format) with a copyright license that complies with the Open Source Initiative's Open Source Definition (OSD). The Open Source Initiative (OSI) established the following criteria for OSS, all of which must be met to comply with OSD:

  • Software is free to be re-distributed
  • Source code is available with software
  • Derived work is allowed
  • Integrity of author’s source code is maintained
  • Contains no discrimination against persons or groups
  • Contains no discrimination against fields of endeavor
  • License must be distributed
  • License must not be specific to a product
  • License must not restrict other software
  • License must be technology neutral

The backbone of OSS is community sharing of ideas and collaborative development of code to serve a common purpose. Some of the key enablers of OSS development are version control systems, mailing lists, wikis, and blogs that help developers collaborate to develop code. Copyright licenses make the resulting software code available to users to use, change, and improve the software, and to redistribute it in modified or unmodified forms.

History of OSS at CMS

CMS is interested in two basic types of OSS—open source frameworks/libraries and open source solutions (CMS Open Source Software (OSS) Thoughts for a Strategy, CMS, March 2010). Open source frameworks / libraries are standalone pieces of code that may have dependencies on other software and are used or embedded as part of a larger software development effort. Examples include the Spring Framework and Apache Struts, both of which are open source application frameworks for Java development, and Hibernate and myBatis, which are persistence frameworks that provide mappings between Structured Query Language databases and objects in Java.

Open source solutions are open source applications that may be part of a larger software distribution, can operate standalone, and serve a single function or set of functions. Open source solutions may come with vendor support. Examples include Linux, an open source operating system; Apache HTTP Server, an open source Web server; and Oracle Glassfish and RedHat JBoss, open source application servers.

CMS uses open source frameworks and libraries and a small set of open source solutions. Since 2014, CMS seeks to expand its use of OSS to include additional open source solutions.

Benefits of OSS within the CMS Ecosystem

OSS provides many advantages, as demonstrated in the commercial industry and the government sector. If used thoughtfully, OSS offers the potential for cost reduction, timelier delivery, and improvement of the overall capabilities of a system. Specifically, OSS can:

  • Offer considerable savings in software purchasing cost and greater efficiency in repurposing and enhancing existing software. OSS has inherent flexibility in its purpose and adds value to its use.
  • Minimize vendor lock-in through open standards and flexibility in the choice of solutions. This is evident of the cost-saving nature of OSS applications.
  • Enable transparency of software implementation to support extension of software (through supported plug-in mechanisms) while also ensuring robustness of design and implementation
  • Increase the productivity of application development through open standards to help achieve interoperability across systems
  • CISA Development Guide OSS Policy, “Publicly available source code enables continuous and broad peer review. Whether simply publishing the completed code or opening the development process, the practice of expanding the review and testing process to a wider audience—beyond the development team—ensures increased software reliability and security. Developing in the open also allows for other opinions to help adjust the direction of a product to maximize its usefulness to the community it serves.”
  • Help keep the enterprise abreast of technology developments and facilitate the early adoption of emerging technology

CMS is committed to meeting the needs of its IT development community by optimizing best practices and creating innovative approaches for rigorously evaluating and enhancing existing software development processes. CMS incorporates OSS as part of its IT development practices to realize the many benefits and capabilities of OSS.

Inbound Open Source

Considerations for Adoption of OSS

Although there are tangible benefits to OSS, the process of realizing these benefits may significantly change how CMS uses IT. For example, adopting OSS alters the Agency’s development and acquisition of software and how the Agency’s IT departments conduct business.

A common misconception about OSS is that it is cost free. The nature of cost with OSS differs from commercially licensed software. Although licensing costs may be minimal to none, CMS must consider the total cost of ownership, as various costs are involved including evaluation, maintenance, installation/ configuration, integration, and support/ operations. One important contributor to OSS cost is the absence of a “productized” solution. Many OSS products do not offer comprehensive documentation, easy installation and upgrade of utilities, friendly migration tools, and Graphical User Interface (GUI) configuration tools. Therefore, operations and maintenance (O&M) can be complex and expensive for some OSS products. To overcome this gap, the enterprise adopting an OSS solution must possess the necessary skills or have access to third-party vendors who provide IT support services for that specific OSS solution.

Another key OSS consideration is managing the adoption of OSS within the enterprise. Free software and full access to the source code of OSS present unique challenges for the enterprise. Individuals might download and install OSS without sufficient oversight. Therefore, the enterprise must carefully manage open source adoption and assess OSS for use across the enterprise, using the inbound review checklist. Criteria for assessing solutions must include the open source software’s maturity, total cost of ownership, licensing implications, alternative solution evaluations, and security & risk review. CMS requires that OSS be managed to the same standard as COTS software. The business owner is responsible to perform due diligence.

Procuring OSS

CMS focuses on OSS solutions that can be used “out of the box,” like a COTS product. This has two implications:

  • That there will be third-party vendor support for the OSS solution (e.g., to install, configure, integrate, operate, and maintain the software)
  • That the OSS solution will require no customization, extension, or modification

It is important to evaluate the total cost of ownership when adopting an OSS solution as various costs are involved including evaluation, maintenance, installation/ configuration, integration, and support/ operations. Contractors using open source frameworks and libraries are responsible for (a) ensuring that the OSS works properly within the CMS environment, and (b) that their staff is well versed in configuring and subsequently operating the software. One major consideration is that the OSS license for a package may encumber any modifications to the source code.

The contractor’s responsibility extends to monitoring the community behind the OSS product for new versions and patches, and support for older versions. Furthermore, contractors will monitor the community to ensure it continues to offer substantial support for the OSS product. The contractor should coordinate with the TRB to understand how and where CMS currently uses OSS today, and to identify lessons learned from successful implementations and resolution of issues involving OSS. This knowledge will shape the contractor’s use of OSS and may also influence CMS’s policy regarding OSS.

The TRB is responsible for approval of OSS licenses. The TRB will rely on the guidance in this chapter to assess the potential adoption of and approval for specific OSS.

Open Source Software Customization and Contribution for CMS

As CMS expands the adoption of OSS, it is possible that specific OSS will require extensions or customization for use within CMS’s environment. In some rare cases, the customization may be specific to CMS’s needs and remain internal to CMS. Because internal forks and patches are costly to manually maintain and accumulate technical debt, projects must provide detailed justification for managing OSS extension and customization exclusively within CMS.

CMS Open Source Strategy

OMB Memorandum M-16-21, Federal Source Code Policy: Achieving Efficiency, Transparency, and Innovation through Reusable and Open Source Software, encourages the use and publication of open source in the federal government.

The Source code Harmonization And Reuse in Information Technology Act or the SHARE IT Act of 2024, ensures that “custom-developed code (i.e., source code that is produced under an agency contract, funded exclusively by the federal government, or developed by federal employees as part of their official duties) and certain technical components of the code such as architecture designs and metadata are (1) owned by the agency, (2) stored at no less than one public or private repository, and (3) accessible to federal employees under certain procedures. Agency contracts for custom-development of software must acquire and exercise rights sufficient to allow government-wide access, sharing, use, and modification of any custom-developed code. The act does not apply to source code that is classified, developed primarily for use in a national security system, or developed by an element of the intelligence community. An agency's office of the chief information officer may exempt source code from being shared or made publicly accessible to protect individual privacy.”

The CMS Open Source Program Office created a detailed guide, Enabling Reuse of Federal Source Code that provides full technical documentation on the subject, mostly in the context of GitHub. Guidance for other repositories are planned.

Outbound Open Source

Contribution of Open Source to the community

CMS may consider release of internally developed software as open source, including software that is an extension to an existing OSS project or software that CMS may want to use to initiate a new OSS project.

Appropriate considerations for developing and contributing OSS to the community include defining policies and processes to release software as open source, guidance for continued development, ground rules for engaging the larger community to contribute to the project, assessment of alternative copyright licensing strategies, and publishing project metadata to support agency-wide software inventory and reuse as per the SHARE IT Act. These outbound checklists outline these processes.

Alongside the software, agencies are required to publish metadata on all custom-developed code after July 22, 2025, which is not subject to exemptions as per the 2024 SHARE IT Act. All repositories must store their metadata in a code.json file.

Employee participation in OSS projects used by CMS is often in the Government’s interest. It is typically a legitimate use of government resources when the Government uses the software in question, either directly, as a component of a larger government system, or as a component of underlying government infrastructure.

Government employees may contribute to existing OSS projects (including being the primary maintainer) as part of their official duties, so long as they consult with their supervisor first to ensure a common-sense approach for contributions that preserve Operations Security (OPSEC) and further the Government's interests. Contractors may contribute to OSS projects at government expense when the government task monitor authorizes such activity as being in the Government’s interest, subject to the scope and provisions of the contract.

Creating a separate, CMS-specific version of any OSS project, for any reason, increases support risk and should be avoided whenever possible. Creating such a CMS-specific fork creates risk by requiring separate maintenance, reducing access to software capabilities from the OSS community, and may require the agency to assume full management responsibility of the software’s source code, thereby increasing costs.

CMS encourages upstream contributions to OSS projects. Any improvements or new capabilities added to an OSS project whether it be code, commentary, bug reports, feature requests, or overall strategic direction should be contributed back to the upstream open source project via pull/ merge request, in accordance with the processes used by that project.

Open Source Licensing

For CMS projects using OSS, each CMS business owner is responsible for assuring that CMS’s use of the open source is compliant with its license. For a CMS-released OSS project, the CMS business owner is responsible for selecting an appropriate license model. In either case, the TRB is responsible for approval of OSS licenses.

Be aware of licensing issues and select a licensing model considering the difference between “permissive” and “non-permissive” licenses, warranty/ guarantee limitations, attribution, and trademark/ IP protections and its impact on the choice of license.

The CMS default LICENSE file for projects acknowledges that CMS work is in the U.S. public domain, and uses Creative Commons Zero International 1.0 (CC0) to waive copyright internationally. The Open Source Initiative maintains a list of licenses that the project team may optionally use. In addition to obtaining TRB approval, the Office of General Counsel must review any proposed open source modification or creation.

Other useful descriptions of Open Source Licenses include tl;dr Legal's verified license resource page, the Wikipedia Comparison of Free Software Licenses, and Civic Commons’ Choose an open source license.

OSS Consideration and Applicability

Project teams at CMS are requested to conduct market research and analyze commercial and other open source alternatives that may meet their business need before venturing into OSS development. To help make this decision, a project team may use the three-step software solution analysis outlined at Digital.gov’s Requirements for achieving efficiency, transparency, and innovation through reusable and open source software, and supplemented by BR-OSS-2, BR-OSS-3, BR-OSS-4, and the CMS OSS Policy.

The CMS OSS Policy will guide project teams that venture into custom code development and intend to release as OSS. If the project teams seek to release as OSS, they should consult the outbound review checklists first. If they have any further questions or need support, they can reach out to the Open Source Program Office via Slack or email. If a team has questions about contractual requirements or legal questions, they can reach out to the Office of the Acquisition & Grants Management (OAGM) to understand the contractual requirements, policies and legal issues.

CMS contractors who develop software for CMS business use are covered by the procurement clauses that assign the copyright of the CMS-funded custom designed software to CMS and prohibits the contractors from reselling it to other federal government agencies. Due to the variety of CMS IT and non-IT contracts, it is the project team’s responsibility to perform all due diligence for their specific contract in consultation with the OAGM.

Business Rules and Recommended Practices

CMS developed the following business rules to guide the adoption, selection, implementation, operation, and maintenance of OSS within CMS.

Business Rules

BR-OSS-1: Products in Use at CMS that Provide the Required Functionality Are Preferred

CMS prefers mature COTS and OSS products that have demonstrated successful implementation at or outside of CMS. The benefits must outweigh the risks of using new or unproven products. Like custom code, CMS business owners using OSS in a CMS application are responsible for its maintenance.

Rationale:

Before adopting any software product, whether OSS or otherwise, it is essential to understand the inventory of licensed COTS and OSS products within CMS’s environment. Identifying products already in successful CMS implementations that can be used in place of the proposed product under consideration will promote reuse of existing capabilities.

See the Inbound Review Checklist: Existing Products and Alternatives

BR-OSS-2: Criteria for Evaluation

Any OSS product considered for use by CMS must be evaluated against the following types of criteria (in addition to the criteria traditionally used to evaluate COTS products):

  1. Product Criteria. These address the software itself. The criteria may include the quality of the software code, quality of the software architecture, and the extent and scope of documentation.
  2. Use Criteria. These address what it takes to use the software. The criteria may include quality of end-user support, quality of packaging, flexibility of the associated copyright license, quality of project site, level of leadership behind the open source project, and vitality and momentum of the community.
  3. Integration Criteria. These criteria address what it takes to make the software work in the enterprise environment. The criteria may include integration with other products, which may include COTS, and support for standards.
  4. License Criteria. These criteria address what licensing limitations exist on the package.

Rationale:

Finding the right OSS that meets CMS’s needs requires assessing the maturity of the software and the level of support provided by the community involved in the OSS project. This inquiry includes the extent to which a support community is active, the quality and stability of the OSS implementation, the responsibilities involved in using the OSS, and the skills needed to use the software and manage risks. CMS is especially interested in the degree to which the OSS is “productized” (in comparison to COTS), which encompasses much of these three criteria. Although most OSS is only partially productized, information is available to support productization. CMS’s primary concern is assessing the burden to overcome the lack of productization and determining whether the benefits outweigh the risks in this regard. Initially, CMS will only consider OSS that has third-party vendor support for productization (Red Hat, for example, is a third-party vendor that offers an “enterprise-ready” distribution of Linux (an open source operating system platform) along with associated support, training, and consulting services).

See the Inbound Review Checklist: Software Evaluation

The table Overview of Common Criteria for OSS Evaluation summarizes the most common OSS criteria and related questions involving open source products.

Overview of Common Criteria for OSS Evaluation
OSS CriteriaPertinent Questions
Existing Products
(BR-OSS-1)
  • Approved Alternatives: Are there other OSS or COTS products in current use at CMS that can serve the same purpose as the OSS?
Maturity
(BR-OSS-2)
  • Market Share: What is the OSS’s market / industry rate of adoption?
  • Credibility: What is the position of OSS’s structured bodies and organizations?
  • Documentation / Books: Are instruction manuals comprehensive and readily available?
  • Community: Is there an active online community behind the OSS? Are there any commercial conflicts? What is the composition of the community (e.g., commercial organizations and government agencies)?
  • Source Code Quality: Does the open source project routinely use source code inspection tools to assess quality and security of the source code? Is the code well documented? Does an independent body review the code?
  • Rate of Change: What is the rate of change of the OSS? This includes the rate of code change and the rate of new releases.
  • Bug and Vulnerability Tracking: Is there a bug / vulnerability reporting and resolution system? How quickly are patches issued to resolve them?
Costs and Risks
(BR-OSS-3)
  • Integration: How well will the OSS integrate within CMS’s environment? Is customization of the OSS required?
  • Talent Pool: Does CMS have the right resources to support development and Operations & Maintenance for the OSS?
  • Prohibitions: Are there CMS or federal security policies that prohibit the use of the OSS? Does the OSS product allow implementation of CMS or federal security requirements?
  • Professional Services Availability: Is there third-party vendor support to install, configure, customize, integrate, operate, and maintain the OSS?
  • Cost Comparison: What is the cost of using and maintaining the OSS compared to a COTS solution?
Licensing
(BR-OSS-4)
  • Licensing Model: Does the copyright license associated with the OSS have restrictions limiting the redistribution of software code or restrictions on merging with proprietary code that are unreasonable for CMS?

 

BR-OSS-3: Total Cost of Ownership

CMS must determine the total cost of ownership and the risks of using open source as a solution before adopting or authorizing adoption of OSS for a CMS project.

Rationale:

It is a misconception that open source is free of cost. The nature of OSS costs is different. Although licensing costs may be minimal to none, licensing represents only a small portion of the total cost of software ownership. It is essential to factor in the following costs to arrive at the total cost of OSS ownership: See also Open Source for the Enterprise: Managing Risks, Reaping Rewards, Woods, D. and Guliani, G., O’Reilly Media, Inc. (Safari Books Online), July 2005.

  • Evaluation Costs. OSS must be installed and evaluated locally to assess its use within the enterprise.
  • Maintenance Costs. IT departments must define and implement a strategy to manage the OSS once installed, such as installing patches and upgrades.
  • Installation and Configuration Costs. Installation and configuration of OSS can be time consuming. Effective installation and configuration depend heavily on the organization’s skill level in using open source development tools, system administration, and operations.
  • Integration and Customization Costs. Open source projects may not be designed to operate within enterprise IT infrastructure. IT departments may have to integrate and customize OSS to suit their environments.
  • Operations and Support Costs. On the surface, OSS operations and support costs do not appear to differ from COTS costs; however, OSS’s lack of productization may introduce additional burdens on individual organizations.

 

Because CMS’s focus is on OSS that has third-party vendor support, the vendor will perform many of the functions associated with these costs, as in the case of many COTS products. Whether these functions are outsourced to a third-party vendor or performed within CMS, they entail additional costs that must be considered as part of the total cost of ownership for OSS. In any instance when OSS is considered acceptable, it will be CMS’s responsibility to ensure that all subsequent government estimates and/or contractor quotes clearly show a detailed breakout of costs associated with using the OSS.

CMS must weigh the support costs of evaluation, installation and configuration, integration and customization, and O&M against the benefits of using specific OSS. In the commercial sector, organizations define value primarily as an increase in revenue or cost savings. For CMS, reducing cost is relevant and a valid driver for using OSS. In addition, the timeliness of delivering functionality for CMS use may be equally important. Therefore, it is essential to compare the cost of OSS versus COTS as well as the potential agility offered by the OSS.

In addition, CMS must evaluate the risks associated with the OSS product, determine whether the risks are acceptable, and establish a clear line of accountability to manage and mitigate each accepted risk. This includes identifying any CMS or federal security policies that prohibit the use of the specific OSS product under consideration, determining whether the OSS product allows implementation of required CMS or federal security requirements, understanding any hardware / software dependencies, and ensuring integration of the OSS product within CMS’s environment.

See the Inbound Review Checklist: Total Cost of Ownership

BR-OSS-4: License Compatibility for Using OSS

When selecting OSS, CMS must ensure that the associated OSS copyright license is consistent with CMS’s objectives and environment.

Rationale:

CMS will consider whether there are any restrictions on linking the open source code with code that follows a different license and whether there are any restrictions on distribution of changed source code.

See the Inbound Review Checklist: License Compatibility

BR-OSS-5: Use OSS Built from a Controlled Source

CMS requires the construction of open source products from configuration-managed source code before introduction into the CMS Processing Environment. This ensures that the binary versions of software do correspond to known source code versions. It also provides CMS the opportunity, but not the obligation, to perform a combination of manual and automated testing of the source code to assess the security of the code.

BR-OSS-6: Binary Package Management Is Mandatory

To control proliferation of open source binary libraries, CMS requires that the OSS built from controlled sources is (1) stored in a package management system or repository, and (2) all CMS products built from these packages use this repository as the source of record (not Internet-based repositories).

Rationale:

This binary pack management affects build systems like Apache Maven and Apache Ivy that automatically import binary packages when referenced as dependencies.

BR-OSS-7: (Deprecated after TRA 2017R1): Modification of Open Source Is Prohibited

BR-OSS-8: (Deprecated after TRA 2017R1): Use Only Published Final Releases of Project Packages in Testing Validation and Production Deployment Systems

BR-OSS-10: CMS OSS Code Released as CMS-Managed Code Requires a Governance and Support Model

The CMS project team must have adequate resources to manage the pull requests from the community and incorporate requested enhancements. The project team should review the CMS OSS policies in detail and create a governance and support model to manage and sustain the open source project.

BR-OSS-11: CMS OSS Code Released as Unmanaged Code Must Be Identified as Such

If the CMS project team chooses to release unmanaged code through an OSS code repository, then the repository should have documentation and readme files to indicate that CMS is releasing the code for public use but does not have resources to incorporate any requested enhancements.

BR-OSS-12: CMS-Released OSS Code Must Include Automated Unit Tests, Build Scripts and Be Checked for Software Vulnerabilities

The software should contain automated unit tests, build scripts and it should be checked for software vulnerabilities, code quality and code coverage using available standard CMS tools. The checks for software vulnerabilities should be included in the automated build.

Ensure that the software code is adequately peer reviewed and is free of security vulnerabilities that can be exploited by malicious actors. There are several Static & Dynamic Application Security Testing (SAST / DAST) commercial and open source tools available to look for security vulnerabilities such as Cross-site scripting, SQL Injection, Command Injection, Path Traversal and insecure server configuration. The project teams should check the effectiveness of DAST tools by checking out the OWASP Benchmark project, which scientifically measures the effectiveness of all types of vulnerability detection tools, including DAST.

Until the software code is adequately reviewed, it should be either 1) maintained in an internal code repository that replicates the intended public repository or 2) checked by publicly available services providing the same functions on all code check-ins to the public source code repository.

BR-OSS-13: CMS-Released OSS Code Must Include Documentation Accessible to the Open Source Community

The documentation should be accessible to the Open Source community, using repository templates and hygiene best practices as recommended in the CMS repo-scaffolder project.

BR-OSS-14: Use the CMS’s External GitHub Repository and Code.gov for CMS-Released OSS Code

Put the source code in CMS’s external GitHub repository (github.com/CMSgov) and follow a consistent naming convention and a unique prefix for code repositories.

BR-OSS-15: Publish project metadata in an open source repository to support agency software inventory

Store metadata on the project in a code.json file and add to the repository. This metadata supports updates to agency software inventory initiatives such as CMS’ metrics website and the enterprise code.gov repository to encourage discovery and dissemination and reuse of software.

Find more details about this and the SHARE IT Act in the CMS Open Source Strategy section.

Provide updates to the CMS official JSON enterprise code inventory, the OSS landing page and, the CMS enterprise code repository at Code.gov to encourage discovery and dissemination of the software.

Recommended Practices

RP-OSS-1: Provide Ample Documentation with CMS-Released OSS Code

For a CMS-managed repository, the project team should include sample documentation with the software code for increased adoption and modification by the community. The documentation should provide the information on the project’s mission, philosophy, goal, design, decision-making process, and product roadmap. It should also provide instructions on how to submit issues, feature requests and how to contribute towards fixes or enhancements.

RP-OSS-2: Implement the Tools to Support the Community Around a CMS-Released OSS Project

Implement some or all of the following tools to support the community around an open source project:

  • Mailing lists
  • Message forums
  • Version control
  • Wiki
  • Tracking mechanisms, such as Kanban boards

Development teams should create and maintain a development roadmap for each project to publicly display development work and plans. Various tools are available to facilitate this which can be integrated into the development workflow. Some options include GitHub Kanban, which is included with GitHub, or tools such as Trello or Jira.

RP-OSS-3: Use codejson-generator and automated-codejson-generator to add and update project metadata

Per BR-OSS-15, project metadata can be generated and updated using the following CMS OSPO tools:

Industry Best Practices

  • Follow or adopt a decentralized governance model.
  • Define the team constituents, their decision-making authority and their roles to support the project in the open source community in the COMMUNITY.md file.
  • Define and staff the roles for active user engagement, product roadmap development, and for accepting new contributions via pull-requests requests in the CONTRIBUTING.md file.
  • Identify and promote active contributors to committer status based on the quality and quantity of code contributions and involvement in day-to-day discussions in 

Portal Strategy

Background

Over the past decade, significant increases in demand for information sharing and collaboration have resulted in new federal standards and mandates for more efficient technical infrastructures that deliver higher levels of program performance. Faced with the challenges of an expanding beneficiary population and the task of implementing federal mandates intended to improve U.S. public healthcare, CMS must identify ways to improve the usability, accessibility, and integration of their user services. This will enable more agile and efficient delivery of healthcare services and allow CMS to more rapidly adapt to changes in our national healthcare system.

The essence of the CMS Portal Strategy is this:

  • The user interface presented by a portal is an “Integration Glass,” a single window through which users may see and access information and applications from multiple sources, based on each individual user’s roles and permissions. To ensure consistency, there should be only one source within CMS for each type/collection of information provided through CMS portals. Although multiple CMS portal systems serve specific communities of users, all CMS users will need only one CMS portal to locate and access all CMS content and services relevant to them.
  • A portal combines and displays content and forms from multiple applications and information sources, supports users with navigation and cross-enterprise search tools, supports simplified sign-on, and uses role-based access and personalization to present each user with only relevant content and applications. CMS content and applications from across the enterprise must be designed to fit into this structured environment, and the environment itself must be structured to present users with a common and intuitive look, format, and navigation tools.
  • All CMS stakeholders involved in portals, portal-based applications, or providing content for portals must work toward a shared enterprise vision of the CMS “Integration Glass” for presentation to users.

 

Portals are one of several components for which CMS is evolving enterprise strategies to meet these rising challenges. The strategic importance of portals to the CMS enterprise is apparent because:

  • As user interfaces and end consumers of Shared Services, portals complement the CMS enterprise Shared Services Strategy and unify geographically dispersed services and users through a common access point.
  • An enterprise approach leverages common CMS user identities for all portals, simplifying access to information, services, applications, communication and collaboration tools, social media, and other emerging technologies. Portal benefits include enhanced productivity, efficiency, workflows, communication, and exchange of ideas among CMS user communities.
  • Portals help shape users’ impression of the agency and its programs. Portals provide a “face” of CMS to users, whether they be the public, beneficiaries, providers, other business partners, other agencies, or even CMS’s own staff and contractors.
  • As critical points in the fabric of CMS’s applications and services, portals provide opportunity and infrastructure for implementing and enforcing certain standards of security, accessibility, metadata, technologies, and procedures.

The CMS Office of Information Technology, and in particular, the Chief Information Officer, the Chief Technology Officer, and the Chief Enterprise Architect (CEA), sponsor the task to create a viable and effective portal program.

Purpose

This chapter describes major facets of portal technologies and the business applications of portals. It provides strategic guidance in the form of high-level goals and guiding principles for quality, performance, and relationships among CMS portals, and describes some options to consider in every facet of tactical planning.

This chapter does not provide specific tactical guidance but leaves tactical planning to the stakeholders responsible for each CMS portal and for CMS portal content. Any tactical effort involving a CMS portal must align to this strategy.

The CMS Portal Strategy generates awareness, discussion, and support throughout CMS for implementing an enterprise portal strategy and vision across all lines of Agency business. This strategy is a first step toward introducing a high-level framework that encourages development of business and technology portals to promote operational efficiencies, meet increasing demands for healthcare information, and align more closely with CMS business processes.

Scope

The CMS Portal Strategy is an enterprise-wide initiative that encompasses both internal and external users of CMS data and services. This provides a framework for CMS and contractors by defining terminology, concepts, and guidance as the initiative moves forward. The strategy does not provide a detailed road map for CMS portals, which will be developed in a separate document.

Introduction to Portals

This topic provides background information on fundamental portal concepts and definitions, portal frameworks and features, and portal models relevant to this Portal Strategy.

Fundamental Concepts

This topic explains fundamental concepts and definitions.

Portal Definition and Characteristics

Definition: A portal is a gateway that provides a given user community with access to an organization’s information, services, and data through a consolidated web-based user interface.

Through an organization’s portal, a user gains access to content and services provided by multiple sub-organizations and applications. The user perceives the multiple applications as one system since the portal implements common mechanisms for navigation, authentication, and layout for all content and services.

Portals share several other important characteristics:

  • Typically, a portal appears to users as a website accessed using a web browser. Portal gateways can interact with other client devices such as smartphone applications, telephones using voice and touch-tones, or client applications supporting cellular text messages, e-mail, File Transfer Protocol, facsimile, or other digital data transmissions.
  • Client devices can communicate with a CMS portal through one or more CMS networks designated for that portal. Thus, a portal can be on an external network (the Internet), an internal network, or both. Each portal has a unique address on each network.
  • A portal system can host multiple portals (each with its own address) to serve different user communities or different types of client devices. The community of users served by a given portal may include multiple types of users with different roles and interests.
  • Portals are used by the public, business partners, and employees. A portal may be available to the public, or it may be limited to specific user communities such as enrolled physicians or CMS staff.
  • Portals are customized for the purpose they serve, such as providing information and data collection for physicians or integrating information and tools from multiple data systems to create an efficient digital workspace for a CMS knowledge worker.

Enterprises with Multiple Portals

Enterprises that operate multiple portals on the Internet, or on any network, must make business-driven trade-off decisions. They can choose to consolidate content into one portal, integrate separate portals, create distinct portals for each line of business or user community, or some combination thereof. Several factors should be considered in these decisions:

  • Multiple unrelated portals may confuse users with multiple userIDs, conflicting duplicated information, or inconsistent branding. Multiple portal systems may also duplicate resources and lead to higher implementation and maintenance costs.
  • Significant differences exist between portal platforms in cost, content management, and compatibility with business applications. Some business domains may see significantly lower operations and implementation costs if they use a portal platform that exactly matches their needs instead of a standard product selected by the enterprise. Security concerns, rapid development requirements, or support for specific web technologies can also drive an enterprise toward multiple portals.
  • Portals dedicated to a specific line of business or to a specific community of users may allow significant efficiencies in marketing and user administration.

 

Portals on the same network (such as the Internet) can be designed to allow users to seamlessly navigate between them. This allows the possibility of several independently operating CMS portals on the Internet appearing to users as a single CMS portal environment with one primary entry point address. This also allows any portals on the same network to serve as entry points to this unified CMS portal environment, if they are so designed.

The CMS enterprise must determine how best to execute optimal branding and marketing of CMS portals and content.

Definition: The scope of a given portal is defined by the community of users that portal serves and the client devices it supports.

An enterprise approach to planning multiple portals requires a clear description of each portal’s scope. For example, a top-down approach to enterprise portal planning begins with creating a three-dimensional map that includes:

  • The user communities to be served by the portal(s)
  • The content and services each user community requires
  • The types of client devices that will provide content and services to each user community

This mapping process will identify which portals are needed and provide a basis for each portal’s charter.

Portal Systems and Portal Frameworks

This chapter makes distinctions between three related terms:

  1. Portal. The interactive digital environment presented to users.
  2. Portal System. The software, hardware, and network infrastructure used to create and manage one or more portals.
  3. Portal Framework. The portal system plus other software, metadata, standards, process workflows, and technologies used to create, manage, and render content for one or more portals. Some portal systems provide a complete portal framework, while others rely on separate components for web content management functions.

The most effective portal framework design for a given portal depends on the purpose and content required by that portal. Existing CMS portals currently use multiple portal systems and frameworks. CMS is developing a portal system reference architecture that will provide guidance for multiple portal frameworks. Over time, CMS will migrate existing web applications to the portal system reference architecture.

Web Applications, Web Services, and Portlets

A “web application” is a business application that uses web technology. The term is most often used to describe applications that use a web browser for the user interface, but it also includes any use of web technologies within an application.

Web applications are often designed as standalone portals that present users with only one application-specific user interface. Web applications may also be designed as either “web services” or “portlets” to be integrated into enterprise portals.

Web services are interfaces of a web application that support transactions using Extensible Markup Language and related cross-platform web service standards. Portal-side scripts present an input form to the user, pass the user’s input to the web service, and render the transaction results for display to the user. Web services require development of portal-side scripts that often must be uniquely customized to each portal platform.

Portlets are pluggable, user-interface software components of a web application that are managed and displayed in a portal, often as a window, frame panel, or image within the portal’s user interface. Portlets may be standalone applications (for example, a calculator); a generic interface for a standards-based application service (for example, a newsfeed reader that displays news headlines); or an application-specific, client user interface such as a form or a dashboard graph. Portlets often use web services as an interface to web applications. Portlets do not require customization of scripts or code for each portal platform, and their installation on a portal is a relatively simple process.

By incorporating web services or portlets, an enterprise portal can present the user with a single user interface to multiple web applications. For both web services and portlets, adherence to industry standards is essential for compatibility and interoperability.

Web Content Management

Web content management (WCM) refers to four functions: creating, managing, storing, and deploying web content. A web content management system (WCMS) is a collection of products, preferably integrated, that supports these functions by providing some or all of the following content services:

  • Authoring
  • Publishing
  • Scheduling
  • Storage
  • Versioning
  • Workflow
  • Archiving and Disposal

 

From an enterprise perspective, all WCM implementations must address existing federal and CMS policies for protection of sensitive information, CMS disclosure and public release requirements, and data retention requirements. WCM implementations are unique to each portal framework and sometimes to each portal. CMS uses multiple WCMS products to support portals across the Agency, subject to the following goals:

  • To ensure consistency for all portals that content appears on, each collection of content should have only one original source within CMS.
  • To avoid unnecessary complexity and inefficiency in web content management, content should flow through a minimum number of WCMS workflow paths to reach portals. Parallel WCMS workflow paths should be consolidated or connected whenever possible.
  • To leverage common skills and IT investments, CMS will standardize on one WCMS product for all portals and all content other than web applications. Other WCMS tools may be used when appropriate or advantageous to CMS.

Three Portal Models

From an enterprise strategy perspective, the biggest difference between portal models is not in their technology, but in their purpose, user interface design goals, owner responsibilities, and content life-cycle workflow. Table - Three Different Portal Models compares three different portal model examples to illustrate these differences: Applications-Focused Portal, Information-Focused Portal, and Team Collaboration Portal. Operations, governance, and portal frameworks will be different for each of these (or any other) portal models.

Table - Three Different Portal Models
ComponentApplications-Focused PortalInformation-Focused PortalTeam Collaboration Portal
PurposeProvide a unified user workspace where all information, resources, and tools are readily available to users to support a given user role, job function, or activity.Inform and educate users, providing access to services and assistance.Provide tools and workspaces that teams may use to share, communicate, and collaborate on documents, information, calendars, and tasks.
Design & Personalization GoalsProvide a highly task-optimized user interface that aids users in their daily roles.Provide an intuitive and supportive user interface presenting content in convenient formats that guide users’ attention to relevant information.Empower a team to customize their workspace while ensuring simplicity and security.
Content Characteristics
  • Content represents a user’s actual work and resources for a given activity or role.
  • Some content may be prioritized for presentation based on timeliness, importance, or relevance.
  • Content other than that generated or retrieved by applications rarely changes.
  • Content has no life-cycle workflow except as implemented by the applications.
  • Content is provided by one or more CMS organizations or services.
  • Content must be intuitively organized, easily navigated, and searchable.
  • Some content may be prioritized for presentation based on timeliness, importance, or relevance.
  • Content is frequently added or changed.
  • Content life-cycle workflow is defined and automated.
  • Content is provided by users.
  • Content must be searchable.
  • Intuitive organization and navigation may be a team responsibility.
  • Content is continuously added or changed.
  • Content life-cycle workflow is ad hoc and manual.
  • Users are allowed some control over access to content and its organization.

Vision and Strategy

Current State of CMS Portals

CMS has many gateways of information that are not unified or interrelated with one another. Many of these independent gateways do not fully leverage portal technologies to support integration, collaboration, or advanced user interface features.

These information gateways have been separately developed as independent systems addressing specific needs of a given program or user community. Because CMS lacked enterprise-wide portal standards for branding, style, and architecture, this collection of CMS information gateways has the following drawbacks:

  • No single entry point exists for accessing the CMS resources and environment.
  • Multiple log-ins, using multiple userIDs, are required for those who need access to different areas of the CMS system.
  • Multiple systems exist for user self-service account management.
  • No aggregation or correlation of information and services exists between multiple CMS sources at the user interface level.
  • No ability exists for users to navigate or search across all enterprise content.
  • Users experience inconsistent CMS interface design and branding.

The CMS Portal Vision

The CMS vision is to develop a one-portal strategy creating consistency in user identity, CMS branding, and user access. Users would have one userID and one entry point to all CMS content and gateways of information they are permitted to access.

The phrase “one-portal strategy” does not mean that CMS would have only one portal, but that CMS users would need only one CMS portal to locate and access all CMS content and services relevant to them. Portal technology makes this possible by rendering the user interface as an “Integration Glass” for user access to CMS information and services. As shown in The Portal as an “Integration Glass” , an “Integration Glass” renders and combines content for display from multiple applications and data sources, supports users with navigation and personalization, and implements simplified sign-on, permissions, and some security functions.

The Portal as an “Integration Glass” (page 35)

Key elements of this Portal Vision include:

  • One enterprise-wide portal strategy
  • A shared vision of the portal—the “Integration Glass” through which users access CMS information and services
  • Links directly to CMS’s mission and goals
  • Leveraging of existing governance bodies
  • Alignment with the CMS Business Reference Model (BRM) in classifying audiences and Communities of Interest (COI)
  • Alignment with and leveraging of the CMS Shared Services Strategy for leveraging applications and services

CMS Portal Strategy Business Drivers

The desire for a common portal strategy across the CMS enterprise is driven by:

  • Rapidly expanding interaction between CMS, CMS business partners, providers, beneficiaries, and the public as well as CMS’s increased reliance on portals to support those business interactions
  • Increased demand for information sharing and collaboration across CMS and the federal government that supports cross-domain / cross-agency trust, improves access to data, and provides semantic interoperability
  • The need to contribute to a more efficient government by finding ways to mutually leverage public and private sector investment to drive business and mission improvements while, at the same time, responding to fiscal constraints through effective use of the Agency’s IT budget (i.e., fostering reuse of existing capabilities and assets)
  • The effort to increase security, transparency, and resilience by ensuring that:
    • Information and data exchanges are secure
    • IT and business assets are available, understood, and governed
    • Continuity of operations is preserved

CMS Portal Strategy Goals

The CMS Enterprise Portal Strategy has the following goals:

  1. Simplify access to government services while supporting Agency integration and growth. Reduce complexity of service-providing processes. Enable intuitive, continuous customer interaction directly through one high-impact channel, promoted and identified by a single agency brand.
  2. Provide a well-defined portfolio of business services. Align all CMS portals to provide users with a full suite of services (e.g., eligibility, enrollment, customer service, and identity management) across business domains.
  3. Enhance productivity for all users. Support role-based contextual services and information filtering, collaboration within and outside of the Agency, and self-service for customers, employees, and partners. Enable users to find relevant information more quickly and, through collaboration, to leverage collective experience and reduce cycle times.
  4. Reduce help desk costs by providing self-service account administration for all CMS portal users. Self-service account administration is defined as an “Online user interface by which a user may reset forgotten passwords, request access to application resources, update some identity information such as location and contact information, and view corporate and organizational identity information.”
  5. Improve internal business agility to more rapidly respond to changing demands.
  6. Deliver new critical mission capabilities more quickly and with less cost by minimizing IT rework, alleviating duplicate efforts, and sharing a single infrastructure.
  7. Provide secure, shared, and governed services by continuously ensuring that content and services accessed through portals and managed across government organizational boundaries are protected and available.

CMS Portal Strategy Objectives

Stakeholders responsible for each CMS portal and for CMS portal content should align their tactical planning and designs with the following objectives:

  • Charter and define all CMS production portals serving the public, beneficiaries, or healthcare providers according to purpose, content scope, and user community.
  • Standardize on one web content management system product for all CMS portals and all content other than web applications. Other WCMS tools may be used when appropriate or advantageous to CMS.
  • Allow only one original source within CMS for each collection of content to ensure consistency of that content on all portals.
  • Provide all end users of CMS portals with:
    • A single point of access: one Universal Resource Locator
    • Self-service identity management
    • Role-based access with content and service filtering
    • A common navigational model, look and feel, and branding

Tactical Challenges

CMS must address four tactical challenges to implement the CMS portal vision:

  • Obtaining the buy-in and sponsorship of senior CMS leadership
  • Developing the requisite governance and standards
  • Continuing to expand and join CMS user identity and authentication capabilities to support CMS user portals across the enterprise
  • Developing migration plans for legacy CMS portals as opportunities to migrate become available

Tactical Recommendations

Stakeholders responsible for CMS portals and for CMS portal content should consider the following recommendations in their tactical planning and designs:

  • Leverage and extend existing CMS portal technologies and standards as well as the CMS Shared Services Strategy, the CMS TRA, and the CMS Business Reference Model.
  • When appropriate, use an incremental approach to extend and integrate with existing CMS portals and websites.
  • Leverage CMS enterprise services for identity management and access control.
  • Leverage industry best practices and standards, enabling efficient and flexible architecture and supporting enterprise-wide collaboration and reuse.
  • Ensure that CMS portal-related standards, governance, and business performance measures differentiate between different portals models and are appropriate within the context of the models for which they are intended.

Benefits of the Portal Strategy

With the adoption and implementation of one enterprise-wide Portal Strategy, CMS lines of business would realize benefits from:

  • Enhanced productivity for all portal end users
  • Increased adoption rate of new services
  • Cheaper and faster deployment of new service channels
  • Reduced number of costly channels to maintain
  • Reduced disruptions from having to patch or upgrade multiple portal implementations
  • Saved money on recouped hardware and administration

Guiding Principles

To align with the CMS Enterprise Portal Strategy, stakeholders responsible for each CMS portal and for CMS portal content must adhere to the following Guiding Principles in their tactical planning and designs or provide adequate justification for exceptions.

Principle 1

All CMS portals on a given network—public internet, CMS Core Network, CMS Private Network (CMSNet), CMS extranet, public telephone system, etc.—must be accessible to users via a single point of access on that network.

Rationale:

One URL, phone number, or other network address will be sufficient for a user on a given network to access all CMS production portals available on that network. An exception may be for portals dedicated to highly sensitive content and services.

Principle 2

Each user will have a single identity for all CMS production portals.

Rationale:

Supporting a single user identity for access to all CMS production portals generates important benefits for security, user administration, user productivity, and user empowerment through communication and collaboration tools, social media, and other emerging portal-based technologies.

Principle 3

Public content on a CMS portal should be available to anonymous portal users.

Rationale:

Users only seeking public information should not have to identify themselves. Public information should be freely accessible. Likewise, users should be able to ask general questions or use public services without having to authenticate their identity when that is not necessary. Without authenticating, users may even provide a name, telephone number, email address, or postal address to receive responses to general questions.

Users will not be required to authenticate themselves (i.e., log in) to gain access to a CMS portal’s public content and services. A CMS portal will require authentication only when users request to retrieve content or access services requiring their personal identity or role information. This includes disclosure or modification of personal or sensitive information as well as identity and contact information associated with a userID.

Principle 4

All CMS Production Portals serving the public, beneficiaries, or healthcare providers will be chartered and defined according to purpose, content scope, and user community. These definitions will be updated annually and reviewed by CMS division representatives and all CMS organizations responsible for operating such portals.

Rationale:

CMS Production Portals are the most visible face of CMS. They should deliver consistent, authoritative content and services that avoid duplication. These same portals need to be recognized as Agency resources with great potential for delivering new content and services across CMS business domains. The annual review described in this principle provides a governance process for working toward these objectives as part of an enterprise-wide Portal Strategy.

Principle 5

All CMS portals and WCM implementations must address existing federal and CMS policies for protection of sensitive information, CMS disclosure and public release requirements, and data retention requirements.

Rationale:

Each business owner is responsible for managing the content that they publish.

Principle 6

For any given collection of content, all CMS portals will use the same original CMS source to ensure consistency of that content on all portals where it appears.

Principle 7

All CMS public portals and content must adhere to CMS branding standards.

Principle 8

CMS portals, web services, and portlets must adhere to CMS and industry standards.

Guidelines for Adopting and Supporting the CMS Portal Strategy

The following guidelines provide the basis for adopting and supporting the CMS Portal Strategy. Stakeholders responsible for each CMS portal and for CMS portal content should align their tactical planning and designs with these guidelines.

Address All Four Perspectives of the Portal Strategy

To be comprehensive, tactical planning for a portal must address four distinct perspectives, as explained in the tableTable - Four Portal Strategy Perspectives.

Table - Four Portal Strategy Perspectives
PerspectiveRelevant Goals and Objectives

User / Consumer:

Concerned with providing portal users with a
high-quality user experience. Scope includes:

  • User awareness of the portal
  • CMS branding and style
  • Navigation, enterprise search, and collaboration features
  • User support and self-service capabilities
  • Accessibility and client device support
  • Provide intuitive, continuous customer interaction directly through one high-impact channel, promoted and identified by a single agency brand
  • Support role-based contextual services and information filtering, collaboration within and outside the agency, and self-service for customers, employees, and partners
  • Enable users to find relevant information more quickly and, through collaboration, to leverage collective experience

Information:

Concerned with ensuring that each portal delivers information and services that are understandable, locatable, complete, and relevant to the portal’s community of users. Scope includes:

  • Management of user communities
  • Alignment of content and services to user communities
  • Manner in which content is described and presented to users
  • Management of web content, including its creation, management, storage, and deployment

Links directly to CMS’s mission and goals:

  • Align with the CMS Business Reference Model to classify audiences and Communities of Interest
  • Reduce the complexity of service-providing processes

Data:

Concerned with using CMS’s authoritative data sources to deliver data-driven content and services that are current, consistent, understood, and high-quality. Scope includes:

  • Standardizing access to CMS’s authoritative data sources
  • Supporting harmonization with other data sources across healthcare and government domains
  • Align with the CMS Shared Services Strategy for leveraging applications and services
  • Continuously ensure that content and services accessed through portals and managed across government organizational boundaries are protected and available

Application:

Concerned with integrating CMS applications with the portal environment. Scope includes:

  • Reference architectures for portal systems and frameworks
  • Standards for portlets, web services, and other portal-related application interfaces
  • Utility functions that portal systems provide for applications (i.e., user authentication, session management, content rendering, navigation, layout, and logging)
  • Improve internal business agility to more rapidly respond to changing demands
  • Deliver new critical mission capabilities more quickly and with less cost by minimizing IT rework, alleviating duplicate efforts, and sharing a single infrastructure
  • Continuously ensure that content and services accessed through portals and managed across government organizational boundaries are protected and available

 

Address the Processes Associated with Portals

Effective tactical planning for a portal must address two broad categories of portal-related processes: Management Processes and Operational Processes. The following describes these at a high level. While some processes will be managed centrally, others may be completely decentralized to maximize efficiency.

Management Processes

Management processes define decision-making practices, establish goal and performance metrics, oversee portal strategy and prioritization, report on portal success, and define and establish policies and standards. Management processes consist of:

  • Strategy Setting and Prioritization – Establishes portal goals and objectives, investigates portal development opportunities and areas of focus, and prioritizes initiatives.
  • Performance Management – Defines measurement objectives, develops a measurement approach, defines measurement categories, identifies indicators and metrics, and collects and reports data.
  • Policy and Standard Setting – Identifies and revises policies and standards around portal development, content creation, and presentation.

Operational Processes

Operational processes involve the daily activities and operations supporting the portal as follows:

  • Web Content Management – Includes content generation, review, translation, testing, posting, quality control, archiving, and appropriate updating.
  • Operations – Includes determining portal production requirements, identifying system monitoring requirements, developing security management approaches, defining service level agreements, conducting system quality control, and managing change requests.
  • Feature Request Management – Captures and assesses requests for offering new services or tools through the portal. Centralizing this process allows for proper prioritization, coordination, and budgeting. Implementing a request management process also avoids unnecessary duplication and enables economies of scale.
  • Continuous Improvement – Includes collecting stakeholder and user feedback, conducting portal reviews, identifying improvement areas, assessing new technologies, and recommending areas for site development.
  • Quality Control – Establishes a validation process for content accuracy and relevance before publishing and a site monitoring process for flagging content for revision.

Address the Roles of Stakeholders

Tactical planning for CMS portals should address the roles and needs of the following major stakeholder roles:

  • User Communities. The user communities served by a given portal are the ultimate judges of how well the portal’s design, content, and accessibility meet their needs. Effective tactical planning for a portal also includes user awareness, training, assessment, and feedback mechanisms.
  • Portal Business Owner. The overall business lead for a given portal is responsible for coordinating that portal’s design and operations, authorizing the portal’s users, and ensuring that the portal’s content and services are relevant and friendly to the user community. Portal owners require guidance, governance, and assessment tools necessary to fulfill their responsibilities and make their portal successful.
  • Web Content Management Teams. Skilled staff or contractors are responsible for technical administration, design, content maintenance, and content workflow of a given WCMS as well as for providing technical design and implementation support to all portals using that WCMS. This team may also be responsible for providing user training.
  • Web Content and Service Suppliers. Individuals and business units supplying content to portals, or operating services and web applications offered to users through portals, are responsible for their data and content quality, user authorization, and implementation of role-based access control (RBAC) within a portal service.
  • Portal System Owner and Maintainer. This is the business owner responsible for a given portal system that hosts one or more portals.

Employ a System of Governance

The CMS Portal program will apply the CMS Integrated IT Investment and System Life-Cycle Framework and other Agency governance processes to portals, portal systems, and portal frameworks. Portals, portal systems, and WCMSs will all follow the CMS ILC as separate projects. Portlets and web services are components of business applications and therefore will follow the CMS ILC and Business Service Management life cycles as part of their associated business application project.

Portlet Services

Introduction

Purpose

This chapter provide the necessary information and guidance for integrating an application with CMS’s Enterprise Portal. This chapter describes the technical specifications to allow application integration teams to successfully integrate individual applications with the CMS Enterprise. Issues addressed in each topic of this chapter will help application integrators produce a specification that can be used to plan, implement, and test the application integration with CMS’s Enterprise Portal.

Scope

This chapter represents the specifications that should be used by CMS and CMS / Contractor partners for the CMS Processing Environments. It addresses specific architectural requirements of interest to IT architects desiring an overview of standards and policies for building J2EE Portlets and Remote Portlets in compliance with the Web Services for Remote Portlets (WSRP) standards.

Enterprise Context

This topic defines fundamental enterprise concepts and terminology used throughout this chapter and describes its alignment within the CMS TRA.

Introduction to Portals and Portlets

A portal provides functionality and content from a diverse and potentially distributed set of applications (called portlets) in a common, unified manner. From a user perspective, a portal aggregates multiple applications (represented as portlets) into a common look and feel while allowing users to customize and personalize both content and layout. A portal also makes available a user interface via a Web server. From an application perspective, a portal provides a common run-time environment (i.e., a portlet container) for portlet instantiation, usage, communication, and destruction. This includes a single framework for user management, authentication and authorization, and inter-portlet communication. Portlet Wireframe and Nomenclature illustrates key components of a portal architecture.

Portal Architecture Overview (page 28)

A portlet generates portlet content as markup fragments. The portlet renders these portlet fragments in a consistent manner. For example, a portal may add a title, specific window controls, and decorations and renders this within a portlet window. Multiple portlet windows aggregated into a complete document represent the portal page. This is illustrated in Portlet Wireframe and Nomenclature , adapted from Java Specification Request (JSR) 168: Portal Specification.

Portlet Wireframe and Nomenclature (page 29)

The Java Portlet Specification (JPS) defines a consistent programming model for developers to use in developing portlets that reside within a portlet container. The JPS was developed as a (JSR, which dictates that any Java-provided content comes by way of a portlet. Version 1.0 of the JPS (JSR 168) specifies portlet modes, window states, data models, packaging formats, and methods of integrating applications with portlets. JPS Version 2.0 (JSR 286) extends JSR 168 to define inter-portlet communication, serving dynamically generated resources, Extensible Markup Language or JavaScript Object Notation data, and portlet filters and listeners. The WSRP specification complements JSR 286 by defining a standard for communicating with remote portlets. The JSR 286 and WSRP standards should be used and are referenced throughout this chapter.

CMS Enterprise Portal Environment

The CMS Enterprise Portal is intended to provide a shared service for exposing CMS applications and content to intranet, Internet, and extranet users consistent with the CMS TRA Multi-Zone Architecture. The outermost zone is the Presentation Zone that supports Web servers and user interfaces. The middle zone is the Application Zone that controls business logic, business rules, and messaging. The third zone is the Data Zone that supports the database servers. The CMS Enterprise Portal spans the Presentation and Application Zones. In the Presentation Zone, the portal serves as a user interface, presenting different CMS applications in a consistent manner. The portal container resides in the Application Zone, responsible for integrating and providing a common program interface to disparate CMS applications made available as portlets. The business logic and data access contained within applications resides in the Application and Data Zones, respectively.

The Enterprise Portal will enable users to interact with CMS through a single website, allowing for personalization as well as integrated security. Users will be able to alter, within specified limits, the experience they enjoy with the portal by selecting from approved portlets and configuring the position and size of portlets to meet their needs for a productive and comfortable experience.

The CMS Enterprise Portal has SSO and a Security Programming Interface (SPI) for managing authentication and authorization. The SPI isolates most details of complying with security checks from the application integrator. For example, the CMS Enterprise Portal will shield portlet applications from the details of SSO with the existing CMS applications. There is no need for specialized programming or use of third-party software on the part of the application integrator; the SPI provides services based on the URL mapping of applications.

Portlet Business Rules

Table - Summary Portlet Rule Statements outlines the Portlet Business Rules, which are explained further below.

Table - Summary Portlet Rule Statements
Rule StatementsLocation
Portlet Life-cycle: Each portlet must include a deployment descriptor as specified in JSR 286.BR-P-1
Portlet Life-cycle: Each portlet must implement specific portlet life-cycle management functions as specified in JSR 286.BR-P-2
Portlet Logging: Each portlet must log events via portlet container logging functions.BR-P-3
Portlet User Interface: If functionality to customize / personalize the portlet user interface is provided, it must be consistent with JSR 286.BR-P-4
Portlet Security: Each portlet must control access to content / functionality based on user and role information via the portal.BR-P-5
Portlet Security: Each portlet must use portal services to authenticate users.BR-P-6
Portlet Security: Each portlet must securely transport sensitive content.BR-P-7
Inter-Portlet Communication: Inter-portlet communication must be performed only by either public render parameters or events as specified in JSR 286.BR-P-8
Remote Portlet Security: Remote portlets must follow Web Service standards defined in this Web Services chapter to secure Simple Object Access Protocol messages, verify the sender’s identity, and exchange authentication and authorization data.BR-P-9

Portal Integration Options

There are two basic approaches to integrating an application with CMS’s Enterprise Portal: URL pass-through and native portlets. This topic presents a discussion of the pros and cons of each approach.

URL Pass-Through

URL pass-through is the practice of using an HTML object to simply proxy some external content. The HTML object is used to place the content, referenced by a URL, on an external server, thereby rendering the content within the portal context. This means that the content will be accessed within the portal’s security and the user interface framework, but all functionality can be provided from an external application. Once the portal opens the HTML object between the user and the external application, the user’s browser will directly interact with the application. Although requests for the application may go through the same Web server that services portal requests, portal server components will not intercept the application requests. This tends to work well with third-party products, legacy applications as well as applications interested in maintaining their own UI. This integration pattern is responsive and allows applications to have their own screen size for their content. Additionally, the CMS Enterprise Portal takes care of the authentication and applications can rely on headers being passed by Enterprise Portal for authorizations.

Native Portlet

Native portlets are those with a main application that is entirely contained within the portal (within a portlet). This normally includes accessing CMS services available to the portal in the Application Zone. In contrast to the URL pass-through approach, the native portlet approach ensures that:

  • The application is customized for a portal environment.
  • Content is displayed optimally for portlet screen size.
  • Inter-portlet communication is possible.
  • The portlet can use built-in portal services such as authentication for SSO.

The remainder of this chapter focuses on guidance for developing a native portlet for use within the CMS Enterprise Portal because, in the case of URL pass-through, the application is managed outside the portal environment and does not require specific life-cycle management within the portal.

Note: This integration pattern is recommended only for those applications that have an existing UI already built using the portlet technology and are unable to move to non-portal technologies. For other applications, the use of URL pass-through integration pattern is recommended.

Native Portlet Integration

A portal provides a common point of access for several applications. It aggregates content across multiple applications and allows end-users to configure their view, including both the portal’s look and feel as well as the content / functionality made available by the portlets. Through heterogeneous application implementations, a portal needs to introduce a degree of uniformity to provide a consistent user experience. This leads to performance, scalability, reliability, security, and presentation considerations.

This topic focuses on addressing each of these areas of integration, particularly in portlet
life-cycle management, personalization, security, inter-portlet communication, and remote portlets. The guidance in this topic is based on JSR286 and WSRP standards. Although actual portal products may include vendor-specific extensions, they should be avoided to prevent vendor lock-in.

Portlet Life Cycle

Portlets are managed within a portlet container. Associated with each portlet is a deployment descriptor, which provides the portlet container with the information necessary to run the portlet. The portlet container instantiates a single object to serve requests targeted to the portlet for each portlet definition in the deployment descriptor file (portlet.xml, XML schema available at http://download.oracle.com/otndocs/jcp/portlet-2.0-fr-oth-JSpec/).

As described in the JSR 286 specification, each portlet deployment descriptor contains the portlet name (portlet-name), portlet class (portlet-class), and initial parameters (init-param) for the portlet. Also contained within the deployment descriptor is the length of time the portlet container should cache the contents created by the portlet (expiration-cache). The content types supported by the portlet and modes supported for that content type are defined by the supports element (supports). The portlet information element (portlet-info) defines information like the portlet’s default title. Finally, default portlet preferences are defined by the portlet preferences element (portlet-preferences). The deployment descriptor file may also include vendor-specific extensions. Sample Portlet Descriptor shows an example of a portlet deployment descriptor (from the Sun Microsystems white paper “Introduction to JSR168”).

Sample Portlet Descriptor (page 32)

Four basic functions serve as a “contract” between the portlet and the portlet container and are invoked through different phases of the portlet life-cycle:

  1. Init(). The init function is called when the portlet is instantiated by the portlet container. This function basically prepares the portlet to serve requests.
  2. ProcessAction(). The process action function is called when the user makes changes to the portlet, such as via the doView() or doEdit() method. This function processes user changes.
  3. Render(). The render function is called when the portlet is drawn by the desktop.
  4. Destroy(). The destroy() function is called when the container destroys the portlet, when the portlet is no longer needed.

As part of the “contract,” each portlet implements these functions used by the portlet container to manage the portlet. A portlet may also publish information to log files through the portlet container—for example, for debugging purposes or error reporting. The portlet container makes available a PortalContext interface through which portlets can access logging functions.

Portlet Personalization

A consistent interface across portlets requires that each portlet implement particular functions to present content to users and for users to customize / personalize the portlet. A GenericPortlet class is provided which implements the render() function and delegates the call to more specific functions, depending on the portlet mode (view, edit, or help). The GenericPortlet class can be extended to implement as many of the mode-specific render functions as necessary. These functions, as described in JSR 286, include:

  1. doView(). The doView function renders the portlet contents when in view mode.
  2. doEdit(). The doEdit function allows users to edit the portlet.
  3. doHelp(). The doHelp function makes portlet help available for users.

Portlets should implement these functions to give users the ability to customize / personalize the portlet.

Portlet Security

Given the potential diversity of both portlets and users accessing the portal, it is important to manage access controls and exchange content securely while supporting a unified user experience. Security is the responsibility of the portlet, although the portal container makes functions available to support secure access to portals.

The portlet container provides information about the user accessing the portal, including the user’s role. The portlet container is not, however, responsible for authentication. In JSR 286, three basic functions support user authentication:

  1. getRemoteUser(). The getRemoteUser function returns the username used to authenticate into the portal.
  2. isUserInRole(). The isUserInRole function determines if a user is in a specified security role.
  3. getUserPrincipal(). The getUserPrincipal function determines the principal name of the current user.

Portlets should use these functions to authenticate and authorize user access to portlet content and functionality.

A second aspect of securing the portlet is protecting the content during transit to ensure content integrity and/or confidentiality is preserved. This is especially important if there may be sensitive data exchanged between the portlet and users. Security constraints are specified based on a portlet collection and user data constraints. The portlet collection describes the portlets to be protected. The user data constraints describe the transport layer security required by the portlet collection. For example, exchange of confidential messages must occur over Transport Layer Security protocol, typically as part of an HTTP/S communication. As part of the portlet deployment descriptor, developers should set a flag to indicate whether communication with the portlet should occur over secure communication.

Each project / application that wishes to integrate with the portal is responsible for ensuring adherence to the CMS ARS and the appropriate security artifacts are completed. Additionally, each portlet may have its own security assurance level as stipulated by the CMS ARS .

Inter-Portlet Communication

Portlets that are developed independently and packaged as separate portlets may coordinate with each other at run-time. For example, there may be a single portlet that allows users to specify their location and, based on user input, other portlets may display various information about that location, such as weather forecast, houses for sale, or traffic conditions. The JSR 286 specification defines two ways in which inter-portlet communication may occur:

  • Public render parameters. This allows for a portlet to set parameters that can be read by other portlets.
  • Events. This allows portlets to issue events / notifications that other portlets may or may not react to.

In the case of public render parameters, the parameters passed between portlets are built into the URL, and the user’s browser causes the event to be triggered. This has security implications; for example, the parameters are visible in clear text. Alternatively, events provide a more secure means for interacting with other portlets. Events go to all portlets, though they are contained within the CMS environment.

At development time, portlets need to define what type of data they understand as part of their deployment descriptor. At run-time, either shared parameters or eventing enable portlets to be combined in a loosely coupled manner, share data, and potentially derive new functionality based on the composite portlet.

Remote Portlets

Portlets often run locally within the same computing environment (i.e., sharing resources with the portal) as the rendering portal, but this is not always the case. Sometimes the portlet runs in a remote portlet container with respect to the portal. A remote portlet works best when there are multiple application owners and/or applications that are not subject to centralized deployment constraints. In this case, the WSRP specification defines a Web Service interface for accessing and interacting with portlets that reside in a remote portlet container. WSRP allows portals to display portlets that are running remotely within their portal, without any additional portal development. The advantage of remote portlets is that they do not compete for portal resources and provide a more scalable implementation.

The WSRP specification does not dictate specific security requirements, but portlet developers should follow Web Services Security (WS-Security) and Security Assertions Markup Language (SAML) when implementing WSRP-compliant portlets, in accordance with the Web Services chapter. WS-Security defines mechanisms for signing SOAP messages (for integrity), encrypting SOAP messages (for confidentiality), and attaching a security token to verify the sender’s identity. SAML provides a standard for exchanging authentication and authorization data between the remote portlet and the portlet container. TLS should also be used for transport level security.

The WSRP specification also describes session management. Session state is stored on the portlet producer side (i.e., the Web service that offers one or more portlets). When a session is created, the producer returns a reference to the session (i.e., session ID) to the consumer that provides a user interface to the portlet. For each subsequent invocation by the consumer, the consumer must include the session ID in the request to the producer.

Portlet Deployment

There are two basic portlet deployment options: local or remote. As described above in Remote Portlets, a local portlet runs within the same computing environment as the portal. A remote portlet runs in a remote computing environment and communicates with the portal via a Web service, based on the WSRP standard. Furthermore, a local portlet runs as a Java application within a portlet container (like a servlet running within a servlet container). A remote portlet based on WSRP, however, exposes a consistent Web service interface so that it can be accessed and viewed by a portal, but the actual implementation may vary.

Trade-offs should be considered when determining whether a local or remote portlet is more suitable. A local portlet should be used, for example, to optimize message exchange performance between the portlet and portal. A remote portlet is more suitable when the portlets are implemented in a programming language other than Java; local JSR286 portlet implementation is limited to Java, while remote WSRP portlets are agnostic to implementation technology. Furthermore, remote portlets may simplify security assessments, because a remote portlet may only require a single security assessment for multiple deployments, while a local portlet may need assessment for each deployment along with the portal.

Portlet Business Rules

CMS has established the following common Portlet (P) business rules for governing the Agency’s portlet environments. These rules apply to CMS.gov and all CMS sub-sites.

Portlet Life Cycle

BR-P-1: Each Portlet Must Include a Deployment Descriptor as Specified in JSR 286

BR-P-2: Each Portlet Must Implement Specific Portlet Life-Cycle Management Functions as Specified in JSR 286

BR-P-3: Each Portlet Must Log Events via Portlet Container Logging Functions

Portlet Personalization

BR-P-4: If Functionality to Customize / Personalize the Portlet User Interface Is Provided, It Must Be Consistent with JSR 286

Portlet Security

BR-P-5: Each Portlet Must Control Access to Content / Functionality Based on User and Role Information via the Portal

BR-P-6: Each Portlet Must Use Portal Services to Authenticate Users

BR-P-7: Each Portlet Must Securely Transport Sensitive Content

Inter-portlet Communication

BR-P-8: Inter-Portlet Communication Must Be Performed Only by Either Public Render Parameters or Events as Specified in JSR 286

Remote Portlets

BR-P-9: Remote Portlets Must Follow Web Service Standards

Remote portlets must follow Web Service standards defined in the Web Services chapter to secure SOAP messages, verify the sender’s identity, and exchange authentication and authorization data.

Future Considerations

This chapter describes advanced techniques such as inter-portlet communication. Additional areas will merit future consideration as this chapter evolves to meet business needs, including the following topics:

  • Integration of Asynchronous JavaScript and XML (AJAX) technologies into the portlet ecosystem
  • Further treatment of inter-portlet communication for security purposes
  • Portlet skins and themes not yet standardized by Java Community Process (JCP)
  • JavaServer Faces (JSF) to portlet-bridging technologies

Although technologies under future consideration (such as AJAX and portlet skins and themes) can be used, this chapter does not cover these topics. Guidance on these topics may change in future versions.

Containers and Microservices

Containers and Microservices Introduction

This topic is based on Research Spotlight Technology Review – The Importance of Input Validation, CMS TRB Technical Topics, 12 October 2017.

Container technology is a method of virtualization and an alternative method to virtual machines for implementing virtualization. Containers provide operating system-level virtualization, which runs on a single control host and accesses a single kernel. Containers use the host OS features to create virtual instances that share the same OS kernel while a hypervisor creates VMs, each of which can run an OS and none-of-which share an OS. This can provide efficiencies over VMs because a kernel does not load for each user session and uses less memory and CPUs than VMs running the same workload.

Container Architectures

There are many options when it comes to deploying containers. Container products are shipped with multiple features and have associated tools that control the scaling and deployment / placement of containers (orchestration) within the infrastructure. (The interested reader is encouraged to read the TRB Research Spotlight paper on the Orchestration of Containers).

Development teams have implemented containers using various levels of sophistication. Some groups have deployed containers without the use of these orchestration features, opting to control deployment and scaling in a more manual fashion. Other development groups have opted for full incorporation of the tools into their DevOps and production environments.

Implementations using full container orchestration platforms typically have separate instances of these tools in each environment (development, test, production); and may also include an additional separate instance to support containerized tools to support DevOps and deployment processes. Although many of these container orchestration platforms are capable of multi-tenancy deployments, this separation helps to enforce access and deployment controls across environments. Sample Implementation of Container Management / Orchestration Components depicts a sample implementation for the production environment; other architectures are possible depending upon a project’s requirements.

Generally, the container orchestration tools handle automated provisioning, deployment, and scaling of containers across the hosts. These tools typically leverage an overlay network that is on top of CMS’s traditional zone-based network. Each of these should be configured to enforce the appropriate routing and traffic flow policies according to CMS policies and guidance. The implementation of containers and overlay networks at CMS are still required to meet CMS TRA and CMS ARS standards to ensure the required level of protection to mitigate many of new exposures introduced in the environment.

Sample Implementation of Container Management / Orchestration Components (page 30)

Security

Virtualization is an abstraction layer that separates that environment from the underlying physical hardware. VMs utilize individual kernels, which limits the attack surface to the hypervisor. In theory, this means vulnerabilities in a particular OS cannot be used to compromise other VMs running on the same physical host. Containers share the same kernel; therefore, extra care must be taken to avoid security issues from adjacent containers.

Business Rules and Recommended Practices

Business Rules

BR-CA-1: The CMS Zonal Architecture Must Be Preserved

The CMS Multi-Zone Architecture (CMS Multi-Zone Architecture) calls for the implementing a zonal architecture with at least three security challenges prior to accessing CMS data. The zones (e.g., Presentation, Application, and Data) must be on separate virtual machines (VMs) and subnets. A VM / physical host may only support one zone. An individual container is not allowed to encompass a complete application (multiple zones). The Presentation, Application, and Data Zone components must be split into separate containers and deployed onto the corresponding VM/physical host for that zone as shown in the Sample Container Implementation diagram below.

Sample Container Implementation (page 31)

BR-CA-2: Lower Environments Must Be Separated from Production

Development, test, and implementation may not be implemented on a VM / physical host serving production.

BR-CA-3: The CMS TRA Zonal Hierarchy Will Be Enforced

The CMS TRA Zonal hierarchy will be enforced. Communications to the data zone requires that the other implemented zones be traversed, not allowing the Data Zone to be accessed directly. This requires Presentation-to-Application and Application-to-Data Zone communication; communication from Presentation Zone directly to Data Zone is not allowed. The traffic must be challenged in each zone and not be a pass-through component such as a load balancer.

BR-CA-4: Interfaces between the Containers Implemented in Each Zone Must Be Locked Down (Source / Destination)

Locking down the interfaces and placement of containers in the CMS TRA Zonal architecture can be accomplished using a combination of:

  • AWS VPC rules and security groups
  • Subnets (each zone is usually sub-netted)
  • Network Access Control Lists – NACL
  • Container Interface and placement rules within the orchestration software. This would include any namespace rules. With network namespaces, each collection of containers (sometimes referred to as a “pod”) gets its own IP and port range, thereby isolating pod networks from each other on the node.

BR-CA-5: Container Traffic to Non-Container or External Destinations Must Follow Existing TRA Rules

Outbound traffic must flow out through the Presentation Zone and internal communication must adhere to the CMS TRA Zonal architecture.

Recommended Practices

RP-CA-1: Force Containers to Write to Container-Specific File Systems

RP-CA-2: Implement Read-Only File Systems Whenever Possible

Docker containers are read-only. Other container systems may not be.

RP-CA-3: Run Your Containers as Non-Root Whenever Possible

Treat root within a container as if it is root outside of the container.

RP-CA-4: Create Containers with the Least Privilege Possible

Dropping privileges is important and still the best practice. Even better is to create containers with the least privilege possible. Remove all capabilities except for what is explicitly required by the container. Be aware of default privileges that are set on container creation. Containers should run as user, not root.

RP-CA-6: Use a Security Mechanism for Mandatory Access Controls

Security mechanisms such as SELinux or AppArmor enforce mandatory access controls (MAC) for every user, application, process, and file.

Related CMS ARS Security Controls include: CM-3 - Configuration Change Control.

Rationale:

Mandatory access controls provide an additional layer of security to keep containers isolated from each other and from the host. Note that the STIG for RedHat Linux requires SELinux.

RP-CA-7: Use Cgroups

Cgroups (control groups) limit, account for, and isolate the resource usage (e.g., CPU, memory, disk I/O, network) of a collection of processes. Use Cgroups to ensure your container will not be stomped on by another container on the same host. Cgroups can also be used to control pseudodevices, which are a popular attack vector.

Rationale:

RP-CA-8: Use a Secure Computing Mode Profile

A secure computing mode (seccomp) profile can be associated with a container to restrict available system calls.

RP-CA-9: Harden All Containers / Components

The containers / components should be hardened to the appropriate guidelines (CIS, DISA, etc.)

Rationale:

RP-CA-10: Validate All Third-Party Containerized Applications before Implementation

Many people / organizations are releasing containerized applications. Before implementing a third-party containerized application, it must be validated to ensure no security risks exist. Use image-scanning tools to detect known vulnerabilities and validate security controls. Use runtime container tools to scan and monitor processes for security risks.

Rationale:

RP-CA-11: Maintain the Immutability of Your Containers

Do not patch running containers; rebuild and redeploy them instead.

Rationale:

To be supplied.

Container Best Practices

The following are best practices for implementing containers:

  • Containers enable a microservices approach to application design, where complex applications are split into discrete units. This reduces the complexity of managing and updating the application because a problem or change related to one part does not require an overhaul of the entire application. Segregating containers by host, by function, or even by data classification or type when possible helps to secure your infrastructure by applying specific and predictable access control rules. This approach also helps detect abnormal activities much faster, enhancing security stance against zero-day attacks.
  • Test, test, test. The only way to learn if containers will work is to build test applications. That means replicating real-world use, including containerization of real workloads.
  • Consider the architecture. Using containers in the cloud means more architecture and planning. The ability to limit dependencies, scale efficiently, and replicate soundly will pay dividends as the application is deployed. This includes the management of resources used, application performance, reliability, and portability.
  • Container images should be tagged appropriately, perhaps including the destination environment name, within a secure image repository. This can help to ensure that images only get deployed to the appropriate environments.
  • Business units at CMS should plan and develop architectures for the possible sharing of container images across projects.

Microservices

This topic is based on Research Spotlight Technology Review – Microservices, CMS TRB Technical Topics, 10 February 2017.

Microservices (MS) are a modern interpretation of SOAs for building distributed software systems. Each component / module of an application is developed and deployed separately. This contrasts with a traditional, “monolithic” application in which all components are developed and deployed as one piece. Microservices are well suited to DevOps methods and tools.

Monolithic versus Microservices Architecture

Monolithic Architecture

A monolithic application is packaged and deployed as a monolith. The actual format depends on the application’s language and framework. For example, many Java applications are packaged as WAR files and deployed on application servers. As an application grows to become a large, complex monolith, maintaining it may become a struggle for the development organization, and agile development and delivery may become unsustainable.

A monolithic architecture has the following disadvantages:

  • Overwhelmingly complex. Fixing bugs and implementing new features becomes difficult and time consuming. Continuous deployment (e.g., DevOps and Agile) is difficult.
  • The larger the application, the longer the startup time.
  • Difficult to scale when different modules have conflicting resource requirements.
  • May be less reliable. Because all modules are running within the same process, a defect in any module, such as a memory leak, can potentially bring down the entire process.
  • Difficult to adopt new frameworks and technologies.

Microservices Architecture

To avoid the disadvantages of monolithic architecture, many organizations, such as Amazon, eBay, and Netflix, have successfully adopted a microservices architecture pattern where applications are split into sets of smaller, interconnected services.

Business Rules and Recommended Practices

Business Rules

None.

Recommended Practices

None.

Microservices Best Practices

Netflix is one company that has successfully implemented microservices architecture on a large scale. The Netflix development team established several best practices for designing and implementing a microservices architecture This topic summarizes some of those practices (and paraphrased from Adopting Microservices at Netflix: Lessons for Architectural Design).

Keep Code at a Similar Level of Maturity and Stability

To add or rewrite some of the code in a deployed microservice that is working well, the best approach is usually to create a new microservice for the new or changed code, leaving the existing microservice in place. This is sometimes referred to as the immutable infrastructure principle, a way one can iteratively deploy and test the new code until it is bug free and maximally efficient, without risking failure or performance degradation in the existing microservice.

Do a Separate Build for Each Microservice

Do a separate build for each microservice by pulling in component files from the repository at the revision levels appropriate to it. This sometimes leads to the situation where various microservices pull in a similar set of files, but at different revision levels. This is a trade-off that can make it more difficult to clean up the codebase by decommissioning old file versions and verifying that a given revision is no longer in use.

Deploy in Containers

Deploying microservices in containers requires just one tool to deploy everything. So long as the microservice is in a container, the tool can deploy it.

Prefer Stateless Services for Improved Resilience and Scalability

Treat servers, particularly those that run customer-facing code, as interchangeable members of a group. The only concern is that there are enough servers to produce the needed work (auto-scaling can adjust the numbers up and down). If one server stops working, it is automatically replaced by another. Avoid “fragile” systems that depend on individual servers to perform specialized functions.

Create a Separate Data Store for Each Microservice

When multiple microservices share the same database, and therefore, if one microservice requires a change to the database structure, the other microservices must be updated as well. This can be avoided by giving each microservice its own data store, thereby reducing the need to update more than one microservice for a simple change in a data model.

The advantages of separate data stores for each microservice are a trade-off. Using different data structures limits reusing code between microservices. Breaking apart the data can make data management more complicated because the separate storage systems can easily get out sync or become inconsistent, and foreign keys can change, making joins between tables invalid.

If separate data stores are used, tools that perform master data management can help to find and fix inconsistencies. For example, an MDM tool might examine every database that stores subscriber IDs to verify that the same IDs exist in all.

Orchestration of Containers

This topic is based on Research Spotlight Technology Review – Orchestration of Containers (Microservices), CMS TRB Technical Topics, 23 January 2018.

Containers and Microservices are sometimes used interchangeably, but they are two different concepts. Microservices is an architectural style where the application is implemented as a suite of small services, each running in its own process. These processes communicate using lightweight mechanisms (usually HTTP). The goal is to make these lightweight services independently deployable in a fully automated fashion so they may scale as needed. Containers are the vehicle for implementing these microservices and provide a means that enables development groups to more rapidly create iterations to address business problems by piecing together a set of smaller services. Containers are lighter weight compared to VMs. Containers make more efficient use of the underlying architecture and can scale applications to meet fluctuating demand. Containers also facilitate the ability to move applications between different environments or even clouds.

There are two options for implementing containers. One, the bare metal implementation, is when the containers run directly on the host. The other implementation, and probably the most likely for CMS, will be when the containers are implemented within a VM. This is because CMS is making use of more cloud-based infrastructure, where bare metal options are not available. The implemented containers must adhere to CMS ARS security and CMS TRA requirements.

Orchestration

Individual containers can be crafted and managed by a runtime APIs. These runtime APIs may meet the needs for managing one container on one server, but the strength of the container and microservice architecture is the ability to automatically scale the systems. Container orchestration (OR) tools are best suited for managing multiple containers on multiple hosts. These tools can treat an entire cluster of machines as a single entity for deployment and management. They can automate and manage the initial placement, the scheduling and deployment of updates, and perform health monitoring functions while supporting scaling and failover.

Business Rules and Recommended Practices

Business Rules

BR-OR-1: Container Images Must Be Hardened

The container images must be hardened according to security specifications, both at rest and during operation. Ensure conventional security tools can peer into running containers to assess hardening and patching compliance.

BR-OR-2: Use Security Monitoring on Containers

The containers must be subjected to security monitoring, including the review of logs. The logs must be given to CMS security or security personnel must be granted access to review the logs.

BR-OR-3: The Deployment Infrastructure for Containers Must Be Hardened and Monitored

The deployment infrastructure and containerization hosts for containers must be security hardened and monitored.

BR-OR-4: Containers Use Must Respect the Multi-Zone Architecture

Container use must maintain the multi-zone architecture—separation of presentation, application, and data services must be maintained (i.e., vertical separation). For example, a single container within the Presentation or Application Zone cannot be used to contain all elements of an application stack (presentation, application, and data storage services).

BR-OR-5: Libraries of Containers Must Be Maintained in CMS-Only Stores

Libraries of containers, as with any software, must be maintained in CMS-only stores, not shared in Internet libraries (CMS private repository, e.g., private Docker Hub).

BR-OR-6: Required Orchestration Capabilities

Container Orchestration tools must provide the following capabilities to support development and operational (DevOps) needs:

  • Configuration. Container configurations must be described in text files.
  • Service Discovery. OR tools must provide a mechanism to enable the runtime discovery of containers or microservices.
  • Policy-based Placement. To support security, performance, and high-availability requirements, container OR tools must provide a mechanism for defining policies concerning the placement and scaling of containers in the environment, including restrictions such as non-colocation. The policies must also ensure that the container placement and security rules meet CMS ARS security and CMS TRA guidelines.
  • Monitoring. OR tools must be able to track and monitor the health of the system’s containers and hosts. OR tools must run specified health checks at the appropriate frequency, update the list of available nodes, and ensure that the current state of the cluster matches the configuration specified. OR tools must provide or support real-time monitoring to help find performance, compliance, and vulnerabilities issues.
  • Resilience. If a container crashes, the tools must be able to launch its replacement. If a host fails, the OR tools must be able to restart the container(s) to a new host.

Rationale:

These required capabilities establish the minimum baseline for application container orchestration. The configuration capability enables DevOps teams to easily manage configurations via configuration management and version control tools.

The service discovery capability allows containers and microservices to discover one another within the orchestrated environment. This also enables runtime reconfiguration and provides for better resiliency and adaptability.

The policy-based placement capability enables the OR tools to enforce the defense-in-depth strategy of the CMS TRA. A system implemented using containers must still adhere to the security and architecture measures that protect CMS assets from attacks. Even though a service is not implemented within a container, it may not be moved to the edge of the network. CMS expects that all the required defense-in-depth mechanisms still front access to Data Zone services. For example, non-colocation is used to prevent all containers for a service to be placed on the same VM or physical host. Policy-based placement also enables adaptable deployments across development, testing, and production environments.

The monitoring capability acknowledges that containers must be monitored to ensure optimum performance. The resilience capability introduces automation to allow for automatic recovery in case of container or host failure.

Recommended Practices

RP-OR-7: Prefer Container Orchestration Tools That Allow for Container Motion

Container OR tools should be able to move containers across environments and clouds. The placement of the containers must still follow CMS ARS and CMS TRA guidelines. Container OR tools that can operate and move containers across multiple infrastructure providers are preferred.

RP-OR-8: Rolling Upgrades

OR tools should allow for the rolling upgrade of containers across the cluster. The tools will ensure that traffic is routed appropriately, ensuring that a minimal number of containers are available as the other containers are progressively replaced, following the immutable infrastructure model of RP-CA-11.

Serverless Architecture (AWS Lambda)

This topic is based on Research Spotlight Technology Review – Serverless Architecture (AWS Lambda), CMS TRB Technical Topics, 19 December 2017.

To use or to obtain more information on leveraging the AWS Lambda Serverless compute service, contact the IUSG’s Cloud Operations Team at cmscloudoperations@cms.hhs.gov.

Introduction to Serverless Architecture

Although it is called “Serverless Architecture,” it is not really serverless. This architecture always requires a physical or virtual server built up with an OS, CPU, and memory to run software. A Serverless Architecture eliminates the b burden on the application developers to own and manage the infrastructure, allowing them to focus on implementing the code itself.

Stateless Computing Architecture Using “Function as a Service” (FaaS)

One approach to Stateless Computing Architecture is using FaaS to implement event-driven mini functions. FaaS is stateless. Examples of FaaS functions are “moving a file from one S3 bucket to another S3 bucket”, “send an email”, “send alerts/notification”, “perform a mathematical operation “, and “perform virus check”. In short, the FaaS life cycle is Load, Execute, Unload.

FaaS is a next level of microservices. FaaS helps break down a microservice further into independent functions and events (pieces of code). For example:

  1. A File Upload Service can be decomposed into multiple, stateless mini functions and events such as file processing, file move, and virus checks.
  2. A Measures Engine Service (used for analyzing clinical data against measures of healthcare performance) can be decomposed into such mini functions such as Measure Eligibility Check and Performance Criteria Evaluation.

Microservices and FaaS can co-exist and work well together. An application can be well-developed using the combination of microservices (for example, features where the state needs to be maintained) and FaaS. Microservices can also be developed as a logical group of isolated small functions.

AWS Lambda

AWS Lambda (LD) is an example of a Stateless Computing Architecture platform. AWS Lambda is a compute service that helps to “execute the code” without the need for the application developers to provision or manage the servers or any other computing resources. The application developers are only responsible for writing the code as functions. Each function will be configured as a “Lambda Function”. When a function is invoked, the AWS Lambda creates a container with all necessary configurations (balanced memory, CPU, network and other resources) to execute the code. AWS Lambda is responsible for creating and deleting these containers. In addition, AWS Lambda will be responsible for the entire operational and administrative activities such as provisioning capacity (scaling), monitoring code and infrastructure health, logging, applying security patches, and loading the code.

Once the code has been configured as “Lambda Functions” within AWS Lambda, the functions can be triggered by the AWS Services’ published events such as DELETE event in a S3 bucket, exposed as HTTP endpoints from an API Gateway, or invoked by using the AWS SDK.

The advantages of AWS Lambda include:

  • Simplified operational and infrastructure management. A serverless platform such as AWS Lambda is responsible for auto scaling and infrastructure management. It is designed to provide high availability for FaaS functions.
  • Reduced operational costs and administrative activities. Automatic operational management such as auto scaling helps to reduce the operational costs.
  • Clear separation of duties. The serverless architecture helps to isolate duties for system and functional developers.
  • Faster Innovation. The FaaS allows developers to update their technologies frequently to experiment with innovative ideas.

The disadvantages of AWS Lambda include:

  • Multi-tenancy and vendor lock-in. The serverless architectural model does not allow a dedicated server or virtual server (or any other compute resources) as the platform. AWS Lambda is responsible for creating or decommissioning the FaaS containers. The design and code must be compatible to support the vendor platform. In addition, AWS offers simplified Lambda integrations (for example, using event source mapping, SDK) with other AWS services such as the S3 buckets, API Gateway, SNS, SQS, and features like Step Functions to orchestrate custom components like microservices. Using these services and features within the Lambda functions may greatly confine you to the AWS platform.
  • Complexity. The distributed computing architecture of AWS Lambda or other serverless computing platforms can create architectural complexity.
  • Limitations. Serverless computing platforms may have limitations in programing languages support, memory allocation, number of threads, or payload size in functions.
  • Potential Latency. The platform is responsible for creating and bootstrapping the lambda containers that may involve potential latency.

CMS Considerations and Recommendations for Using AWS Lambda

The Serverless Architecture / AWS Lambda technologies may offer various technical and business advantages to the CMS enterprise. This architecture also poses some constraints such as multi-tenancy and vendor lock-in.

CMS recognizes the following recommendations for using AWS Lambda for CMS applications:

  • Non-Cloud Based Solution. If cloud hosting is not an option for your application, using the AWS Lambda or any other Serverless platform is not a choice. At present, the Serverless architecture is only supported within the AWS environment.
  • Unpredictable Load. AWS Lambda is best suited for tasks that require instantaneous scaling to support unpredictable load. AWS Lambda supports instant scaling up to a large number of parallel processes. For example, the AWS Lambda is best suited for applications when you are expecting a massive hit on your application for a very short period and need containers to scale up and down based on the traffic.
  • Long-running Processes. FaaS / AWS Lambda functions are not good fit for long-running processes.
  • Quantity. It is preferable to not implement significant / larger portions of an application using AWS Lambda or other serverless technologies because this may lead to vendor lock-in with that specific vendor.

CMS Measures Engine Example

In the referenced file upload measures engine example, the AWS Lambda service will be responsible for the entire operational management of the following stateless functions: File processing, File move, Measure eligibility check, and Performance evaluation. The uploaded code should not have any affinity to the underlying infrastructure (i.e., stateless). If a Lambda Function is required to maintain state, it should use one of the supported persistent stores such as S3, Dynamo DB, and NFS.

Managing Lambda Functions

The Life Cycle of a Lambda function is shown in Life Cycle of a Lambda Function. An AWS account with identities is required to set up and manage the AWS Lambda service. The identities are AWS account root user, IAM user, and IAM roles. The chosen identity must have valid permissions to work with the resources in a service.

In AWS Lambda, the primary resources are Lambda Functions and Event Source Mapping. Permission for these resources are managed by the permission policies. The permission policies can be attached to the IAM identities (users, groups, and roles) and to resources of AWS services like AWS Lambda. Each AWS Lambda function is controlled by the IAM execution role and this execution role must include the required permission policies to access other resources such as a S3 bucket. For example:

AWSLambdaBasicExecutionRole and AWSLambdaVPCAccessExecutionRole

AWS Lambda runs the Lambda functions in a VPC by default and supports connecting to resources inside a private VPC by providing the additional configuration settings (VPC Subnet and Security Groups). To support Lambda functions auto scaling behavior, the VPC resources must be configured with sufficient capacity.

Important Note: AWS Lambda does not support connecting to resources within “Dedicated Tenancy” VPCs where all EC2 instances / hosts that are launched in a VPC run on hardware that is dedicated to a single customer. Also, AWS Lambda does not support functions accessing multiple VPCs (in this case VPC peering is required).

Life Cycle of a Lambda Function (page 33)

Invoking Lambda Functions

Invoking a Lambda Function illustrates how Lambda functions may be invoked. AWS Lambda supports both synchronous and asynchronous invocation of a Lambda function. A Lambda function can be invoked / triggered from AWS Services or from HTTP Endpoints (using the AWS API Gateway or AWS SDK). The AWS provides the “invoke” API with the option for application developers to choose synchronous or asynchronous invocation types. When a lambda function must be invoked / triggered from AWS Services’ event sources (published events by the respective AWS Services), the invocation methods are predetermined by the respective AWS Services.

Invoking a Lambda Function (page 34)

Monitoring Lambda Functions

The AWS Lambda uses its Amazon CloudWatch to monitor the AWS Lambda infrastructure and Lambda functions and reports the real-time metrics. The Lambda “functions” can also be monitored using other enterprise logging tools. Some logging libraries are integrated with AWS Lambda to allow lambda functions to send logging events to them through HTTP event collector features. The application developers need to write lambda functions to perform this task.

Business Rules and Recommended Practices

Business Rules

BR-LD-1: Avoid Using Lambda for CMS Sensitive Data

At present, the AWS Lambda does not support dedicated hosts or instances or even connecting to resources within the dedicated tenant VPC. The AWS Lambda platform is shared and a multi-tenant environment. Therefore, any handling / processing of CMS sensitive data such as PII or PHI within the serverless platform must be avoided. If any sensitive data handling / processing is unavoidable within AWS Lambda or any other serverless platform due to other factors, the project ISSO must work with the ISPG to implement the required security controls to protect CMS sensitive data.

BR-LD-2: Integrate with CMS Enterprise Security

CMS uses an enterprise logging and event platform for application, infrastructure and security logging, and for other services such as threat analysis. The AWS Lambda supports various enterprise logging platforms for application logging (not for infrastructure). Consider writing a common logging function and use it with other functions for logging purpose. Also, integrate your application and infrastructure log files with the CMS CCIC platform for further analysis. Please contact ISPG for more information on CCIC.

BR-LD-3: Configure Lambda to Control the Cost / Budget

When using AWS Lambda, carefully evaluate your business needs, costs, and operational requirements, and configure conditions such as the number of parallel executions, timeouts, and memory allocations to control the cost / budget.

BR-LD-4: Do Not Use the Local Volatile Storage of Lambda Containers for Persistent Data

AWS Lambda platform creates containers to execute the functions. The containers include resources such as the local volatile storage. It is always a best practice to use the container resources efficiently. For example, consider using the local volatile storage spaces temporarily for any potential performance increment opportunities, but definitely not for permanent data.

BR-LD-5: Use Lambda Only for Stateless Transactions

Maintaining state within the Lambda functions requires using other services such as databases, S3 buckets, and/or vendor specific features such as Step Functions. These other services will add complexity to the architecture and operational management and may lock you into AWS.

Recommended Practices

RP-LD-1: Avoid Complex Application Programming or Workflows in Lambda

When using FaaS functions, try to avoid writing complex application programming or workflows using vendor specific services that could lock in with a specific vendor. FaaS functions are best fit for parallel execution of mini / small tasks. They are also best suited for executing event-driven tasks in an application, such as batch processing, file transfer, virus check, email delivery, scheduled backup services, and monitoring services.

Serverless Architecture and Lambda Best Practices

The following are best practices for using Serverless Architecture and AWS Lambda:

  • Orchestration / Sequencing. AWS Lambda / FaaS functions are not a good fit for orchestrating/sequencing sub-tasks. Orchestration / sequencing tasks will add complexity to the architecture and operational management.
  • Hybrid Approach. Microservices and FaaS functions may be used in a hybrid approach. Microservices can be implemented using a group of FaaS functions. Microservices helps to decouple functionality by domains, and the independent/isolated sub-tasks within a microservice are eligible candidates for developing as FaaS functions.

Input Validation

Input Validation Introduction

This topic is based on Research Spotlight Technology Review – The Importance of Input Validation, CMS TRB Technical Topics, 12 October 2017.

The Equifax data breach is old news, but its lessons raise basic questions on effective IT processes for monitoring common vulnerabilities and the timely patching of the servers. The root of the problem Equifax faced lies in the vulnerability of the Open Source Apache Struts framework that allows an attacker to execute random code inserted into Content-Type of the HTTP header. It is possible to perform a remote code execution attack with a malicious Content-Type value. A workaround to this problem is to implement a servlet filter to validate Content-Type and throw away requests with suspicious values not matching multipart / form-data.

Some commercial vulnerability monitoring services have observed a new Apache vulnerability that is being actively exploited. The vulnerability (CVE-2017-5638) is a remote code execution bug that affects the Jakarta Multipart parser in Apache Struts. They have observed simple commands (e.g., whoami) as well as more sophisticated commands, including malicious code execution. The adverse effects of this inserted code include stopping the Linux firewall and subsequently downloading and executing a malicious payload from a web server simply by injecting this code into the HTTP header. All input must be validated to assure that malicious code will not be executed.

CMS TRA and CMS ARS Requirements

The CMS Services Framework (CMS Services Framework) and the TRA Multi-Zone Architecture (CMS Multi-Zone Architecture) require challenge requests for validation, authorization, and malicious content. This generally occurs within the Presentation Zone, where business applications first receive data from external sources. Specifically, the Application Development, Web-based UI Services business rule for Input Validation (BR-UI-16) suggests that the CMS websites must validate all user input at a minimum on the server side, although validating on both the client and server side is recommended.

According to the Common Weakness Enumeration®:

Use an allowlist of acceptable inputs that strictly conform to specifications. Reject any input that does not strictly conform to specifications or transform it into something that does. Do not rely exclusively on looking for malicious or malformed inputs (i.e., do not rely on a denylist). However, denylists can be useful for detecting potential attacks or determining which inputs are so malformed that they should be rejected outright. When performing input validation, consider all potentially relevant properties, including length, type of input, the full range of acceptable values, missing or extra inputs, syntax, and consistency across related fields, and conformance to business rules.

The CMS ARS Security Control, SI-10, Information Input Validation, provides supplemental guidance as follows:

Structured messages can contain raw or unstructured data interspersed with metadata or control information. If software applications use attacker-supplied inputs to construct structured messages without properly encoding such messages, then the attacker could insert malicious commands or special characters that can cause the data to be interpreted as control information or metadata. Consequently, the module or component that receives the tainted output will perform the wrong operations or otherwise interpret the data incorrectly. Prescreening inputs prior to passing to interpreters prevents the content from being unintentionally interpreted as commands. Input validation helps to ensure accurate and correct inputs and prevent attacks such as cross-site scripting and a variety of injection attacks.

Implications for CMS Development

Many CMS applications have used the Apache Struts framework extensively, and a few applications have used functionality specific to the file upload / download capability of this framework. The CMS data centers provision boundary appliances enable the header inspection and validation before the payload is accepted for further processing in the deeper tiers of the application.

The CMS application development community must always adhere to the CMS TRA and CMS ARS security controls and validate all requests (and response messages) as well as the results of any message transformations or alterations. There are many viable options and threat protection-related settings available for use in the CMS Processing Environments. The application development team should include the review and tuning of these options and settings as a part of the application acceptance process. An application that does not use a dedicated appliance can still validate the Content-Type as a Servlet filter in the Application Zone and discard the requests that do not pass validation.

Configuration Management

Introduction to Configuration Management

Background

The CMS Policy for Configuration Management, CMS-CIO-POL-MGT01-01, April 2012 (hereafter simply “CMS CM Policy”) formalizes the management and control of CMS’s inventory of IT assets in a disciplined manner to ensure their integrity and availability to support the CMS mission.

Toward that end, CMS has tailored or customized the ISO / IEC 12207 standard, Software Life Cycle Processes, to meet CMS’s specific needs, and documented these processes in the CMS Target Life Cycle framework.

As CMS implements components of this IT Modernization, it needs policies, standards, charters, processes, and procedures for planning, managing, governing, executing, and controlling the enterprise CM process.

Purpose

This Configuration Management (CM) chapter establishes business rules and recommended practices for IT configuration management of the CMS enterprise. CM processes are both managerial and technical activities. The contents of this chapter apply to both disciplines as appropriate. As noted in the CMS CM Policy, these business rules cover the automated systems, software applications and products, supporting hardware and software infrastructure, as well as contractor deliverables, associated documentation, and services that are used behalf of CMS. The business rules expressed in this chapter support the Agency’s CM framework and provide guidance to the CMS enterprise for applying and developing CM plans, processes, and procedures that describe the effective management and implementation of CM practices within CMS.

This chapter integrates modern configuration management practices from the DevOps movement with traditional practices defined in the IT Infrastructure Library (ITIL). This allows CM projects to begin adoption from familiar concepts and terminology.

Scope

This chapter represents the required standards for conducting CM activities performed by CMS and CMS contractor partners across all CMS Processing Environments. The guidance stated in this chapter reflects the CMS agreed-upon industry and government best practices to support the most viable approach for CMS that meets legislatively mandated security and privacy requirements as well as current technical standards and specifications.

In concert with the CMS CM Policy, the concepts, strategies, business rules, and recommended practices in this chapter promote the alignment of IT activities and IT assets owned or controlled by CMS, including those of CMS’s agents, contractors, or other business partners when acquired or supported by CMS funding. Accordingly, this chapter applies to all hardware, software, supporting infrastructure, services, interfaces, data, and associated documentation—regardless of origin, nature, or location (e.g., contractor, in-house, development, operations, internal and external systems, and all hosting data centers)—unless otherwise specified.

Compliance with Existing Federal, HHS, and CMS Policies

The CM provisions of the following documents supersede this chapter. CMS projects need to comply with these CM provisions as appropriate, especially those associated with system-level infrastructure CM.

The following configuration standards are mandatory for devices and systems:

Note: The foregoing three citations form standard security baselines against which compliance is tracked and monitored. These are systems engineering responsibilities. This is different from configuration management at a project level where the goal is controlled evolution driven by business goals.

Business Rules for Configuration Management

CMS has established the following business rules for governing CM activities within the CMS environments.

BR-CM-1: Projects Must Produce a Configuration Management Plan

In compliance with the CMS CM Policy, CMS requires the development, tailoring, and use of configuration management plans that describe the CM roles, responsibilities, and activities to meet the needs of specific CMS projects.

Related CMS ARS Security Controls include: CM-09 - Configuration Management Plan.

Rationale:

A CM plan documents the roles, responsibilities, and activities that the project will conduct in order to ensure that all configuration items are baselined and managed accordingly. Having a plan ensures that all participants understand their roles in the process.

The CM plan can also provide guidance for how to conduct activities. For example, the CM plan could explain how software CI’s would be stored in the CMS GitHub, documents are stored in Confluence, and hardware assets are to be managed in Chef. In this instance the CMDB is collectively GIT, a Wiki, and a Chef repository. Each CI should be managed in an appropriate configuration management tool, collectively known as a Configuration Management Database (CMDB).

The CM Plan can be maintained in a standalone document, a set of wiki pages, or document management system. It should be easily accessible by team members to increase communication and promote transparency.

BR-CM-2: Projects Must Identify Items to Be Placed under Configuration Control

Projects must identify and baseline all items placed under CM control. CM items considered for baseline control must be important project artifacts that are subject to change during system development and operation, and include:

  • Charters and policies
  • Project schedules and budgets
  • Business and technical requirements
  • System, subsystem, and product interfaces
  • Architectural, hardware and software designs
  • Developed code
  • Hardware infrastructure documentation
  • Test artifacts
  • Critical security items, including security settings
  • Media assets (e.g., video, audio, and images)
  • Software configuration settings
  • Service dependencies
  • Trouble Tickets or Issues

 

Related CMS ARS Security Controls include: CM-03 - Configuration Change Control.

Rationale:

Identifying CIs is the process of identifying important, durable assets needed for the lifecycle of a project and system. Temporary assets would not typically be identified as configuration items. For example, most logs are not configuration items. The contents of a database are not a configuration item. The database engine configuration, on the other hand, is a CI. A rule of thumb is to consider any item produced by a human as a CI. Accordingly, if it is the output of an automated process, it is generally not a CI. Thus, a source file is a CI, but an object file (i.e., build output) is not.

Once CIs are identified, they are ready to be managed. Management simply means tracking changes to the item and being able to readily know its state.

With the advent of “Infrastructure As Code,” the ability to manage IT infrastructure such as systems and devices has been greatly improved. In most cases, a CI can be represented a set of one or more files. Baselining these CIs becomes a simple application of version control.

BR-CM-3: Significant Changes to Configuration Items of a System or Component Managed by a CCB Requires the Approval of That CCB

Configuration Control Boards establish baselines on those items under CCB control. Once established, CCBs manage and control changes to the controlled items. CMS formally charters its CCBs with specific thresholds for their change approval authority.

BR-CM-4: Projects Must Maintain Accurate and Reliable CI Information

Projects are responsible for tracking the configuration status of the CI’s that were identified and baselined. This includes status of proposed changes and the implementation of approved changes. Projects must use persistent repositories, known collectively as a configuration management database (CMDB), for storing up-to-date and accurate configuration information. The repositories must include audit trails because projects must be prepared to provide status on CM activities to project staff, senior program management, CMS Security, and CMS auditors.

Related CMS ARS Security Controls include: CM-08- Information System Component Inventory.

Rationale:

Maintaining accurate and reliable information on CIs critical to managing assets. Modern configuration management software can be configured to provide such information on demand. However, it is responsibility of the project team to ensure that CIs and associated metadata (such as AWS or Git tags) are accurate and verifiable.

BR-CM-5: Projects Must Conduct Periodic Audits of CM Activities and Products

Conduct audits on CM activities to verify compliance to both the project-specific configuration management plan and the CMS Risk Management Handbook: Configuration Management (CM), processes, and procedures. Audits must identify discrepancies to project leadership.

Related CMS ARS Security Controls include: CM-05 - Access Restrictions for Change and
CM-06 - Configuration Settings.

Rationale:

Security CM audits must be conducted by authorized security audit personnel in accordance with the policies and procedures established in the CMS Risk Management Handbook. Frequency of audits needs to be defined as part of the Configuration Management Plan. Without audits to verify that the expected configuration matches the actual, it is possible for project assets to be unknowingly misidentified or unidentified. Conducting audits, reviewing, and addressing the results ensures that baseline information is accurate, enabling confidence in both business and technical decisions.

BR-CM-6: Projects Must Maintain Configuration Baselines

Project develops, documents, and maintains under configuration control, a current baseline configuration of the information system. Project must establish baseline configurations for information systems and system components, including communications and connectivity-related aspects of systems.

Related CMS ARS Security Controls include: CM-02 - Baseline Configuration and CM-03 - Configuration Change Control.

Rationale:

Baseline configurations are documented, reviewed, and agreed-upon sets of specifications for information systems or configuration items within those systems. Baseline configurations serve as a basis for future builds, releases, and/or changes to information systems. Baseline configurations include information about information system components. Baselines are typically specific to the configuration management system used to manage configuration items.

For example, in Git, a baseline is a tag or commit id since either of these can uniquely identify a set of software configuration items (i.e. source code, files, etc.)

One measure of successful baselining is the ability to reliably reproduce a system’s configuration, quickly and accurately, using only that information which is baselined. In other word, reproducibility depends only on CIs.

BR-CM-7: Follow the CMS ARS Hierarchy for Security Configuration

The CMS ARS Security Control CM-6 1 (b) provides the authoritative CMS hierarchy for applying security configuration guidelines. Apply this standard across all CMS systems and devices.

Related CMS ARS Security Controls include: CM-6 - Configuration Settings

Rationale.

The CMS ARS Security Control CM-6 1 (b) provides the authoritative CMS hierarchy for applying security configuration guidelines. This ensures that CMS systems and devices are hardened at a minimum according to the best available and applicable standards developed by Federal agencies and security organizations.

Best Practices and Recommendations

Recommendation 1:
Establish Configuration Control Boards to Manage CM-Controlled Items

Configuration Control Boards (CCB) can be established to manage significant changes to CM-controlled items. CCBs must review, approve, disapprove, defer, escalate, or remand change requests (CR) to baselined items.

A CCB must be able to identify the impact of changes across multiple projects and coordinate roll-out. The CCBs conduct impact assessments on a project’s requested changes to determine the following:

  • Cost to implement the change
  • Schedule impacts and time to implement the change
  • Performance characteristics of the change
  • Impact on other items within the system
  • Impact on system interfaces
  • Security impacts

The baselines must be kept current as controlled items are changed.

Recommendation 2:
Projects Should Consider Using CM Automation

The advent of pervasive virtualization and “infrastructure as code” has enabled API-driven automation. Projects should strive to use automation to eliminate inconsistency and variability from processes.

Automation tools such as Chef and Ansible can be used for automating system configuration management activities, while declarative infrastructure automation tools such as AWS CloudFormation can be used to automated platform configuration. Projects are encouraged to use COTS configuration management products rather than developing their own.

When using automation, the configuration files used as input to the tools become the configuration items.

Configuration Management Processes

There are two basic processes in performing configuration management: Configuration Identification and Configuration Control / Change Management.

Configuration Identification

Configuration Identification involves identifying the configuration of items such as hardware, software, and documentation within a system as well as their physical, functional, and performance characteristics. Configuration Identification also involves the identification of items that do not necessarily have physical, functional, and performance characteristics such as project schedules, budgets, and plans. The following are some examples of these configurations:

  • Requirements
  • Architecture, software, and hardware designs
  • System, subsystem, and product interfaces
  • Test plans
  • Test procedures
  • Code
  • Hardware infrastructure
  • Schedules
  • Budgets

Configuration Items

A Configuration Item (CI) is the identified configuration of an item, or a portion of its parts, that is designated for CM and change control. CIs are important program or project items that are subject to change during their life. One would not identify temporary items as configuration items.

Identifying Configuration Items

CI identification involves the analysis of the identified items to determine their importance to the project and its products, and an analysis to determining if the items are subject to change during development and operations.

The identification of CIs includes:

  • Assigning unique identifiers to each CI
  • Technical documentation describing each item’s configuration
  • Establishing naming conventions for CIs
  • Establishing and maintaining associations between CIs and their descriptive information
  • Describing the product structure through the selection of CIs and identification of their internal and external relationships

Configuration Control / Change Management

Configuration control / change management is the systematic evaluation, coordination, approval or disapproval, and implementation of changes to CIs. Change control is the sub-process of making changes in a planned fashion, where the objective is to correct defects, add capability, and more effectively implement new and improved methods and systems on a project or in an enterprise. The CR is the typical means to initiate a change; some changes may originate as a Problem Report (PR) or an Engineering Change Proposal (ECP).

Establish Baselines

A baseline is the approved and fixed (immutable) configuration of a collection of one or more CIs at a specific time in the collection’s life cycle that serves as a reference point for change control. For example, a Git commit can be used as a baseline since it represents an immutable collection of files at a specific point in time. Not every commit is used as a baseline, however, because not every commit is suitable for release.

The baseline is a specification or product that has been formally reviewed and agreed upon, that thereafter serves as the basis for further development, and that can be changed only through formal change control procedures. A change to a baseline requires a CR, which ensures that an appropriate entity addresses implications to cost, schedule, and technical baselines for the project. A specific version of a single CI by itself or a set of functionally related CIs can be established as a baseline.

A baseline is established at the proper time—namely, when the CI is mature, stable, and has been reviewed and agreed upon by all stakeholders. As CIs change during product development, a series of baselines is established to enable assessment of the evolving product’s maturity at different points in time. This task includes identifying:

  • Events that establish a baseline
  • Items to be controlled in the baseline
  • Procedures used to establish and change the baseline
  • Authority required to approve changes to the approved baselined items

During each iteration / phase of a development project, newly developed items and new versions of pre-existing items may be identified as CIs. At the close of each iteration or phase, approved CIs may be baselined for the project.

Configuration Control Boards

CCBs manage and control changes to the controlled items. CCBs review, approve, disapprove, defer, escalate, or remand CRs for CIs. CMS formally charters its CCBs with specific thresholds for their change approval authority. Security-related changes can only be approved by an ECCB (or higher authority). CM ensures that all updates, deletions, and additions to baselined CIs are performed only as an outcome of the change control process.

CCB membership consists of management and stakeholders, and is supported by subject matter experts (SME). A CCB may exist at the enterprise and/or project level, with an approved charter and operating procedures. A cross-section of disciplines need representation on a CCB.

When a CR impacts multiple baselines that may be the responsibility of other CMS CCBs, it is necessary to coordinate the assessment and approval processes among the affected CCBs.

Change Requests and Continuous Changes

Concepts related to continuous changes such as DevOps move towards the ideal state of continuous standard changes. That is, the ideal state is a continuous flow of small, discrete, low impact changes.

Any request to change baselined CIs must be documented in a Change Request form, continuous standard changes. There are at least two classes of changes:

At a minimum, the CR should contain the following fields:

  • CR requester
  • CI to be changed
  • Class of change (e.g., normal, emergency)
  • Description and reason of proposed change, including rationale, and purpose
  • Affected baseline
  • Analysis of impact on the project and other entities
  • Historical resolution
  • Approval or disapproval
  • State of the change (e.g., open, approved / rejected, implemented, and tested)
  • Date closed

Impact Assessment

All stakeholders need to assess impacts against requested CRs. The assessments should include at least the following items:

  • Physical, functional, and performance characteristics of CIs
  • System, subsystem, and product interfaces
  • Security issues
  • Cost (may be against a cost threshold)
  • Schedule (Master Schedule or equivalent)
  • Shared data by systems
  • Environmental Impacts

Security CRs

In addition to project or program CRs described above, security-impacting CRs have additional requirements—in particular for the conduct of security impact assessments—imposed by HHS and the federal government as follows:

  • Security Impact Assessment (as required by HHS Policy for System Security and Privacy Handbook requirement, P-CM.8)
  • Security Impact Analysis (as required by NIST SP 800-53, Recommended Security Controls for Federal Information Systems and Organizations and the CMS ARS)
  • FAR Subparts 39.1, 42.3, and 52.248-3 (available from https://www.acquisition.gov)

Updated Baselines

The CCB establishes an initial baseline for a CI once it is deemed mature and the stakeholders have approved it. Any update of the baseline requires submission of a CR, completion of an impact assessment, CCB approval of the requested change, and implementation of the change. The process for updating a baseline may take days, weeks, or even months depending on the complexity and degree of anticipated impact.

Full Life-Cycle Configuration Management

Configuration Management continues during the full System Life Cycle. Continuous changes made to the baseline as part of DevOps results in large number of standard changes that take place during the O&M phase. The main emphasis during O&M is on change control, although other CM sub-processes are important.

The CCB is involved in CM during O&M because all items are baselined and under change control. Since the systems are in operation and subject to changes and security infrastructure threats, CM of system security issues and artifacts is a major activity during O&M.

Once a system reaches the retirement and decommissioning lifecycle stages, configuration management continues to provide value by identifying critical items that must be dispositioned appropriately.

Note that while a system may be decommissioned, its configuration managed assets may fall under data archival laws, regulations, and rules. Projects should coordinate with CMS OSORA to determine data archival requirements for CM data.

 

DATA AT CMS

Data Management

Enterprise Data Environment Overview

These sections of the CMS Technical Reference Architecture (TRA) consolidate guidance and policy information regarding data — its management, storage, and consumption by users and application systems.

The content is organized within the following sections:

Introduction to Data Management

The Centers for Medicare & Medicaid Services (CMS) Enterprise Data Environment (EDE) data strategy addresses data management and the Enterprise Data Mesh. The strategy seeks to incorporate shared costs across multiple CMS Centers to build shared services. It depicts a notional framework that organizes data management capabilities, data governance policies, and data user support services. It includes core components that support CMS’s infrastructure and enterprise shared data. The overall framework seeks to keep costs down while encouraging data reuse, better data quality, faster DevOps, advanced security management, and improved adaptability. In addition, this framework supports part of the risk-based management framework and the requirements of National Institute of Standards and Technology (NIST) Special Publication (SP) 800-37 Rev. 2., Risk Management Framework for Information Systems and Organizations: A System Life Cycle Approach for Security and Privacy

Effectively securing and managing enterprise data is accomplished by implementing consistent data management methods to include data governance, architecture, quality, and security, as outlined in the HHS Policy for Enterprise Data Management guidelines. Data architecture and consistent data management methods should include, but are not limited to, scalable solutions, minimizing data redundancy, and considerations for data virtualization to explore data synchronization and integration needs.

Monitoring and managing data security are essential to protect the confidentiality, integrity, and availability (CIA) of data. CMS policies and procedures must include data security requirements that comply with Department of Health and Human Services (HHS) and federal mandates. Such mandates include, but are not limited to, required controls to protect data collected and shared across the enterprise, to maintain comprehensive inventory of databases and their contents, to encrypt sensitive data at rest and in transit except data approved for public release, and to protect data against unauthorized use. Data sharing agreements must comply with HHS policy, Office of Chief Information Officer (OCIO) Rules of Engagement for Security, Monitoring, and Collaborative Systems, NIST SP 800-47 Rev. 1, Managing the Security of Information Exchanges, and NIST SP 800-53 Rev. 5., Security and Privacy Controls for Information Systems and Organizations

Major Related Services and Data Sources

 PREFERRED - CMS has invested heavily in the maturity of these solutions, and strongly recommends their use where feasible.

Integrated Data Repository Cloud

The Integrated Data Repository Cloud (IDRC) is a high-volume data warehouse integrating Medicare claims—Parts A, B, C, D, and Durable Medical Equipment (DME) with beneficiary and provider data sources, as well as such ancillary data as contract information and risk scores. This robust, integrated data supports much needed analytics across CMS.

IDR services include:

  • State-of-the-art capabilities for business intelligence and reporting, along with additional data access capabilities
  • Automated Finder File and Data Extract Process
  • Data dictionary, data limitations information, and source-to-target mappings
  • Customer support and assistance

IDR Enterprise Data Product (EDP)

The IDR Enterprise Data Product (EDP) supports the functions of the decommissioned Enterprise Data Mesh (EDM) using the same Snowflake-based capabilities of the IDR. This robust, integrated data supports business intelligence and analytics across CMS.

Center for Medicaid and CHIP Services (CMCS) DataConnect

CMCS DataConnect (Internal Link) is an all-in-one analytics platform for the Center for Medicaid & CHIP Services (CMCS). DataConnect is built on Databricks and Amazon QuickSight (sites no longer active on Confluence) dashboards. It provides read-only access to an expanding set of enterprise datasets, integrated with tools.

Center for Medicare and Medicaid Innovation (CMMI) Analysis and Management System (AMS)

The Center for Medicare and Medicaid Innovation (CMMI) Analysis and Management System (AMS) (Internal Link) provides a business intelligence and reporting tool that provides key information about CMMI models and demonstrations. This includes each model’s design, quality measures, and participating providers. The business intelligence tool is CMS Tableau.

CMS Master Data Management

The Master Data Management (MDM) system is a CMS enterprise shared service. MDM performs Identity Resolution on multiple CMS sources to provide singular, consolidated, and ID-resolved authoritative sources of data for use within CMS and by external agencies and organizations.

 

Enterprise Data Sharing & Governance

The Centers for Medicare & Medicaid Services (CMS) is required to protect the integrity and privacy of its enterprise data, whether within CMS authorization boundaries or outside them. Enterprise data includes data containing PII and/or PHI, but also includes sensitive and proprietary information that CMS must protect. Even public information for which CMS is authoritative must have its integrity maintained. This section discusses various use cases for data sharing.

  • Among CMS authorization boundaries with the same authorizing official (CIO). Managed with data use MOUs.
  • Between CMS and its contractors, Application Development Organizations (ADOs), and researchers. Managed with DUAs
  • Between CMS and other Federal Agencies. Managed with ICSAs and CMAs
  • Between CMS and State & Local organizations. Managed with ICSAs and CMAs

Existing Guidance

The CMS TRA currently contains these business rules that relate to data governance, but each is stated in a limited context:

  • BR-F-5: Any System That Processes CMS Data Must Be Covered by a CMS ATO
  • BR-CCIC-01: Security Authorization of Systems
  • BR-EFT-12: CMS Data in ATO’d Environments May Not Be Transferred to Non-ATO’d Environments
  • BR-EFT-13: CMS Data May Not Be Transferred Outside of CMS Processing Environments without a Prior Agreement
  • BR-SAAS-2: SaaS Must Have a CMS ATO

 PREFERRED - CMS strongly recommends that all CMS data remain within CMS authorization boundaries, except for public data released by CMS. Any CMS data used outside a CMS boundary must be protected in accordance with CMS privacy and security requirements and data release policies, as specified in a Data Use Agreement (DUA) or other governance measure. Colloquially, the preference is for this data to remain “inside CMS firewalls.” Its replication to external contractor facilities may only be permitted with additional governance measures.

Business Rules

BR-DG-1: All CMS enterprise data must be stored within a CMS authorization boundary

This is a corollary of BR-F-5. CMS data may be shared under the terms of a data governance vehicle (see below), but any persistent storage must be within a boundary authorized by the CMS CIO. If under an ATO from a different authorizing official, it is still subject to security controls required by CMS. The only exception being public data released by CMS.

Rationale:

CMS is required to protect the integrity and privacy of its enterprise data. Whether a SaaS or PaaS FedRAMP environment — or other contractor-owned/ contractor-operated facility, this requirement still applies. In addition to PII and PHI, CMS must manage other types of sensitive information.

RP-DG-2: Any CMS enterprise data sharing beyond CMS authorization boundaries should “share-in-place” where feasible, for example, a workspace or a remote API , avoiding file export or replication

Cloud-based deployments can now mitigate the capacity, processing, and network limitations that made it necessary to copy entire datasets for local processing. A workspace, accessed remotely, can provide analytics for particular views or datasets. Multiple workspaces could also help to segregate costs for individual external organizations.

Rationale:

Implementation of sharing as a virtual workspace or API enables dynamic authorization, which is a key Zero Trust principle. CMS capabilities that support this are mature enough for most use cases.

References

Essential data governance vehicles include:

  • CMS Data Use Agreement (DUA) defines how Protected Health Information (PHI) will be disclosed to organizations requesting data from CMS. Applicable CMS TRA Business Rule: BR-EFT-13
  • CMS Information Exchange Agreement (IEA) for Business Owners and Privacy Advisors working together to determine the terms of sharing PII with other federal or state agencies
  • CMS Interconnection Security Agreement (ISA) defines the relationship between CMS information systems and external systems. Applicable CMS TRA Business Rule: BR-CCIC-01
  • HHS Computer Matching Agreements (CMA) is created when CMS records are matched with records from another Federal or State agency and the results of such match may have an adverse impact on an individual in relation to a Federal benefit program.

Further information is found at Access to CMS Data & Application: CMS Contractor Data Communications Support Policy.

 

Enterprise Data Business Rules

These Enterprise Data business rules provided in this topic serve as the CMS standards and conventions for implementing CMS data mesh solutions. They previously appeared with the decommissioned EDM, but apply in general to CMS data mesh solutions.

BR-DM-1: CMS TRA Compliance 

All production data mesh systems must comply with the CMS TRA.

BR-DM-2: Data storage is to be separated from compute

Separating the data from the application and analytic tools that access it means that the data does not have to be duplicated for each organization that wants to use it. Data contributors can focus on managing their data, while users are allowed to work with tools that are native to their understanding. This supports “share-in-place” implementations.

BR-DM-3: Data assets are not to be copied or moved

The Enterprise Data Mesh does not seek to centralize the data. Using a very light footprint, the data mesh works with the contributors to align their existing data to a set of common standards and integration patterns. The data mesh exposes those datasets through a centralized metadata catalog, which makes it accessible and discoverable based on permissions.

BR-DM-4: Shared data assets are to be registered in Snowflake and a user-facing data catalog where available

Rather than moving their data to a central location, data contributors simply publish the information about their metadata to the IDR Snowflake Metadata Layer. The compute metadata catalog allows data sets to be automatically discovered by database tools.

BR-DM-5: The data mesh does not share raw data or unstructured data. All data in the EDM is fully structured and immediately consumable

A data mesh is closer in design to the industry term Data Lakehouse and customizes the design and approach to CMS requirements using Data mesh and Data domain principles.

BR-DM-6: Data sets are to remain within the data owner’s security boundary

Because the data is not moved or copied, the data remains within the data owner’s security boundary and under the data contributor’s control. Only the metadata is exposed to the Data Layer. The data contributor remains in control of who can access their data.

BR-DM-7: Data owners are required to curate their data assets and manage freshness and usability

Data contributors continue to manage their data throughout its lifecycle (curate) in its current location, as they always have done. As part of their data management, they will also update information about the data as things change.

BR-DM-8: Data consumers bring their own compute resources

“Separation of storage from compute” means that Data Consumers can point their own tools (computes) at different types of storage, accessing the data wherever it lives, rather than having to load it all into one database. Each user is allowed to bring their own skills and explore the data in ways that make sense to them. Our goal is to democratize our data. Democratizing data means making data accessible to the average non-technical user of information systems, without having to require the involvement of IT.

BR-DM-9: The data owner is responsible for determining the users, groups, roles, and policies that govern data access

In the System to System Design Pattern, the consuming system is responsible for implementing the role based security required by the data owner. But in the End User to Data Mesh Design Pattern access roles are applied to the data and maintained by the data owners.

 

Enterprise Data Facilities

Integrated Data Repository Cloud (IDRC)

The Integrated Data Repository Cloud (IDRC) is a high-volume data warehouse integrating Medicare claims—Parts A, B, C, D, and Durable Medical Equipment (DME) with beneficiary and provider data sources, as well as such ancillary data as contract information and risk scores. This robust, integrated data supports much needed analytics across CMS. In addition, the IDRC now incorporates the CMS Enterprise Data Mesh (EDM) in the Enterprise Data Product (EDP). The IDRC is managed by the Division of Enterprise Information Management Services (DEIMS).

IDR Cloud services include:

  • State-of-the-art capabilities for business intelligence and reporting, along with additional data access capabilities, including integration with CMS Business Intelligence Tools
  • Automated Finder File and Data Extract Process
  • Data dictionary, data limitations information, and source-to-target mappings
  • Customer support and assistance

IDR Cloud data access capabilities include:

  • Snowflake Data Shares
  • Snowsight web interface for data load
  • Bring-Your-Own-Computer (BYOC) flexibility
  • Cloud Data Dictionary and Catalog
  • Data-as-a-Service APIs

IDR Cloud Architecture

The IDRC High-Level Architecture diagram below shows the three major components of the IDRC:

  • IDR Cloud Data Lake
  • IDR Cloud ETL
  • IDR Cloud Data Warehouse

IDR Cloud Major Components (page 39)

The IDRC Conceptual Architecture diagram below shows the IDRC's integration:

IDRC Conceptual Architecture (page 38)

 

Integrated Data Repository (IDR) Enterprise Data Product (EDP)

The IDR Enterprise Data Product (EDP) (Internal Link) leverages the Snowflake External Table feature to reference external data product(s) maintained by the Data Product Owner(s) and Contributor(s) natively in cloud storage such as AWS S3. In addition to External Tables, the IDR EDP also contains data contributed via Snowflake Data Shares, providing access to data in other Snowflake accounts without requiring data replication. This investment and the existing Snowflake feature, enables the IDR to immediately onboard any new CMS data product(s).

The IDR EDP features include:

  • Extends Data Mesh capability
  • Integration with CMS Business Intelligence Tools
  • Seamless search, query, and join across all objects
  • User access is controlled by EUA job codes

IDR EDP Architecture

The IDRC High-Level Architecture diagram below shows the Snowflake metadata layers of IDR, including the Enterprise Data Product:

  • EDPs — Enterprise Data Product(s)
  • ADMs — Analytic Data Mart(s)
  • VDMs — Virtual Data Mart(s)
  • IDR Proper: Database Schemas

IDR EDP Architecture Context (page 37)

 

Center for Medicaid and CHIP Services (CMCS) DataConnect

CMCS DataConnect is an all-in-one analytics platform for the Center for Medicaid & CHIP Services (CMCS). DataConnect is built on Databricks and Amazon QuickSight (Links have been removed from CMS Confluence) dashboards. It provides read-only access to an expanding set of enterprise datasets, integrated with tools and the DataConnect Query Engine (Internal Link). Existing data insights include Dashboards (Internal Link).

Information about Data Quality

The Data Quality (DQ) Atlas has data quality and usability assessments using T-MSIS data on Medicaid and CHIP enrollment, claims, expenditures, service use, and more. These can help stakeholders to determine whether the data can meet their analytic needs. You can explore data by topic and state.

Medicaid and CHIP Data

The table below lists the datasets that are currently available. Find current and complete details at Get Started with Medicaid and CHIP Data

Datasets Available through DataConnect
DatasetNameDescription
TAFT-MSIS Analytic FilesAn enhanced version of T-MSIS (Internal Link & Password Required) data tailored for research.
T-MSIS (Internal Link)Transformed Medicaid Statistical Information SystemA relational database that agencies use to submit their Medicaid and CHIP data.
PI  (Internal Link)Performance IndicatorsMedicaid and CHIP eligibility and enrollment activity.
SEDS (Internal Link)Statistical Enrollment Data SystemThe unduplicated number of children ever enrolled in the federal fiscal year.
NPPES (Internal Link)National Plan and Provider Enumeration SystemHealthcare providers and their unique identifiers.
DMFDeath Master FileDeaths that were reported to the Social Security Administration.
CCS/CCSRClinical Classification Software / RefinedProcedure and diagnosis codes classified into clinically meaningful categories.
CARTSCHIP Annual Reports Template SystemCHIP program changes, goals, operation, financing, challenges, and accomplishments.
GeocodingTAF Beneficiary and Provider Geocoded AddressesGeographical coordinates for other datasets to analyze local and regional differences.
Eligibility Processing ReportPublic Health Emergency UnwindingEligibility and enrollment to address after the COVID-19 public health emergency (PHE).
FDBFirst DatabankA third-party provider of drug and medical device databases.
UPLUpper Payment LimitHow states expect to reimburse hospitals, nursing facilities, clinics, and other providers.
MCPARManaged Care Program Annual ReportAnnual state reports with MCO, PIHP, and PAHP data.
MMAMedicare Prescription Drug Improvement and Modernization ActData exchange between Medicaid agencies and CMS on dually eligible beneficiary status.
EPSDTEarly and Periodic Screening, Diagnostic and TreatmentComprehensive and preventive health care services for Medicaid children under age 21.
CAPCorrective Action PlanNarrative of steps taken to identify the most cost-effective actions to correct error causes.
MLRMedical Loss RatioAnnual state reports disclosing insurance companies expenditure.
MFPMoney Follows the PersonState and territory data on long-term services and support (LTSS) and home- and community-based services (HCBS).
Medicare FFSMedicare Fee-for-ServiceParts A and B service and billing information for Medicare reimbursement and enrollment.
Medicare MBSFMedicare Beneficiary Summary FileMaster list of all Medicare beneficiaries.

Upcoming Datasets

DataConnect is continually adding new datasets to become the trusted one-stop-shop for all Medicaid and CHIP data. DataConnect will be adding the following datasets in the future:

  • COREset Measures (Core set of health care quality measures)
  • State Vital Records Data
  • CHIP Data Budget (CMS-21B)
  • State Portfolio Tracking Tool (SPoTT)
  • Waiver Management System (WMS)
  • Medicare Enrollment and Eligibility Data
  • Medicaid Drug Rebate Program (MDRP)
  • Provider Link Key (PLK)
  • T-MSIS Analytic File - Research Identifiable Files (TAF RIF)
  • Medicaid Program Budget Report (CMS-37)
  • Annual CHIP Expenditures Report (CMS-21)
  • Medicaid Expenditure Data (CMS-64)
  • Medicaid Model Data Lab (MMDL)
  • Federal Upper Limit Program (FUL)

 

CMMI Analysis & Management System

The Analysis and Management System (AMS) (Internal Link) was created by the Center for Medicare and Medicaid Innovation (CMMI) to complement existing tools to provide analytics and reporting on the CMMI model portfolio. AMS is the single repository for these innovation models, programs and initiatives. Its analytics and reporting functionality supports CMMI leadership and stakeholders.

AMS provides the following key information (Internal Links):

 

The AMS data flow is shown in the diagram below:

AMS Data Flow Diagram (page 36)

 

Research Data Assistance Center (ResDAC)

ResDAC (Research Data Assistance Center) provides technical assistance to researchers interested in CMS Medicare and Medicaid data. Various data sets are available; however, access is restricted to approved research requests. Public use data files (which contain no protected information) are available via data.CMS.gov.

Some of this data is provided in the Chronic Conditions Data Warehouse (CCW), which supports the CMS Virtual Research Data Center (VRDC) analytic environment. CMS data products are managed by the CMS Office of Enterprise Data and Analytics (OEDA).

 

Master Data Management (MDM)

Master Data Management and Identity Resolution

The Master Data Management (MDM) system is a CMS enterprise shared service application with a focus on eliminating redundancy, inconsistency and fragmentation of CMS data and increasing efficiencies. MDM performs the function of Identity (ID) Resolution:.

  • ID Resolution is the process of identifying and merging records from different data sources to create and maintain a single trustworthy view of a person, organization, or entity.
  • MDM’s ID Resolution relies on probabilistic algorithms to match and link records to deliver an accurate and complete view of the same entities that organizations can trust.
  • MDM’s ID Resolution helps CMS to quickly perform analyses on unified critical data improving decision-making and customer experiences; thus, delivering value from ID-resolved data products.

Master Data Management and Identity Resolution

Master data management (MDM) harnesses the use of identity resolution to link sets of records to a unique entity even in the presence of inconsistent or ambiguous identifying attribute values

  • Records are matched by comparing the degree of similarity
  • Matched records above a threshold similarity score are presumed to represent the same entity
  • Linked entities are assigned a unique “enterprise identifier” (EID)

Records can be linked:

  • Within a Data Source (e.g., Duplicate records for the same provider in PECOS)
  • Across Data Sources (e.g., Determine that a PECOS provider is the same as an NPPES provider)
  • Across Programs (e.g., Determine that a Medicare provider is the same as a Medicaid provider)
  • Across States (e.g., Determine that a Virginia provider is the same as a Maryland provider)

This enables MDM to:

  • Create an index containing same-source and cross-source linkages
  • Leverage an EID that is persisted in produced master indexes (e.g., the PMI/SPP data products)
  • Provide access to ID-resolved data from one location via Cloud services, APIs, and BI reporting

MDM is scheduled to be retired in February 2027. Support of this data management functionality is migrating to the CMMI Analysis and Management System (AMS), Integrated Data Repository (IDR) Cloud, IDR Enterprise Data Product (EDP), and CMCS DataConnect. ID Resolution will be performed as needed within these systems, as well as by the Center for Clinical Standards & Quality (CCSQ) and the Center for Program Integrity (CPI).

For more information about MDM services, contact Office of Information Technology (OIT) / Enterprise Architecture Data Group (EADG) / Division of Data Enterprise Services (DDES) at MDMTeam@cms.hhs.gov.

File Transfer

File Transfer Introduction

This chapter provides guidance for “File Transfer,” which is the process or action of transferring files from one system to another and governs transferring files among CMS datacenter and cloud environments as well as between CMS environments and external partners. “Enterprise File Transfer (EFT),” a.k.a. Electronic File Transfer, as defined in this chapter, is a product or system purpose-built to perform file transfers between systems, applications, and platforms by employing encryption and authentication to maintain data integrity and confidentiality. It is used to securely share large volume of data.

Any operating system can transfer files directly without the aid of an EFT system. In complex or challenging applications, EFT systems may provide advanced scheduling, workflow, and management of file transfers, support multiple sources and destinations, offload other system components, and improve the scalability, reliability, auditing, and security of file transfers.

 PREFERRED - CMS strongly recommends that all CMS data remain within CMS authorization boundaries, except for public data released by CMS. Colloquially, the preference is for this data to remain “inside CMS firewalls.” The implication is that the need for outbound file transfer ought to be limited where possible.

The concepts, strategies, and guidelines in this chapter align principally with the EFT Team’s EFT User Guide, Version 2.1, May 9, 2023.

Although the foregoing document defines the details of specific EFT products and services at CMS, this chapter provides guidance and business rules for implementing file transfer processes at CMS with or without an EFT product.

File Transfer and Enterprise File Transfer

File Transfer includes all file transfers between CMS data centers or cloud environments (i.e., intra-CMS file transfers) as well as those between CMS environments and external partners. In many situations where scale, security, or reliability are a concern, file transfer may be implemented using EFT products and services.

Enterprise File Transfer, a.k.a. Electronic File Transfer, refers generically to enterprise file transfer products and technologies used or implemented by CMS applications and data processing environments to transfer files. In a few cases, this chapter explicitly refers to specific EFT products by name or to the “CMS EFT System” or the “CMS EFT Infrastructure” when discussing the shared EFT services CMS provides across the Agency.

EFT products and services support management, authentication, verification, scheduling, logging, and/or notification and post-processing of file transfers between organizations or destinations. File transfers do not necessarily require use of an EFT service. Using an EFT product or service may provide needed capabilities such as error-detection, retransmission, multiple destinations, and audit trails.

The CMS ePortal and other CMS external portals offer file upload services for accepting files from external users, and file download services for allowing external users to request and receive files. Some portals are also capable of offering “managed” file download and upload services that use client-side software to ensure a successful transfer.

Scope

This chapter addresses all file transfers between CMS data center or cloud environments (i.e., intra-CMS file transfers) as well as those between CMS data centers and external partners. This chapter does not apply to transfers occurring within the authorization boundary of a given system . File transfer between production and non-production environments, as well as upload and download services for end users are outside the scope of this chapter.

 

File Transfer Business Drivers

As a participant in this nation’s healthcare system, CMS needs the capability to securely and reliably exchange data files with its business partners, government agencies, and other stakeholders. Files may be transferred electronically, or in some instances, using physical media such as encrypted digital tapes and disks.

File-oriented processing often requires a sequence of file transfers followed by application processing. For example, a file received from a partner is processed by an application and the results are sent to another partner for subsequent processing.

The CMS EFT system is the primary mechanism for coordinating the workflow of file transfer and subsequent application execution from various CMS business partners. Applications may also implement or use other file transfer solutions, such as the file transfer capabilities of the CMS ePortal.

CMS Enterprise File Transfer is a critical enabler of major programs like Medicare and Medicaid. These exchange data with a wide variety of business partner organizations, large and small. New legislative or regulatory changes may drive transfers with new partners and changes in data transfer volume.

Goals

The business goals of an effective, secure file transfer infrastructure include:

  • Exchanging files with CMS partners
  • Adapting to changes such as new partners, new applications, and new CMS services
  • Securing all data transfers in accordance with the current CMS ARS on an ongoing basis

Facilitate Exchanging Files

The primary business goal of an EFT system is to facilitate the secure exchange of data between CMS and its partners as well as between partners (as a pass-through) consistent with current CMS ARS requirements.

The CMS EFT team is responsible for all in bound and outbound file transfers that involve the CMS EFT infrastructure.

Adapt to Change

The CMS EFT infrastructure enhances productivity by isolating customers from changes in physical transfer locations. The Sweeps system used by CMS EFT is a store-and forward-transfer rather than a simple point-to-point transfer. This allows customers to be unaffected by changes to the other side of the transfer. This can be either a new server location, such as migrating from the mainframe to the cloud, or a new transfer product such as moving from Connect:Direct to SFTP. The sender and receiver may also each use different transfer products.

Objectives and Processing Environment Requirements

CMS data centers manage file exchange in accordance with prescribed management, security, and operational objectives. This CMS Processing Environment handles more than 1,200 customers and sites and an estimated three million transfers per month. In such an environment, robust automation is essential.

Managing Audit Trails

Maintaining an audit trail of file transfers is an essential capability of an EFT system. Audit trails are used in diagnosis as well as security. Both the HIPAA and HITECH acts require traceability of Protected Health Information for Covered Entities and Business Associates (CMS is a hybrid entity). Audit logs, including audit trails of all EFT administrator activity, must adhere to the current CMS ARS requirements, undergo regular review, ensure non-repudiation, and be appropriately retained.

From a security and privacy perspective, the audit trail is useful in determining when and what happened in a sequence of events, the party’s identities, and which data are involved. By examining audit logs, it is possible to determine when files were transferred, which trigger scripts were executed, and whether such activities were successful. When working with CMS partners, it is useful to know when events occurred to help diagnose problems at either end of a transaction.

Coordinate Application Execution

Coordinating application execution is a critical function of any EFT infrastructure. An EFT workflow management system performs this function by detecting successful in-bound file transfers and coordinating with the batch scheduler to execute corresponding applications. The mapping of data file to trigger script is recorded in the EFT routing tables.

Managing EFT Processing Issues

Detecting, reporting, and managing EFT processing errors to resolution is another critical capability. An EFT infrastructure can report the status of file transfers and detection of errors that occur during file transfer. An EFT infrastructure does not, however, detect or manage application processing errors; error reporting of application processing errors is an application responsibility.

Report on the EFT Process

It is important that the EFT system produce reports about the file transfer and application triggering process (when controlled by the EFT) that include information such as timestamps of file transfer, application triggering, file name, file sizes, and destinations. Reports can be produced on an on-demand and scheduled basis, with delivery via email or online.

Decouple File Transfer from Application Execution

One objective of EFT is to decouple file transfer operations from application execution. Without an EFT service, partners transferred files and were responsible for initiating trigger-script execution. With a CMS-managed EFT service, the control of initiating trigger-script execution may remain within CMS rather than relying on a partner to initiate application execution.

Prevent Data Overlays

Some EFT solutions, such as the CMS EFT service, introduce file name timestamps during a file renaming process that occurs at the end of file transfers. Timestamping helps retain data integrity by preventing accidental data file overlays.

Keeping Archival Copies

Applications are responsible for archiving copies of files transferred. Archiving files sent to external customers is needed to save communications subject to financial or legal review. Archiving also eliminates the need to repeat application execution to generate a file that was previously generated, as well as a variety of administrative duties such as resetting databases to prior condition, restoring backups, and other tasks that re-processing would entail.

The CMS EFT infrastructure saves copies of any file received or transmitted via Sweeps for up to 7 days. This allows for re-transmittal of files by the EFT admins without an application resending the file to Sweeps, which provides the benefits of archiving for a short term for all files transferred. The CMS EFT infrastructure is not responsible, however, for long-term archival or records management — this is an application owner’s responsibility. Files transferred using the Store-and-Forward or Pass-Through mechanisms are not archived.

Business Rules for File Transfer 

To guide EFT within the CMS Processing Environments, CMS developed the following business rules.

BR-EFT-1: (Deprecated after TRA 2018R1): File Naming Conventions Must Be Obeyed

BR-EFT-2: Limited Protocols Are Permitted in the CMS EFT System

Table Permissible File Transfer Protocols lists the only protocols permitted in the CMS EFT system.

Table - Permissible File Transfer Protocols
ProtocolVendor or StandardDescription
HTTP/SOpen StandardHypertext Transfer Protocol / Secure, using a Web browser to download or upload files
S/FTPOpen StandardSecure File Transfer Protocol (FTP) over Secure Shell (SSH) protocol
Connect:DirectIBM SterlingProprietary secure file transfer protocol (not compatible with FTP, S/FTP, or FTP/S)
TIBCO Managed File Transfer (MFT) ServerTIBCOProprietary secure file transfer protocol (not compatible with FTP, S/FTP, or FTP/S)

New users are encouraged to use S/FTP as the preferred EFT protocol.

Rationale:

The CMS EFT system only supports the listed protocols.

BR-EFT-3: Use Mailboxes between Business Partners Only

CMS allows EFT mailboxes to store collections of files transferred between external business partners and CMS. CMS prohibits using mailboxes as work areas or as a communication facility between internal CMS applications. Pushing files to the partner with SFTP is preferred.

Rationale:

The EFT mailboxes consume resources and are not designed as work areas. Once information is transferred, the mailboxes must release storage resources.

BR-EFT-4: Registration Is Required When Using the CMS EFT System

If using the CMS EFT system, each new file must be registered with the CMS EFT system authority to ensure that the routing tables are modified correctly and that the correct jobs will execute upon file identification.

Rationale:

The EFT infrastructure requires pre-registration of the files to be transferred to ensure proper routing.

BR-EFT-5: Internal Integrity Validation Is an Application Responsibility

The CMS EFT infrastructure guarantees proper transfer of files using permitted protocols. Validating the internal integrity of a file is the responsibility of the receiving application. The EFT infrastructure does not inspect the internal contents of a file to determine data integrity.

Rationale:

The CMS EFT infrastructure has no knowledge of the internal data schemas for files transferred. Thus, it cannot identify schema problems. It only guarantees that the same bytes of information are transferred.

BR-EFT-6: File Encryption Is an Application Responsibility

The CMS EFT infrastructure encrypts the transport mechanism (typically TLS), but does not encrypt (or decrypt) the data files individually. Encrypting data files must be done by application owners themselves and must, on an ongoing basis, be consistent with current CMS ARS requirements and relevant federal law, regulations, and executive orders. (Please refer to BR-EFT-11.)

Rationale:

This rule removes the burden of performing file-level encryption from the EFT infrastructure. The EFT infrastructure does not manage file encryption keys.

BR-EFT-7: Secured Transmission Is Required

Files must be transmitted over a CMS ARS-approved secure protocol, as determined by the current version of the CMS ARS at the time of transmission, such as TLS v1.2 (or later). Unsecured protocols such as FTP or telnet must not be used unless tunneled over a secure protocol, as permitted by the CMS ARS.

Rationale:

File transfer software must be configured to ensure secure file transmission to prevent inspection by unauthorized parties.

BR-EFT-8: IP-Based File Transfer Protocols Only

IP-based file transfer protocols are now mandatory. All currently approved protocols are IP based.

Rationale:

CMS infrastructure no longer supports non-IP based network protocols.

BR-EFT-9: (Deprecated, replaced By RP-EFT-1 after TRA 2018R1): Transfer Large Files Using Check-pointing Protocols

BR-EFT-10: Encrypt Files Residing in EFT Mailboxes

Files sent to or received from mailboxes must be encrypted by the application (not the EFT infrastructure) in accordance with the current CMS ARS-approved cryptography rules and stored in encrypted form while residing in any EFT mailbox.

Rationale:

Because external parties can access these files, the files must be encrypted at REST while residing in the EFT mailboxes.

BR-EFT-11: CMS Data Files May Only Be Transferred to the Data Zone

CMS data can only persist in a zone within the multi-zone architecture which has been appropriately secured to house sensitive data. Within CMS Data Centers, CMS data cannot be transferred using EFT to a system’s Application or Presentation Zone because these zones are not allowed to host persistent sensitive data. In the cloud environment, where explicit zones are not as clearly defined, the implementer must ensure the appropriate security mechanisms have been applied. These mechanisms include authentication, authorization, security certificates, routing rules, and filtering as discussed in CMS Services Framework Mediation Principles.

Rationale:

This is an extension of CMS TRA rules about storage of persistent data outside of the Data Zone. Sensitive data may only persist in an appropriately secured zone (data zone). Sensitive data may exist temporarily outside the data zone for processing, such as data validation, prior to persistent storage. If longer duration processing is required, consider placing the processing system in the Data Zone (CMS data center) or a zone appropriately secured for sensitive data.

BR-EFT-12: CMS Data in ATO’d Environments May Not Be Transferred to Non-ATO’d Environments

EFT cannot be used to move CMS data from an environment with an ATO to an environment that does not have an ATO unless that data has been de-identified.

Rationale:

This rule is intended to prevent data movement to lower environments. Only ATO’d environments have been formally verified to meet stringent CMS security controls. Projects or applications should not be seeking to move data from an ATO’d environment to a non-ATO’d environment without first consulting the CMS TRA and CMS ARS rules.

BR-EFT-13: CMS Data May Not Be Transferred Outside of CMS Processing Environments without a Prior Agreement

CMS data may only be transferred under a Data Use Agreement, Memorandum of Understanding, or similar agreement that establishes how third parties may use data.

Related CMS ARS Security Controls include: CA-3 - Information Exchange, CA-3(6) - Supplemental: Transfer Authorizations.

Rationale:

CMS data may only be shared with third parties under agreement unless the data has been identified as public data.

RP-EFT-1: Transfer Files Using Check-Pointing Protocols When Performance Impacts Are a Concern [Formerly BR-EFT-9]

In some circumstances, file transfers may be susceptible to failures and require resending. Repeated retransmissions may impact system or network performance. When such impacts pose a significant concern, an engineering and risk-based decision may be made to use check-pointing protocols to transfer files. Check-pointing protocols allow for interrupted transfers to be restarted from the last confirmed checkpoint rather than from the beginning of the file. This reduces the burden on CMS networks and systems and helps ensure a successful transfer.

Rationale:

A check-pointing protocol ensures much more efficient file transfer for both sender and receiver. It also reduces the burden on CMS networks and systems. Check-pointing works if supported and permitted by both the client and the server.

Data Storage

Data Storage Services

Data Storage Services Introduction

This chapter addresses Data Storage Services generally as any form of online data storage, whether physical, virtual, shared, dedicated, or cloud. The services may include SAN, NAS, object stores, remote or network backup devices, and databases (when used to store files). The Data Storage Services, as defined in this chapter do NOT include dedicated, locally attached storage devices.

A special focus in this chapter is on a subset of data storage services called “File-Level Storage Services” or “File Storage Services,” which store data as files.

Storage Service Types, Access Methods, and Terms

This chapter classifies all data storage services and access methods as follows in Tables 10 and 11. The important terms used in this chapter are:

  • Storage Service Client – Any component that directly accesses a data storage service. A storage service client may be a component of a business application, or a component of a storage mediation service. The component may be a driver, a data access layer, a utility, or a service.
  • Storage Mediation Service – A service that uses a data storage service to store, retrieve, or manipulate files to meet the storage needs of one or more business applications. A storage mediation service hides from its clients the API and/or file directory structure of the data storage service and implements its own access controls.

Storage Service Types

This chapter classifies all data storage services into the generic storage types described in the table below: Generic Types of Data Storage Services.

Generic Types of Data Storage Services
Data Storage Service TypeDescriptionDirectory and Access ControlAccess Enforced By
Tape / Archival StorageUsed for offline storage, files are catalogued; requires multiple steps to locate and retrieve a given file in a catalog and in the storage media.The service maintains the catalogs and locates files. File-level access control may not be available.The service.
Block StorageThe device or partition is mounted by a customer’s OS, which implements a file system.The customer’s OS implements a file system on the block device or partition.The customer’s OS provides access controls over the files; the storage service may or may not provide access control over the block device or partition.
File-Level Storage (a.k.a., Object Storage)Files are stored as file or blob objects.The service maintains a file system to locate and retrieve files.The service.

 

Access Methods

This chapter classifies all methods of accessing data storage services into the following five types of generic access as shown in Generic Access Methods for Data Storage Services.

Generic Access Methods for Data Storage Services
Access Method TypeDescriptionDirectory and Access ControlAccess Enforced By
Block Device Mount

The remote storage is mounted or mapped to appear as a block storage device or drive to the client OS.

Example protocols: iSCSI

The customer’s OS implements a file system.The customer’s OS provides access controls over the files; the storage service may or may not provide access control over the block device or partition.
Network File Share

May be used to access individual files or directories via URIs.

The remote storage may also be mounted or mapped to appear to the client OS as a device or drive with a file system.

When mounted, it can be multiuser, single user, or single session within the customer’s OS or browser.

Example Protocols: NFS, CIFS, FTP, WebDAV

The service implements the remote file system.The service enforces access control, and authorized users can change the access controls through the service.
Mirrored Network File Share

Similar to a Network File Share, except that a mirror copy on the client’s local drive periodically syncs with the remote storage and may be used when the remote storage is unavailable.

Example Protocols: rsync or local client software using Network File Share protocols.

The service implements the remote file system. The file system of the local mirrored copy is managed by the customer OS and may be a different file system format.The customer’s OS manages access control at the file system level; the service enforces access control.
Query-based

The service is queried like a database for files matching given criteria.

Example protocols: SQL, ODBC

The service maintains a file system to locate and retrieve files.The service enforces access controls. Authorized users using the service may change the metadata of files including access controls.
Uniform Resource Locator (URL)-based

Same as the Network File Share and Query access methods but using a URL.

Example Protocols: WebDAV, SMB, FTP, Amazon S3

The service implements the remote file system.The service enforces access control, and authorized users can change the access controls through the service.

All five access-method types are network based and describe shared, dedicated, or cloud storage services. Regardless of the method of access used, access to the file storage service, and to the CMS files managed by that service, must be limited to authorized authenticated users. The service must comply with all CMS TRA requirements in the effective Defense in Depth of CMS systems and CMS files.

Securing Access

In all storage service access methods, except for the Block Device Mount method, the access control is tied to user identities and/or roles defined in a directory domain used by the storage service. The directory domain may be local to the storage service or the CSP providing the storage service, or the storage service may use the CMS LDAP service.

To comply with the CMS TRA, users from external networks, CMSNet, or the CMS LAN must not have direct access to the CMS files persisted by the storage services in a Data zone and must use intermediate services in a Presentation and/or Application zone to access the persisted file data indirectly. This requirement applies for other services needing access to files. For example, an intermediate service may provide access to a temporary copy of the persisted file, with the temporary copy in the same zone as the intermediate service.

In addition, users from external networks, CMSNet, or the CMS LAN must be authenticated to access CMS files, unless those files are intended for the public. Therefore, the intermediate services or the storage service itself must authenticate users using a CMS directory domain.

A file storage service must be configured to deny access to all unauthenticated users and all unauthorized users. If the intermediate services, or the storage service itself, cannot directly use a CMS directory domain for authentication, there must be a CMS-approved trust relationship or federation mechanism between the service’s local or CSP-provided directory service and a CMS directory domain service.

The File URL topic has business rules concerning defense-in-depth protections against unauthorized access to the CMS files managed by file storage services.

Some file storage services offer a way to address individual files using a network URL. If the URL is accessible from external networks, CMSNet, or the CMS LAN, additional defense-in-depth protections are required as described in File URL.

The reader should refer to other related CMS TRA and CMS ARS requirements that, in addition to the guidance presented here for File Storage Services, prohibit direct access to CMS data, require data encryption at rest and in transit, require files be stored in the Data Zone, prevent access to the Data Zone from other than the Application Zone, and prevent network traffic from traversing zones.

Storage Service Concerns and Considerations

System designs that include Data Storage Services should consider and address the following issues at a minimum:

  • Mutual authentication. In the default configuration for most Data Storage Services, the service trusts its authenticated clients and users but the clients and users do not have a way to authenticate the service. Mutual authentication or other mechanisms should be considered even in circumstances when the CMS TRA does not require mutual authentication.
  • Storage mediation services. A storage mediation service can hide the storage service interface and file system directory from other business application components. A storage mediation service or other mechanisms should be considered even in circumstances when the CMS TRA does not require storage mediation.

A common misconception is that the CMS TRA always requires separation of storage clients from storage services that use network zones and network gateways or firewalls. Although separation is a good practice, an acceptable alternative is to keep the client and storage service together in the same zone and allow the client to be accessed as a mediation service by other CMS services. By eliminating a gateway between the client and service, the alternative offers the following advantages:

  • Improved performance and reliability
  • The client’s credentials for accessing the service are not stored or exposed outside the immediate zone.
  • The storage service’s interface and API are not exposed outside the immediate zone.

Naturally, when the storage service is external, it must be separated from storage clients by gateways and firewalls. There can be other circumstances where separation may be technically necessary or where it is easier to implement defense-in-depth mechanisms. The choice becomes an engineering tradeoff.

Please refer to RP-DSS-3 for additional related guidance.

Impeding Attacks

CMS TRA guidance and recommendations do more than lock down vulnerabilities and minimize attack vectors—,they also impede the progress of successful attacks. Impeding means to limit how far into CMS systems or data stores an attack may reach, or to force the attacker to engage in activities that increase the likelihood of detection. For storage services, the threat is an attack that succeeds in gaining access to a business application server or process that, in turn, has an interface with a storage service or a storage mediation service. With enough determination and sophistication, such an attack may eventually succeed in gaining access to CMS files, but it is more likely the attack will grab what is readily available from the compromised business application server or process. Therefore, impeding the progress of a smash-and-grab attack should be a design consideration for any system component that has direct access to a storage service or to a storage mediation service.

To protect the storage service from a smash-and-grab attacker on a compromised business application server or process, CMS requires establishing the following attack impedance priorities:

  • Prevent the attacker from detecting the existence of the storage service.
  • Prevent the attacker from having file system-like access to the files, such as through a mirrored, mounted, or mapped network file share.
  • Prevent the attacker from having access to the file system directory, so the attacker does not know the existing files or directory structure.

To do this, on all application servers, CMS requires limitations on the users, local user, processes, or sessions that have access to a mounted or mapped network file share or block storage device. CMS prohibits mounting storage services at boot time or allowing all OS users and processes access to mounted storage services. When possible, the network file share should be mounted only within a user session, or only within a process, or only by specific non-root users.

This same principle of least privilege also applies to file transfer services. If an application uses a file transfer service, CMS recommends against using root or other privileged user accounts. Any user account should only see the least number of files necessary, and in the case of file upload, the user should only see files that the user has uploaded. In the case of file download, the user account should only see files the account is authorized to see.

Example Design Patterns

The following commonly used design patterns or approaches provide access to files managed by a File-Level Storage Service while maintaining compliance with related CMS TRA requirements. Other approaches may be possible, but the most common examples are:

  • Mediated. Configure the File-Level Storage Service so the files are only available to an application server in the Application Zone. Configure the File-Level Storage Service to deny all except for a CMS-defined role held only by the application server. User requests are made to the application server, which retrieves a copy of the file for the user into either the application server’s memory or the application server’s dedicated local virtual block storage drive and file system. The copy of the file must be deleted or transferred to the File Storage Service when the user session is complete or the file is no longer needed by the application server.
  • Federated. Configure the File-Level Storage Service to deny all except for a CMS-defined role. Then use federation (e.g., SAML) to give CMS users (authenticated through ePortal, for example) a temporary credential with the CMS-defined role that exists only during the user session.

The Mediated and Federated example approaches have these important characteristics:

  • The File-Level Storage Service is configured to deny all except for a CMS-defined role.
  • No one can access the file without first authenticating to CMS.
  • Data encryption at rest and in transit is still required.
  • An intermediary application accesses the files managed by the file storage service.

If the File-Level Storage Service permits file access by URL, the URL may safely be open to external networks (if necessary) when unauthenticated, and unauthorized disclosure via those networks is still prevented by ACLs, encryption, and other defense-in-depth mechanisms.

Data Storage

Data Storage Services Business Rules

This topic provides a core set of Data Storage Services business rules that are binding on all CMS business applications. To achieve their intended purpose, these business rules and recommended practices depend on compliance with BRs found elsewhere in the CMS TRA and controls in the CMS ARS. The BRs for this topic address Data Storage Services (DSS), File Uniform Resource Locators (URL), and Common Commercial Services, such as Amazon Web Services (AWS).

The business rules and recommended practices in this chapter depend on adherence to CMS ARS and CMS TRA mandates, and have the following specific dependencies:

Configure File Storage Services to Deny All Access by Default

Consistent with CMS TRA Network Services, BR-ACID-1 and CMS ARS Security Control AC-6 - Least Privilege, File-Level Storage Services must be configured to deny all access by default, and only permit authorized access. This least privilege rule is here for emphasis because failure to properly configure cloud-based file-level storage services is the reason behind many recent highly publicized data thefts. The Example Design Patterns described above rely on this rule to function properly with CSP storage services.

Rationale:

The CMS ARS requires systems to be configured for least privilege. Disclosure of CMS data to unauthorized users may result in harm to individuals, the Agency, or the government.

Deny Unauthenticated Access to Non-public CMS Files

CMS TRA Network Services, BR-ACID-1 and BR-ACID-2 prohibit unauthenticated access to non-public CMS files from external networks, CMSNet, or the CMS LAN. Unauthenticated downloading may be permitted if the files are intended for download by the public. To enforce this BR and maintain compliance with related CMS TRA requirements, file-level storage services should be configured to deny access to unauthenticated users, and an intermediary application must be used to access the files through the file storage service.

Rationale:

Disclosure of CMS data to unauthorized users may result in harm to individuals, the Agency, or the government.

Deny Anonymous Access to Non-Public CMS Files

CMS must know the identity of users accessing CMS files. CMS systems must be designed to prevent anonymous access to CMS files from external networks, CMSNet, or the CMS LAN, and to log the identity of external users or systems that access CMS files.

Rationale:

Refer to CMS TRA Network Services, BR-ACID-2.

In some environments, it may be possible that a file-level storage service and associated logging services may not know the identity of a user accessing a file even though the user is authenticated. Anonymous access may be a result of the systems’ architecture or a recent change in the user’s profile such as their name or email address. Anonymous downloading may be permitted if the files are intended for download by the public.

For further discussion about the causes of authenticated, but anonymous users accessing files, please refer to “Anonymous or unknown people in a file” and “Anonymous animals” in the online support and documentation for Google Drive or Google Docs (https://support.google.com/drive). The problem is not exclusive to Google services, but their support documentation provides insights into how legitimate users may appear to an application or file log as authenticated but anonymous users.

Data Storage Services

CMS developed the following business rules and recommended practices for Data Storage Services within the CMS Processing Environments.

BR-DSS-1: Do Not Use UDP for Data Storage Services or File Transfer Services 

Do not use the User Datagram Protocol (UDP) for File Storage Services or File Transfer Services.

Rationale:

UDP can be easily spoofed or intercepted.

RP-DSS-2: Use a Trust Relationship or Federation for Authenticated Access to CMS Files

To enforce BR-FSS-1 and maintain compliance with related CMS TRA requirements, configure file-level storage services to deny access to unauthenticated users and use an intermediary application to access the files managed by the file storage service. If both the intermediary application and the file storage service are unable to authenticate users directly with the CMS LDAP services, a CMS-approved directory trust relationship or federation method may be used to provide CMS users with temporary authentication to access the CMS files managed by the file storage service. Users must first authenticate using CMS credentials and a CMS authentication service before receiving temporary access to the CMS files managed by the file storage service.

Rationale:

A CMS-approved directory trust relationship or federation method enables CMS to know and log the identity of users accessing CMS files.

RP-DSS-3: Avoid Unnecessarily Separating Storage Clients and Storage Services into Different Zones

Using zones to separate storage clients from storage services may not be necessary. Mounting a storage service in one zone to a storage service client in another zone negates an important goal of CMS zone data separation because the storage service interface (API and file system directory) would be the same in both zones. Therefore, having a storage service client in a different zone from the storage service does not hide the storage service interface from the client’s zone, which is a goal of the CMS TRA Multi-Zone Architecture. Any protections this separation may provide may be outweighed by performance issues or other considerations. An exception is external storage services, such as those provided by remote storage providers or CSPs.

Please refer to Storage Service Concerns and Considerations for additional related guidance.

Rationale:

CMS recognizes that, in some circumstances, the design tradeoff decision for security and performance may be worth the risk to allow location of storage clients and storage services in the same zone.

RP-DSS-4: External Storage Services Are External to Any CMS Zone

External storage services, such as those provided by remote storage providers or CSPs, are external services. They may be accessible from a storage client in a zone, but are not themselves in a CMS zone, and must be treated as external services for CMS TRA compliance purposes.

Outbound communications with external services must traverse the CMS zones in the CMS TRA Multi-Zone Architecture. An exception may be allowed for a CSP-provided storage service when the storage client is in the same CSP hosting environment or if there is a secure VRF between the storage client and the storage service.

Rationale:

To be supplied.

RP-DSS-5: A Storage Service Data Store Should Not Be Accessible from More Than One Zone

A storage service data store (such as a single volume, bucket, or directory) should not be accessible from more than one zone. To provide access to data from more than one zone, a storage mediation service may be used. A storage mediation service may serve clients in multiple zones.

Rationale:

If the service is accessible from multiple zones, then it becomes a potential conduit for data transfer between zones. Having multiple storage mediation services for the same data store is allowed but may not be a best practice for data integrity or performance.

File Uniform Resource Locators

CMS developed the following BRs for File URLs within the CMS Processing Environments.

BR-URL-1: Authentication Is Required to Access a Non-Public CMS File

Unauthenticated downloading may be permitted if the files are intended for download by the public.

Rationale:

CMS must know the identity of users accessing CMS files. Securing Access above describes some approaches to systems designs that comply with this BR.

BR-URL-2: Authentication and Authorization Are Required to Upload a File to a CMS Location Referenced by a URL

Rationale:

CMS must know the identity of users sending or uploading files to CMS.

BR-URL-3: Logs Must Identify the Users Who Access a CMS File Referenced by a URL from the External Networks or CMSNet

Rationale:

CMS must know the identity of users accessing CMS files.

BR-URL-4: Logs Must Identify the Users Who Upload a File to a CMS Location Referenced by a URL from the External Networks or CMSNet

Some remote storage services allow the username of a user to change; therefore, the username alone is not adequate for logging purposes. Logged data from that service should include a user identifier that is not subject to change.

Rationale:

CMS must know the identity of users sending or uploading files to CMS.

BR-URL-5: URLs May Not Include Passwords or Decryption Keys

URLs may not include passwords or decryption keys.

Rationale:

CMS cannot control the protection of URLs in external networks or non-CMS environments, nor can it ensure only authorized use or sharing of URLs within CMS office automation systems. URLs with passwords or decryption keys pose the risk for unauthorized access to CMS services or data and may disclose sensitive passwords or decryption keys.

RP-URL-6: Use Signed URLs to Indicate the Source and Creation Time of a CMS File

Signed URLs for CMS files may be used to indicate the source, creation time, or other characteristics of a CMS file. A signed URL provides the user with some confidence that CMS is the source of the file. Including other characteristics of the file may help the user in processing the file or determining its validity. Signed URLs do NOT provide user authentication as required by BR-URL-1.

Rationale:

Signed URLs are an optional mechanism that may aid the user in processing CMS files or determining a file’s validity.

RP-URL-7: Use Time Limits and/or Click Limits to Control How Long a CMS File Is Available to Unauthenticated External Users for Download through a URL

Unauthenticated downloading without time limits and/or click limits may be permitted if the files are intended for download by the public. For non-public files, it is preferable to set time limits and/or click limits when providing external users with an externally accessible URL for a CMS file.

Although authentication is required to access a CMS file referenced by a URL (as specified by BR-URL-1), time limits and/or click limits provide additional Defense-in-Depth by impeding unauthorized users from taking advantage of a misconfiguration or compromised user account to try to download a CMS file.

CMS permits storing sensitive, PII, or PHI data temporarily in the Application Zone. Temporary files and cached data must be removed from the Application Zone once the transfer or processing is confirmed and successful or time limits and/or click limits have been exceeded.

Rationale:

Limiting the number of clicks on a URL (or the file downloaded) helps prevent unauthorized users from accessing the file. Having a time limit on how long the file is available for download ensures the file becomes unavailable to unauthorized users even if the authorized user never uses the URL to download the file.

Common Commercial Services

This topic provides vendor-specific guidance for implementing CMS applications using certain commercial data storage service providers. This guidance neither endorses nor recommends using these providers. The guidance in this topic will be expanded in future CMS TRA releases to include additional data storage service providers frequently used by CMS applications.

Amazon Web Services – S3

The Amazon Simple Storage Service (S3) can provide CMS systems with a cost-effective means for storage access and retrieval. This topic is based on Research Spotlight Technology Review – Amazon S3 Guidance, CMS TRB Technical Topics, updated 26 April, 2023.

 PREFERRED - The CMS preferred solution is the CMS Cloud implementation. Information and guidance is available at:

The Amazon S3 service can be used in either of two ways: for internal storage of resources in the CMS enclave, or for external data dissemination open to the Internet for access.

Internal Storage within the CMS Enclave

S3 storage is attached to the CMS Zonal environment. The S3 storage may be attached to the Presentation, Application, or Data Zones. S3 must adhere to all CMS TRA and CMS ARS security requirements.

External Data Dissemination

S3 storage may also be used to disseminate data to external parties, including access to data directly from the Internet. Any decision to provide direct access to data, without the defense-in-depth mechanisms provided by the CMS Zonal architecture, should be made with great care. CMS should consider all factors, such as the type of data (e.g., sensitivity) and types of users (public vs. a limited known set of interfaces), before deciding to provide access to the data directly to the Internet.

CMS is committed to providing the highest level of protection for the data. This includes ensuring the perception of security. If a non-sensitive document interface were hacked, the perception would be that CMS security failed, even if no sensitive data were lost.

All development groups should also consider user policies in addition to limiting the availability of the URL to the S3 resource.

AWS S3 Business Rules

CMS developed the following business rules for Amazon Web Services within the CMS Processing Environments.

BR-AWS-1: Encrypt Amazon S3 Data at Rest 

The CMS ARS prescribes that data must be encrypted at rest, which applies to data in AWS S3 storage. AWS S3 must be configured (in the console) because it is not a requirement from AWS to encrypt. The configuration for encryption should be validated using cloud validation services such as CloudTrail. Since January 2023, by default, Amazon S3 automatically encrypts all new objects added on buckets on the server side.

Rationale:

CMS recommends that teams use AWS services for encryption. CMS recommends that the development teams employ the SSE-KMS AWS service, a service that combines secure, highly available hardware and software to provide a key management system scaled for the cloud. KMS allows creation and management of master keys.

BR-AWS-2: The S3 Storage Must Be Attached / Accessible from Not More Than One Zone

S3 storage may be attached to the CMS Zonal environment and must comply with CMS ARS and CMS TRA requirements. If sensitive data is attached to a zone which does not meet the security posture to store sensitive data, then the sensitive data may only exist in that S3 storage temporarily, while being processed (e.g. file scanning) before moving to an appropriately secured zone (i.e. one where access has been restricted via appropriate security challenges.)

Commentary:

For example, S3 storage can be used to support data storage requirements in the Data Zone, and additional S3 services could be attached to the support the Application Zone. Additional recommendations:

  • Assigns S3 to a specific zone. This can be accomplished by setting up access to S3 to be restricted by subnet. Each of the CMS zones within the Virtual Private Cloud (VPC) is its own subnet.
  • Limit the configuration of the S3 and the requirement for access by zone. This configuration and access MUST be validated during the configuration of the VPC and as part of the monitoring tools of the VPC (e.g., CloudTrail). In addition to restricting the access by zone, the developers should consider additional security mechanisms.

BR-AWS-3: Do Not Allow Cross-Zone Access to S3

The TRA requires that S3 storage only be attached to a specific zone (see BR-AWS-2). Multiple resources from the same zone type (e.g., data) are allowed to connect to the S3, but cross-zone access is prohibited. Resources are considered within the same zone type when the security posture for access is the same. For example, CMS prohibits allowing the Application Zone to place data in S3 and then allowing the Data Zone access to the same data because the security posture for application zone resources is not the same as the data zone. A data zone S3, housing sensitive data, can only be accessed from resources that have met the security requirements for accessing sensitive data.

BR-AWS-4: External Content in S3 Requires a Data User Agreement Process

For access by external sources, projects must follow the established Data User Agreement (DUA) process.

BR-AWS-5: Users Must Be Authenticated by AWS S3 Using a CMS-Managed IAM Userid When Accessing Data in CMS S3 Buckets

To access data in CMS S3 buckets directly through the S3 API or via an S3 URL, users must be authorized and authenticated by AWS S3 using a CMS-managed AWS Identity and Access Management (IAM) UserID. CMS must provision and manage this UserID. CMS is responsible for vetting the user’s identity and authorizing the user. Depending on security requirements for the data and application, multi-factor authentication (MFA) or network restrictions may also be required.

Users from external networks, CMSNet, or the CMS LAN must not have direct access to the CMS files persisted in a Data Zone.

A strongly preferred alternative to permitting user access to CMS S3 buckets directly through the S3 API or via an S3 URL is to provide a CMS service that mediates access and performs the storage and retrieval of data in S3. Such a service would not necessarily require users to be provisioned with IAM userids.

Anonymous read-only access to public data in S3 via an S3 URL is permitted without authentication.

For additional related guidance, please refer to BR-URL-1, RP- BR-URL-2, AWS-1, RP-AWS-2, BR-AWS-6, and BR-AWS-7.

BR-AWS-6: External Data Upload to S3 Is Not Permitted

CMS allows S3 to be used for storing collected data only after that data was submitted through a secure collection process.

BR-AWS-7: No Direct Internet Access to an Application’s Main S3 Storage

CMS prohibits all direct Internet access to an application’s main S3 storage. All Internet-based access to S3 Storage may implement a dedicated and separate “S3 bucket” to host the data temporarily. The data should only temporarily exist in the S3 bucket and be accessed by a short-lived URL. For example, if CMS is granting access (via the short-lived Internet URL) the data should first be copied to a new temporary location, and then the URL generated. Upon access, or expiration of the URL availability time window, the data object should be deleted from this temporary storage.

When hosting static content on Amazon S3, configure CloudFront distribution to restrict access to an S3 bucket so that users can access objects only through the distribution. This can be done by using an S3 REST API endpoint as the origin. Restrict access to Amazon S3 by setting up an origin access identity (OAI), a special CloudFront user that is associated with the CloudFront distribution. Add permissions on the S3 bucket to allow access to only this OAI. When the users access the S3 objects through CloudFront, the OAI gets the object on behalf of the users. If users try to access the Amazon S3 URL directly, their access is denied. This makes sure that the client can access objects in the S3 bucket but only by CloudFront.

BR-AWS-8: Remove Data from S3 When They Are No Longer Needed

Unused, unmanaged, and unmonitored data in S3 buckets may become targets of abuse. It is more important in shared or public cloud environments to remove unneeded data and storage than in traditional or private data centers. When unneeded data becomes untraceable and orphaned, it may continue to incur costs beyond the life of an application or even a line of business.

BR-AWS-9: Remove Unneeded S3 Buckets

Unused, unmanaged, and unmonitored S3 buckets and their contents may become targets of abuse. It is more important in shared or public cloud environments to remove unneeded data and storage than in traditional or private data centers. When unneeded data becomes untraceable and orphaned, it may continue to incur costs beyond the life of an application or even a line of business.

AWS S3 Recommended Practices

CMS recommends that service developers adhere to the following recommended practices to ensure the most effective implementation of AWS S3.

RP-AWS-1: Use Amazon AWS Security / Access Features for S3

AWS S3 security/access features include, but are not limited to:

  • Bucket policies and rules that apply broadly across all requests to the S3 resources
  • Identity and Access Management policies to assure fine-grained control to their Amazon S3 bucket or objects while also retaining full control on what the users do
  • Link to DUA Account either through capturing the DUA number at time of registration or through IAM logging policies
  • Access Control Lists (ACL) and query string authentication
  • With ACLs, grant specific permissions (i.e., READ, WRITE, FULL_CONTROL) to specific users for an individual bucket or object
  • Enable Amazon Macie for monitoring data security and privacy.
  • Use VPC Endpoints for S3 Access: Amazon S3 bucket policies are defined to control access to buckets from specific VPC endpoints, or specific VPCs. The VPC endpoint routes requests across the Amazon network to S3 and then routes responses back to the VPC, ensuring that the traffic stays on the Amazon network.

RP-AWS-2: Use Short-Lived URLs for External Access to S3

CMS recommends limiting the exposure of the data by using short-lived URL access to S3. This includes configuring a time limit and a specific number of clicks for which the URL is valid.

Amazon provides the capability to limit the time window a S3 URL is available. CMS recommends restricting a time within which the URL is available to minutes. If the data must be available for a longer period, use additional mechanisms, such as encrypting the data with a password that is provided to the end user through a medium other than the URL.

CMS recommends limiting the number of allowable clicks to access the S3 URL. CMS suggests limiting that access to one click, assuring that once the URL has been accessed, it is no longer available.

These parameters should be explicitly defined and established by the business requirements.

Business Intelligence

Introduction

Purpose

This chapter defines the Agency’s enterprise-wide initiative to provide a consolidated, secure gateway to the wealth of CMS data where users can leverage business intelligence (BI) software tools to access, manipulate, analyze, and share integrated data and then create information they need to ask business questions and derive accurate, clear answers. This chapter articulates the BI guidance and standards that should be used by CMS and CMS Contractor partners for all CMS Processing Environments.

 PREFERRED - While CMS supports all of these tools, CMS prefers that project teams consider these factors:

  • Use case
  • Cost at scale
  • Ability to reuse developed capabilities for the given data source(s)

Scope

The concepts, strategies, and guidelines discussed in this chapter align with the CMS Business Intelligence Strategy, Version 1.5, December 9, 2008 that prescribes the BI tools and processes forming the conceptual framework of the BI environment. This chapter inherits the strategies, guidelines, and capabilities from the CMS Business Intelligence Strategy, Version 1.5, December 9, 2008; CMS Cognos ReportNet Guidelines, March 3, 2006; and CMS MicroStrategy 8 Guidelines, March 16, 2006.

This chapter is also updated with content from the Technical Review Board (TRB) Research Spotlight Data Analytics & Business Intelligence (BI) Tools, May 5, 2025.

This chapter provides architecture guidance, incorporates best practices, and defines the infrastructure of the BI Environment using components compatible with the CMS TRA. It includes the BI Reference Architecture and services that form a blueprint for helping CMS build BI solutions. This document represents the CMS Business Intelligence Reference Architecture in two virtual views—the “Business View” and the “Technical View.”

Why Data Analytics & BI Are Important

Data is an asset, and Analytics​​​/​BI extracts value from it. Both data analytics and BI are used interchangeably, with BI being the generalized term encompassing analytics, but there are some distinctions. Data analytics is the process of primarily collecting, inspecting, cleansing, transforming, storing, modeling, and querying data. Its goal is to produce insights that inform decision-making. There are 4 main types of data analysis:

  • Descriptive: Informs that an event “A” occurred
  • Diagnostic: “A” occurred because of an event “B”
  • Predictive: What could be the future of “A” if “B” continues
  • Prescriptive: What is the best course of action?

While BI provides Descriptive and Diagnostic and can be Predictive based on history only, it doesn't take future trends into consideration. Thus, BI provides a progress report, while data analytics also provides data-driven insights into what are the prescriptive changes to progress.

CMS applications collect and store vast amounts of data. When this data is presented visually in a graphical form, it is easier to interpret, understand, and quickly observe data patterns than to query the data and parse the results. This is the power of data visualization and the primary reason for teams to use it. Data visualization is an integral part of BI, and it helps teams to visualize their data and interact with them. This makes it easier for business users to spot patterns and trends in a much better way. Here are some of the reasons why Analytics and BI are crucial for any application.

  • Ability to gain customer insights
  • Greater visibility on business operations
  • Get actionable insights
  • Improved efficiency across the division / group / center
  • Real-time data availability
  • Better marketing efforts
  • Gives the business a competitive advantage

Hence, it is beneficial for CMS Systems and Business Owners to analyze the need for data analytics and BI and to look closely at the various tools used at CMS.

Business Intelligence Environment

CMS has successfully implemented the Agency’s enterprise-wide BI Environment that provides a consolidated, secure gateway to the wealth of CMS data where users can leverage a suite of standard COTS BI tools. These BI software tools are made available for use by CMS offices and centers, analysts, developers, project managers, external researchers, law enforcement, partners, government contractors, and the public.

 PREFERRED - While CMS supports all of these tools, CMS prefers that project teams consider these factors:

  • Use case
  • Cost at scale
  • Ability to reuse developed capabilities for the given data source(s)

These tools have been integrated with various data repositories and frameworks. Some are for use by the enterprise and external entities, while others are specific to particular Centers or projects. The latter are called out because their capabilities can be replicated for other use cases.

CMS staff are directed to transition from SAS software to CMS CIO-approved alternative solutions by December of 2026.

CMS has also identified Microsoft Power BI as a preferred tool within the CMS Microsoft 365 SaaS tenant. It is also available in AWS and Azure.

The analytic and BI tools implemented and supported at CMS include the following COTS products (comparisons below):

  • Esri ArcGIS Enterprise provides an AWS Cloud based enterprise solution that enables geospatial mapping and analysis. ArcGIS is used by the Center for Program Integrity (CPI) to enable geographic mapping and analyses to help meet its business objective of finding and reducing fraud, waste, and abuse in Medicare and Medicaid. It uses geocoded data from the Integrated Data Repository Cloud (IDRC) Provider and Beneficiary data.
  • AWS Services – AWS provides Query services like Amazon Athena and Redshift Spectrum, data visualization tools like Amazon QuickSight, data warehouses like Amazon Redshift, and sophisticated data processing frameworks like Amazon EMR (Elastic MapReduce). Each of these services addresses different needs and use cases, and are used by multiple centers at CMS. The table below compares them and provides guidance to help project teams choose one or more services based on their requirements.
  • SAP Business Objects (SAP Business Intelligence Platform) supports the Center for Program Integrity (CPI), for example, using the CMS Integrated Data Repository Cloud (IDRC) and provides ad hoc and OLAP reporting capabilities.
  • IBM Cognos Analytics supports the Part D and Drug Data Processing System (DDPS), Medical Appeal System (MAS), Unified Case Management (UCM), and Eligibility Appeals Case Management Solution (EACMS) (among others) using the Teradata data repository and provides ad hoc and online analytical processing (OLAP) reporting capabilities.
  • Informatica PowerCenter and Data Engineering Integration (DEI) are data integration tools used to integrate data warehouses, data marts, and operational data stores (ODS) with the capability to extract, transform, and load (ETL) data from virtually any business data source in CMS. PowerCenter is used for ETL with DB2, Teradata, Oracle databases, and flat files.
  • MicroStrategy is a data analytics and business intelligence reporting tool that can generate interactive dashboards, high-end formatting in reports, scorecards, and several other features relating to the generation, sorting, and automated distribution of reports that provide insights into past, present, and future business trends. It supports Part B Analytics Reports (PBAR), for example, using the Teradata data repository and provides ad hoc and OLAP reporting capabilities.
  • Microsoft Power BI is a unified cloud-based business intelligence and analytics service that provides a full overview of the application's most critical data. By connecting to all the application data sources, Power BI simplifies data evaluation and sharing with scalable dashboards, interactive reports, embedded visuals, etc. while maintaining data accuracy, consistency, and security. It integrates with Microsoft 365, Azure, and AWS using data connectors. All users of CMS E-mail have a Microsoft 365 subscription that includes Power BI Pro. More information about Power BI is available at the CMSConnect / App Hub
  • Python, R, and Scala are programming languages that are very helpful for data scientists since they offer powerful libraries to support data science use cases. The BI tools team provisions CMS Python AWS WorkSpaces with connectivity to IDRC , ACO-OS (DB2), and Informatica server for teams to use. Users can utilize these virtual desktops to run Python scripts and big data computations without storing data on their local machine (see comparisons).
  • SAS Enterprise Business Intelligence (EBI) combines the strengths of SAS Analytics and SAS Data Management and provides business users with a powerful tool for making better decisions. It also boosts data consistency and streamlines the administration. SAS offers a unified platform for teams to prepare the data, analyze it visually, and build, operationalize, and manage data science, and AI/ML models in an augmented design experience. SAS is a statistical software suite for data management, advanced analytics, multivariate analysis, business intelligence, criminal investigation, and predictive analytics. CPI and HHS OIG use SAS for accessing IDRC and IDRC (Snowflake) to perform analytical reports to identify Fraud, Waste, and Abuse for CMS.

Note that CMS staff are directed to transition from SAS software to CMS CIO-approved alternative solutions by December of 2026.

  • Tableau provides all types of users with intuitive business intelligence (BI) tools to enhance data discovery and understanding. With simple drag-and-drop features, a user can easily access and analyze key data, create innovative reports and visualizations, and share critical insights across the agency. Tableau is currently used by multiple CMS centers including CMCS/DSG/DIS, CMMI, OA, OIT/ISPG, OIT/IUSG, and OSPR. The CMMI Analysis and Management System (AMS) uses Tableau as its primary business intelligence tool.

Business Intelligence Integrations and Platforms

These services provide integration with data sources for business intelligence and data analytics tools.

  • Databricks is a unified set of tools for building, deploying, sharing, and maintaining enterprise-grade data solutions at scale. The Databricks Lakehouse Platform integrates with cloud storage and security in CMS cloud and manages and deploys cloud infrastructure on the project’s behalf. Databricks is currently used by CMCS for its DataConnect. It is also available as Software as a Service (SaaS).
  • Snowflake is a cloud-native data warehouse platform that enables data storage, processing, and analytic solutions that are faster, easier to use, and massively scalable. Snowflake is currently used by OIT/EADG for the IDR Cloud program as well as the IDR EDP.
  • BI Workspaces are provided to individual data analysts and scientists within CMS with a versatile and tailored virtual desktop environment equipped with the necessary tools for data analysis, wrangling, and visualization. The technology stack and architectural choices can be customized to ensure a secure, efficient, and personalized analytical workspace for each user. The Workspace offering is designed as an all-inclusive AWS Windows or Linux workspace tailored to cater to the unique preferences and requirements of individual users. The architecture features AWS WorkSpaces, providing a secure, cloud-based virtual desktop environment.
    • These virtual desktops are integrated with Active Directory and Okta SSO for seamless authentication and authorization with CMS EUA and IDM user accounts.
    • The CMS EDM team ensures that connectivity is established to CMS Business Intelligence tools, as needed, to ensure data scientists and data analysts can generate reports utilizing the metadata and data accessed through the EDM.

Background

CMS collects and maintains a vast amount of healthcare data. With the successful implementation of the CMS Integrated Data Strategy, the strategic, tactical, and operational importance of this data is now solidified as a critical enterprise asset. The CMS IDR Cloud has become the primary, authoritative enterprise data asset through the consolidation and modernization of various CMS data warehouses, data marts, and applications.

CMS has implemented an enterprise-wide BI Environment that provides the front-end query, analytics, and reporting solutions that make data more accessible. These solutions also empower users to make important decisions more efficiently and with greater confidence. The CMS IDR Cloud and BI Environments are major components of the CMS BI Reference Architecture, providing shared access to consolidated, reliable data and information across the CMS enterprise as well as with other government agencies and external business partners. The CMS BI Reference Architecture provides more detail.

Objectives

The CMS user base needs access to an integrated, layered CMS BI Reference Architecture that meets the following objectives:

  • Integrates authoritative data from diverse internal and external sources into a primary repository for access and sharing by common user communities
  • Adheres to the current CMS Minimum Security Requirements (CMSR) at the Moderate security level, as published in the CMS ARS , and requirements of the privacy and data use agreement (DUA) processes
  • Ensures that source data collected in the common CMS data repositories meets Agency standards for data quality and consistency
  • Develops self-service methods for easy, quick, and secure access to BI data by providing users with tools for reporting and analyzing data, whether that data resides within the IDR Cloud or in applications and data stores outside the IDR Cloud
  • Enables use of metadata and common enterprise-wide semantics to fully employ the self-service model, giving users access to information about the data they use
  • Provides an environment with the capacity to analyze integrated authoritative data in a secure CMS BI Environment

CMS BI Production Environment

The current CMS BI Production Environment consists of a multi-zone architecture that aligns with the CMS TRA. The CMS BI Production Environment supports the Presentation, Application, and Data Zones and the specific access and protection requirements for meeting the needs of users in each zone. CMS Business Intelligence Production Environment (As-Is) illustrates the BI tools and processes that form the framework of the CMS BI Production Environment.

CMS Business Intelligence Production Environment (As-Is) (page 18)

The BI components are implemented in the CMS BI Production Environment in accordance with the CMS TRA Multi-Zone Architecture.

Non-Compliance with CMS TRA Multi-Zone Architecture

Although the BI tools implemented at CMS are enterprise-class BI software suites, the product manufacturers did not develop them to meet the unique demands and design requirements of the CMS infrastructure. None of the software suites, as implemented at CMS, are fully compliant with one aspect of the CMS TRA—i.e., data transport between the Application Zone and the Data Zone.

The CMS TRA requires that, “When an Application Zone application requires information or actions from the Data Zone, the application makes the request and processes the response using a ‘redundant messaging facility’.”

Instead of messaging, the BI software in use at CMS separates business processing and data processing by using ODBC or Java database connectivity (JDBC) as standard database access methods.

Only the manufacturers of the CMS BI software can alter this aspect of their products. The current versions of the implemented CMS BI tools have no accepted methods for mitigating this conflict with the CMS TRA Multi-Zone Architecture.

CMS Enterprise Portal Framework

In the current CMS BI Production Environment, each BI tool supports its own portal enabling application users to access the BI contents—the metadata, database data, queries, and reports—created for each specific BI application. Because information in the CMS BI production environment is distributed across many BI applications, information sharing and collaboration among CMS BI applications is quite limited.

CMS has implemented an enterprise-wide secure gateway, known as the CMS Enterprise Portal, that enables all BI software tools to access, manipulate, analyze, and share data. This portal provides a one-stop site where users can access and analyze data, helping them make important decisions more efficiently and with greater confidence. The BI Portal Framework leverages CMS’s Portal platform and architecture, providing aggregated reports and visualizations of data retrieved from various BI tools interacting with the CMS Enterprise Data Warehouse, which comprises any CMS data warehouse or data mart built for decision support.

CMS Enterprise Portal Framework illustrates the CMS Enterprise Portal Framework.

CMS Enterprise Portal Framework (page 19)

The CMS Enterprise Portal has the following additional features:

  • Aligns with the CMS Enterprise Portal Strategy
  • Integrates with the CMS Access Manager and CMS Enterprise Lightweight Directory Access Protocol
  • Migrates the existing BI Portal from Solaris 10 to Z/Linux Operating System
  • Implements single sign-on

The initial release of the CMS Enterprise Portal supports:

  1. Parts A, B, D, and Enrollment Dashboards
  2. Part B Analytics Reports
  3. National Level Repository (NLR), Health Information Technology for Economic and Clinical Health Act (HITECH) Electronic Health Record incentive program
  4. Affordable Care Act (ACA) Program Dashboard
  5. MicroStrategy Web
  6. Teradata ViewPoint

 PREFERRED - Enterprise and other Data Repository Frameworks

These provide access to source data repositories housed in the cloud, and have been extensively integrated with enterprise BI tools. Depending on use case, they can be used as appropriate, or their designs can be reused for similar capabilities.

  • The Integrated Data Repository Cloud (IDRC) is a high-volume data warehouse integrating Medicare claims—Parts A, B, C, D, and Durable Medical Equipment (DME) with beneficiary and provider data sources, as well as such ancillary data as contract information and risk scores. This robust, integrated data supports much needed analytics across CMS, as well as external agencies such as DOJ, FBI, OIG, etc.

It consists of a data lake, ETL components, and a Snowflake data warehouse. Access to the IDR Cloud includes SAS EBI and Snowflake user interfaces, as well as the Data-as-a-Service (DaaS) API (Internal Link).

  • IDR Enterprise Data Product (EDP) (Internal Link Password Required) is a new type of data object that uses the Snowflake External Table feature. The EDP incorporates the content of the former Enterprise Data Mesh Hive Metadata.
  • CMCS DataConnect is an all-in-one analytics platform for the Center for Medicaid & CHIP Services (CMCS). DataConnect is built on Databricks and AWS QuickSight dashboards. It provides read-only access to an expanding set of enterprise datasets, integrated with tools. Existing data insights include Dashboards and Data Spotlights (Internal Link Password Required).
  • CMS Master Data Management (MDM) is a CMS enterprise shared service application with a focus on eliminating redundancy, inconsistency and fragmentation of CMS data and increasing efficiencies. MDM provides a single point of access to a singular, synchronized, comprehensive, and ID-resolved authoritative source of Beneficiary, Provider, Organization, Program and Relationship data for use within CMS and by external organizations and agencies.

 

Business Intelligence Tools Comparisons

Below are comparisons of analytics and BI tools in use at CMS

CMS Enterprise BI Tools Comparison
CriteriaArcGISBusinessObjectsCognosMicroStrategyPower BISAS-EBI*Tableau
VendorEsriSAPIBMMicroStrategyMicrosoftSASSalesforce
Use Cases Best Suited forGeospatial mapping and analysisReporting, visualization, seamless integration and native connectivity within the SAP ecosystemReporting, data exploration, modeling, analytics, dashboards, scorecards, and event monitoring and metricsInteractive dashboards, scorecards, formatted reports, thresholds and alerts, automated report distributionUnified cloud-based business intelligence and analytics serviceStatistical software suite for data management, analytics, multivariate analysis, predictive analyticsIntuitive visual analytics for all user types, data discovery, predictive analytics
CMS User BaseDevelopers and users of maps and dashboardsDevelopers, data analysts, and report & dashboard consumersDevelopers, data analysts, and report & dashboard consumersDevelopers and report & dashboard consumersDevelopers, data analysts, and report & dashboard consumersData analystsDevelopers and report & dashboard consumers
CMS User Capacity
  • 150 Creator users
  • Unlimited Viewers
  • 2 ArcGIS Pro licenses in Amazon Workspaces
  • 5 ArcGIS Desktop Pro (CPI users)
  • Approximately 1,400 users
  • Concurrent user capacity 200
  • Approximately 5,000 users
  • Concurrent user capacity 100
  • Unlimited user license allocations
  • Approximately 3,500 users
  • Concurrent user capacity 100
  • User license allocation limited
  • All Microsoft 365 G5 subscriptions
  • No concurrent user limit
  • Approximately 1,200 users (Web and Citrix Microsoft Client)
  • Concurrent user capacity 200
  • SAS licenses are limited
  • Approximately 350 users
  • Concurrent Users Capacity 30
  • User’s Licenses Allocation:
    • 105 Creators
    • 18 Explorers
    • 545 Viewers
Web-based Tools used at CMS
  • ArcGIS Enterprise
  • ArcGIS Dashboards
  • ArcGIS StoryMaps
  • ArcGIS Web AppBuilder
  • SAP Business Objects
  • Web Intelligence Reporting Tool
  • SAP Lumira Dashboard Tool
  • Cognos Report Authoring
  • Cognos Query Studio
  • Cognos Dashboards
  • MicroStrategy Web
  • MicroStrategy Dashboard
  • MicroStrategy Document
Power BI Pro (SaaS)
  • SAS Information Delivery Portal
  • SAS Web Report Studio
  • SAS Studio

Web and Desktop tools

  • Tableau Desktop
  • Tableau Server
  • Tableau Prep Builder
  • Tableau Services Manager
CMS IntegrationCMS Enterprise IDMCMS Enterprise PortalCMS Enterprise PortalCMS Enterprise PortalMicrosoft 365 TenantCMS Enterprise LDAPCMS Enterprise LDAP and Okta MFA
Developer Access
  • ArcGIS Enterprise
  • ArcGIS Pro
  • ArcGIS for Adobe Creative Cloud
  • StreetMap Premium
  • AWS Workspaces for desktop tools
Windows based client tools on Citrix to create data structures (Universes), and Dashboards (Lumira)AWS workspaces with Okta MFAMicroStrategy Developer client toolPower BI Pro

Access via Citrix:

  • Enterprise Guide
  • Enterprise Miner
  • Information Map Studio
  • OLAP Cube Studio
  • SAS Add-in for MS Office
AWS workspaces with Okta MFA
Some CMS Applications & Data SourcesProvider and Beneficiary Geocode data located in IDR Cloud for CPI workflows usesIDRC Cloud
  • Medical Appeal System (Oracle)
  • System for Tracking Audit & Reimbursement (Oracle)
  • Unified Case Management (DB2 & Postgres SQL)
  • Eligibility Appeals Case Management Solution (Athena)
MDX (Oracle) Part B Analytical Reporting/IDR CloudAll sources available to Data Connectors
  • IDRC (Snowflake)
  • TMSIS (AWS Redshift and Databricks)
  • MSIS (DB2)
  • MQM (SQL Server)
  • HCRIP (Oracle)
  • OA/EW_Dashboard (Excel and CSV data)
  • OA/CFM (Amazon Redshift)
  • ISPG/CEDE (MS SQL Server)
  • HCDR (MS SQL Server)
  • DataConnect (Databricks)
Pricing Modelnamed userper userper userper userper user or per workspaceper userper user
CostEnterprise License Agreement funded by CPIsFunded by CPI, limitations on how many users outside of CPI can use itFunded by OIT/EADG, no cost to the projectFunded by OIT/EADG, no cost to the projectIncluded in Microsoft 365 subscriptionFunded by OIT/EADG, licenses are limited and assigned to existing usersFunded by OIT (ICPG & IUSG), licenses are limited

* SAS-Enterprise Business Intelligence (SAS-EBI) will be decommissioned in December 2026. CMS’s SAS EBI support ends in December 2026, and CMS has decided to migrate away from SAS. Project teams are urged to consider other great alternatives, such as Python and R, for their needs. CMS is providing SAS migration support service by offering the SAS2PY5 tool, which converts SAS code to Python and/or Spark SQL.

Comparison of AWS Services
CriteriaAthenaRedshift SpectrumQuickSightRedshiftEMR
What Is It?Amazon Athena is an interactive query service that makes it easy to analyze unstructured, semi-structured, and structured S3 data directly with SQLAmazon Redshift Spectrum allows efficient querying and retrieval of structured and semi-structured S3 data without having to load the data into Amazon Redshift tablesAmazon QuickSight is a fast, easy-to-use, business analytics service in the cloudAmazon Redshift is a fully managed, petabyte-scale data warehouse service in the cloud. Uses SQL to analyze structured and semi-structured data using AWS-designed hardware and machine learning (ML)Amazon Elastic MapReduce (EMR) is a cloud big-data platform to run distributed big data processing frameworks such as Spark, Hadoop, Presto, or Hbase
Use Cases Best Suited forAd-hoc interactive SQL queries against data on Amazon S3Employ massive parallelism to run very fast against large datasets in S3Interactive visualizations and dashboards, ad-hoc analysis and embedded analytics via APIs & SDKsA data warehouse solution that pulls data from many different sources / systems into a common formatCustom big data processing, analytics, ML
Setup & ManagementServerless: no setup or server managementServerless: no setup or server managementServerless: no setup or server managementFully managed service along with the capability to manage clusters/nodes. A serverless option is also available.Greater effort to setup & manage clusters and the software installed on them
PerformanceWell suited for small/medium datasets. Not suited for large datasetsWell suited for medium to large datasets – multiple clusters can concurrently query the same dataset in Amazon S3 without the need to replicate the data for each clusterAutomatically scale to tens of thousands of usersWell suited for large datasets. Resources are automatically provisioned and scaled based on needFull control & flexibility over the configuration of clusters & performance
Pricing ModelPay only for what you use – price per queryPay only for what you use – price per queryPay only for what you use – price per sessionPay for the provisioned capacity by the hour as long as the cluster is runningAmazon EMR cost is added to the EC2 cost (the underlying servers) and Elastic Block Store (EBS) usage
CostLow to MediumMediumMediumMedium to HighMedium to High

 

Comparison of Python, R, and Scala
CriteriaPythonRScala
What Is It?Python is a simple, open source, general-purpose language and is easy to learn. Many data analysis, manipulation, ML, and deep learning libraries are written in Python, and hence it has gained popularity in the big data ecosystem and is one of the de facto languages of Data Science.R is a language and environment for statistical computing and graphics and hence rightly called the language of statisticians. R Studio can be used for research, statistics, plotting, and data analytics applications. R is also used for building data models to be used for data analysis.Scala is a hybrid functional programming language since it has both object-oriented and functional programming features. Scala is a machine-compiled language that runs in a Java Virtual Machine (JVM). Scala is highly scalable and is the native language of Apache Spark.
Learning CurveEasiestModerateSteep learning curve
Used byBeginners & Data EngineersData Scientists/ StatisticiansBig Data Programmers
Use Cases Best Suited forData Engineering, ML, Data VisualizationData Analysis, Data Visualization, StatisticsSpark Native
Type of LanguageGeneral PurposeSpecifically for Data Scientists. Needs conversion into Scala/Python before productizingObject-Oriented & Functional General Purpose
ConcurrencyDoes not Support ConcurrencyN/ASupports Concurrency
Type SafetyDynamically TypedDynamically TypedStatically typed (except for Spark 2.0 Data frames)
Interpreted Language — Read-Evaluate-PrintLoop (REPL)YesYesNo
Mature ML LibrariesExcellentExcellentLimited
Visualization LibrariesExcellentExcellentLimited
Web Notebooks SupportJupyter Notebook SupportR NotebookApache Zeppelin Notebook Support
PerformanceSlowerSlowerFaster (about 10x faster than Python)
CostFree, open sourceFree, open sourceFree, open source

 

Data Usage

CMS Business Intelligence Reference Architecture

This CMS BI Reference Architecture presents a framework for developing CMS BI solutions, provides architectural guidance and standards, incorporates best practices, and defines the BI environment using components compatible with the CMS TRA. It serves as a blueprint for building CMS BI solutions.

This chapter represents the CMS BI Reference Architecture in two virtual views—the “Business View” and the “Technical View.” Each view is intended to communicate with a different audience.

The Business View

The Business View, the first view of the CMS BI Reference Architecture, is also referred to as the BI Analytical Framework. This view is for communicating with a business or end-user community, and is represented in business-related, non-technical terms.

As shown in Business Intelligence Analytical Framework , the BI Analytical Framework consists of an iterative and gradual progression with the following steps:

  1. Collection of data
  2. Understanding relationships in data, which leads to information
  3. Understanding patterns in information, which leads to knowledge
  4. Understanding principles in knowledge, which leads to intelligent actions

The Business View is represented by a multi-layered BI Analytical Framework and includes a range of categories, from information consumers to levels of data sourcing.

Business Intelligence Analytical Framework (page 20)

Users Layer

The Users Layer represents different types of information consumers at various levels of the organization. This architecture supports each type of consumer within the organization as well as those external to the organization, such as CMS business partners, contractors, and U.S. citizens. The key principle of the BI Analytical Framework is its use of role-based authorization, which delivers only the needed information to each type of information consumer. The BI Analytical Framework ensures that the right information is given to the right consumer, at the right time, through the right tools.

Analytics Layer

The second layer of the BI Analytical Framework is the Analytics Layer, comprised of the Analytical Subject Areas and the Analytical Techniques.

The Analytical Subject Areas encompass the analysis and reporting of key business process categories: Claim, Beneficiary, Plan, Provider, and Quality. CMS has built analytics applications for each business process area to help CMS information users understand more about their areas of interest (beneficiaries, claims, etc.).

CMS Analytical Subject Areas represents the current CMS integrated analytics environment and its supported Analytical Subject Areas.

CMS Analytical Subject Areas (page 21) 

Each Analytical Subject Area includes BI tools and technologies to provide historical, current, and predictive views of its business operations. The following analytics functions are performed in their respective analytical subject areas:

  • Claim Analytics
    • Monitoring benefits and claims
    • Identifying payments, overpayments, and recovery
  • Beneficiary Analytics
    • Assessing enrollment and participation
    • Verifying eligibility and entitlement
  • Provider Analytics
    • Identifying fraud, waste, and abuse
    • Monitoring service and performance
  • Plan Analytics
    • Assessing premium and cost sharing
    • Assessing education and outreach
  • Quality Analytics
    • Monitoring quality of service and performance
    • Understanding quality outcomes and efficiency

The analytics circles in CMS Analytical Subject Areas overlap and form a Venn diagram showing how the BI Analytical Framework enables cross-organizational analytics in an integrated analytics environment. In such a system, information consumers from diverse organizations can ask more detailed, involved questions of their enterprise BI Environment.

As the CMS IDR Cloud becomes the primary, authoritative, enterprise-wide data asset through consolidation of CMS data warehouses, data marts, and applications, CMS can broaden its scope of analytics to include new subject areas.

The Analytics Layer also includes the Analytical Techniques, which are essentially the analytics applications. Analytical techniques available for use in the BI environment are:

  • Standard query and ad hoc reporting
  • Online analytical processing implemented as relational online analytical processing (ROLAP), multidimensional online analytical processing (MOLAP), and hybrid online analytical processing (HOLAP)
  • Dashboards and scorecards
  • Data mining
  • Predictive modeling
  • Data visualization
  • Geographical information systems (GIS)

The BI Analytical Framework provides support for CMS BI tools that are integrated with Microsoft Office applications such as Excel, PowerPoint, and Word and with collaboration and social networking tools like SharePoint.

Some BI software tools have strategic alliances with geographic analysis application vendors to provide support for GIS. These GIS capabilities provided in the BI software tools are also supported in the BI Analytical Framework.

The TRA Glossary provides a definition of each analytics application.

Data Layer

The Data Layer comprises data sources and data warehousing layers. The data sources layer represents the operational application systems environment that supports high-volume transaction processing. This environment should not be disturbed for serving any kind of BI user request. The raw data found in the operational systems is also not suitable for direct query and reporting purposes.

The data from the operational systems is selected to represent the source data system of record, meaning the data that is most accurate, complete, up-to-date, trusted, and accessible. The selected raw source data system of record is first stored in the data staging area where it is transformed into standardized CMS formats and stored in the data warehousing layer for analytic processing. The data in the data staging area is not accessible by BI users.

Data transformation is the process of mapping the source data to the destination target environment and transforming the data based upon specific business rules. The business rules represent the correctness or quality of data. When the source data is cleaned, transformed, and cataloged, it is stored in the target data warehouse.

Data warehouses are non-volatile and are periodically updated to reflect changes in business requirements. IDR is an active data warehouse that supports more frequent updates. Data marts are specific subsets of data extracted from a data warehouse for a specific purpose and user group.

Operational Systems

The CMS operational data is found in legacy systems and stored in standard mainframe data storage facilities like the Virtual Storage Access Management (VSAM) and flat files. The operational data categories are Claim, Beneficiary, Plan, Provider, and Quality.

The Data Layer also includes data sources from external agencies, such as the Eligibility, Entitlement, Enrollment, and Death File from the Social Security Administration and census data from the U.S. Bureau of the Census. Each operational system represents an application silo in the CMS operational environment.

CMS is actively integrating its disparate operational data systems into enterprise data repositories (e.g., data warehouses, data marts, and ODSs), which are further integrated into an enterprise-level IDR. CMS uses Informatica PowerCenter to extract, transform, and load data from the various operational systems into integrated CMS data repositories.

The following data repositories are available for use in the CMS BI Environment.

Data Warehouses

  • Integrated Data Repository Cloud

Data Marts

  • Medicare VDM
  • Drug Data Processing System VDM

Operational Data Stores

CMS uses the following ODSs to build Agency data warehouses and data marts:

  • Common Working File (CWF)
  • Health Insurance Portability and Accountability Act (HIPAA) Eligibility Transaction System (HETS) 270/271
  • Medicare Advantage Prescription Drug System (MARx) Inductive Use Interface (IUI)
  • Health Integrated General Ledger Accounting System (HIGLAS)

Collectively, these data warehouses and data marts become the data presentation area where CMS data is organized, stored, and made available for direct querying by users, report writers, and other analytic applications. Technical details of the CMS data warehouses, ODSs, and data marts are provided above in Data Layer.

Security, Data Privacy, and Data Use Agreement

Important components of the CMS BI Analytical Framework from the Business View are security and data privacy. As CMS integrates data from disparate internal and external data sources for access and sharing by common user communities, it must adhere to current Moderate Level CMSRs published in the CMS ARS, the CMS Policy for the Information Security Program (PISP), and CMS privacy and DUA policies. For additional information on privacy and DUAs, please refer to the CMS website on Privacy.

Security, privacy, and DUA are discussed further below in Cross-Infrastructure Layer.

The Technical View

The second view of the BI Reference Architecture is the Technical View, as shown in CMS BI Reference Architecture . This view provides more technical detail and is intended for a technically oriented businessperson or someone implementing, maintaining, or operating CMS IT systems. The guiding principle of the Technical View is that components of the BI framework are implementable segments categorized as technology, hardware, or software.

CMS BI Reference Architecture (page 22) 

The Technical View is represented by the same fundamental concepts as the Business View and comprised of the same three layers—Users, Analytics, and Data—found in the BI Analytical Framework. In addition, the new concept of a Cross-Infrastructure Layer is included in the Technical View, representing the enterprise IT infrastructure management, services, technology, and components. This layer consists of the following categories: Security, Privacy, and DUA; System and Data Management and Administration; Network Connectivity, Protocols, and Access Middleware; and Hardware and Software Platforms.

Users Layer

The Users Layer in the Technical View represents the BI user interfaces. A user interface is the system by which BI users interact with the BI Environment and includes interfaces for Web browsers, portals, devices, and web services.

Web Browsers and Portals

The new CMS web browser-based BI Portal is the primary interface for all business users of the CMS BI Environment and is much more than a layer of screen and report presentation programs. The BI Portal includes services for accessing and retrieving data from the BI / Semantic Layer (described below in detail in Metadata) and from CMS data repositories—all triggered by user requests. Once successfully logged into the BI Portal, the user can enter a new business question, select a query or data mining function to answer a business question, or view BI training and other supporting documents.

The CMS Enterprise Portal is the common user presentation layer that provides a centralized, browser-based, secure point of entry for BI users to access BI data. The BI Portal logically consolidates information and business functions, enabling consistent delivery and presentation of information across the user base. Specifically, BI users can:

  • Collaborate and share queries and reports
  • Use browser-based reporting applications
  • Manipulate data and information
  • Save data and information in the BI Portal layer

Portal users can perform these specific BI analytic functions in various BI applications without having to exit the portal.

The CMS Enterprise Portal is also discussed in the previous topic, CMS Enterprise Portal Framework.

Devices

Devices are technology components used as vehicles to receive information and interact with the digital environment. Uses and features of devices vary widely. Examples of devices include personal computers, personal digital assistants (PDA), and mobile and Smart phones. Although CMS approves the use of these devices, their functionality and interaction with the CMS Enterprise Portal system have not yet been approved by the TRB.

Web Services

Web services are technologies and processes for facilitating communications between two applications. For example, custom-coded applications can communicate with a BI tool using web services. In the CMS BI environment, web services enable the use of selected (published) analytical information by other BI applications.

Analytics Layer

The Analytics Layer provides access to analytics applications central to the BI environment. A variety of applications may be supported, from static reporting and balanced scorecards to sophisticated quantitative models embedded in an operational process. This layer typically consists of various technological components designed to meet specific needs.

The CMS BI Reference Architecture supports the technology components required in the business analytics applications described in detail above in The Business View.

Data Layer

The Data Layer represents all the data sources, data integration processes, and data repositories that support the BI Environment. CMS Technical Data Architecture presents the CMS Technical Data Architecture, showing the flow of data in the Data Layer and the integration processes necessary to move data from the data sources into usable formats in the BI Environment.

CMS Technical Data Architecture (page 23)

 

This topic examines the components and services of the CMS technical data architecture, as well as the primary data sources available in the CMS operational systems. The operational systems are found in legacy systems and stored in standard mainframe data storage facilities like VSAM and flat files.

Operational Data Stores provide other sources of data available for building CMS data warehouses. CMS has implemented numerous ODSs, including CWF, MARx, and HIGLAS, that contain detailed, transaction-level data. ODSs are primarily used as data feeds to build data warehouses and not as data sources in the BI Environment for user query and reporting.

Data is extracted from the operational systems based upon specific BI application business requirements. The extracted data is placed in the Data Staging Area, where much of the data transformation takes place and the added value of a data warehouse is realized. The raw data is stored in the Data Staging Area in simplified and accessible forms—e.g., flat files, relational tables, or proprietary structures used by the Informatica PowerCenter Extract, Transform, and Load (ETL) tool. The data found in the Data Staging Area is not available as input to analytic processing in the BI environment.

Once the data in the Data Staging Area are transformed, combined, and cleaned using specific business rules, it is simple for load utilities to load the data into relational databases in the Data Repositories. Load utilities are provided by the relational database management system (RDBMS) software or by the Informatica PowerCenter.

Because the ETL process is iterative, the source for a specific load process may also be the target data warehouse itself, which becomes the data feed for building dependent data marts.

The following subtopics describe each component of the Data Layer in detail.

Data Sources

The Data Sources Layer identifies all sources of data available within and outside of CMS that are accessed and used as part of the BI Environment. This data may include either structured or unstructured data.

Operational

The appropriate Medicare and Medicaid data is obtained from the operational systems: Claim, Beneficiary, Provider, Plan, or Quality. Operational systems data is extracted and transformed into standardized CMS formats based upon specific business rules, converted into database records, and stored in relational database management systems.

Unstructured

Unstructured data are captured as text in email or documents, as audio or video files, or as images. In some cases, unstructured data may be stored in the RDBMSs as binary or character large objects (BLOB or CLOB). CMS BI applications may simply retrieve unstructured data as BLOBs or CLOBs for use in the BI Environment.

Unstructured data may also be indexed and stored in the CMS Enterprise Content Management (ECM) system for use by the BI application.

Informational

Informational data sources are the output of analytic processes. Output data are stored as informational data in the RDBMSs or multidimensional databases and reused for further analysis.

External

External data sources are those available external to CMS’s data assets, such as Social Security Administration and Census Bureau data. External data is transformed and loaded to enhance the CMS data warehouses and data marts.

Data Integration

The Data Integration Layer contains all technology components and processes that support the processing and movement of data to prepare it for storage in the Data Repositories Layer or to share it with other analytical applications and systems. This layer may process data in scheduled batch intervals or in near real-time / “just-in-time” intervals, depending on the nature of the data and its business purpose.

Extract, Transform, Load / Apply

Extract, Transform, Load / Apply refers to the technologies and processes by which the Data Sources Layer is accessed, extracted, transformed, and loaded into storage in the Data Repositories Layer. This involves extracting the operational data, transforming it based on specific business rules and common data warehousing best practices, and loading it into a target data store or stores.

This is generally referred to as the ETL process. In the context of BI architecture, the key point is that properly designed and executed ETL allows for both the organization of data according to subject matter and the enforcement of business rules as the data is loaded. Information about data movement and the associated business rules is then stored in the metadata repository. ETL software is used multiple times in the life of data movement. Initially, ETL can be used to move (or extract) data from its original source into an ODS, data warehouse, or data mart. During each of these moves, transformations and/or business rules may be applied to clean, scrub, remove duplicates, and/or standardize the data.

In addition to loading the scrubbed data from the ODS databases into the data warehouses, additional ETL processes can aggregate data into a data mart to provide a summarized view of the information. When the loading process is complete, the data is ready for business users to access through the BI Portal.

ETL processing occurs on regular schedules (i.e., daily, weekly, or monthly) to meet business reporting or analysis requirements and consistently refreshes the data loaded into the data warehouse.

Integrity / Quality

Integrity / Quality represents the technology stage in which the operational data that has completed the ETL process is further assessed for quality, reliability, completeness, timeliness, accuracy, and missing values. Achieving and sustaining a high level of data quality requires an effective enterprise data governance program as well as diligent analysis, planning, implementation, and monitoring.

Agency-wide data quality standards for source data are applied in the data cleansing, consistency checking, completion, and profiling activities conducted by the ETL software as part of the transformation process. Data profiling uses specific tools to “measure” the data quality level expected from a specific source system, and the process may be quite complex. In this stage, additional edits or business rules may also be applied to data to meet user requirements.

CMS uses validation tools to measure and monitor key data quality attributes across all data types and sources. This information is used to create and support a business culture that values data quality across the CMS enterprise.

Synchronization

Synchronization is the process of sharing data by copying it across physical storage repositories while still retaining its authoritative validity. Through this process, CMS provides a database that could be used by mid-tier processes, Java-based applications, and Procedural Language / Structured Query Language applications without having to cross the network to access the DB2 data repositories in the production environment.

Data Repositories

The Data Repositories Layer contains the databases and components providing most of the storage for the data that supports a BI Environment. Data repositories are not a replacement or replica of operational databases that reside on the Data Sources Layer, but rather, a complementary set of databases that reshape data into formats necessary for responding to ad hoc queries and helping to make business management decisions.

Data Staging Area

The Data Staging Area is where raw data from the Operational Systems is loaded, cleaned, combined, and exported to one or more data warehouses / marts. The raw data is copied into simple, accessible formats (e.g., relational databases and flat files) for transformation. The transformation programs from ETL software reformat the source data records, rectify any discrepancies, and delete duplicate records. The resolution process for data record issues is predefined in the ETL software according to CMS-created business rules.

Operational Data Store

The Operational Data Store is a hybrid environment in which operational data is transformed by ETL into an integrated format. Once placed in the ODS, the integrated data is available for online updates.

The ODS has a dual purpose—it serves as a point of integration for operational systems, and supplies current, detailed data for data warehouses in the CMS data repositories

The ODS is not available for direct query and reporting by users in the BI Environment.

Data Warehouse

A data warehouse (DW) consists of various databases in which the previously cleansed ODS data is stored in standardized formats for quick access by multiple BI tools and approved users. DWs have different attributes than transactional databases and are designed to optimize response time for user queries and reports, normalize source data, and eliminate data redundancy. Data may be sorted by specific query tables or dimensions, such as geography, time, or claim type, and is sharable by BI tools and users. Data may also be aggregated into a table such as total claims by geographic region for the week, month, or year.

This layer includes several large, user-defined databases as DWs, the largest of which is the IDR. Over time, as existing data from various sources is integrated into the IDR and more users rely on it as the primary, authoritative source of enterprise data, disparate query results from users across CMS will be reduced.

Data Mart

Because the source data extracted, transformed, and loaded into the BI Environment is used in different ways by a variety of users, data marts are used to reorganize source data to meet different user purposes. Data marts are, in effect, smaller versions of the predefined database tables developed in the DWs. Data marts are basically division- or branch-sized DWs or databases, making data available to specific user groups. For instance, one data mart may contain all the data needed to quickly answer Part D queries, while another may be focused on building a data structure to view beneficiary information by month and year. Views or VDMs provide a “window” into the data and can be used to filter out data elements that the user does not want or need to see. These VDMs are pre-defined in each data mart. The data marts and OLAP constitute all the predefined, tailored databases specifically designed to optimize response time for meeting various demands of user groups.

Metadata

Metadata is information common to all layers in the BI Environment that is crucial for ensuring the integrity of data as it moves from being raw data to structured formats accessible by end users. Metadata, or “data about data,” provides a consistent description of discrete data by defining common data names, definitions, and integrity rules across BI tools, ETL tools, and databases. Metadata supports all aspects of BI delivery capability, including:

  • Capturing business conversion and transformation rules
  • Facilitating the user’s ability to understand the “business” meaning of data elements and relationships for creating BI queries and reports
  • Assisting an auditor’s ability to trace data lineage from sources to reports

Metadata also captures metrics for data usage and retrieval, providing insight into BI performance.

Metadata management is a key aspect of the BI Environment and relies on the agency’s strategy that defines and maintains the enterprise-wide metadata source. Each tool and database within the BI Environment houses metadata relevant to its function and purpose.

Metadata captured in the BI Environment is generally known as the BI Semantic Layer. This layer isolates business users from the technical complexities of databases by using everyday terms to describe the business environment. By using the BI Semantic Layer to create a query, users can retrieve exactly the data that interests them while communicating in their familiar business terminology. The BI Semantic Layer empowers users with a variety of tools and applications by providing deeper insight into enterprise data assets while hiding the complexity of data and analytics.

Some BI tools, such as MicroStrategy, call this semantic layer its Metadata, while others, such as Business Objects, call it the Universe.

Cross-Infrastructure Layer

The BI Environment requires interactions across its multiple layers that are relevant to all layers. Many benefits of the dynamic and actionable nature of the BI Environment are a result of these additional component interactions.

The Cross-Infrastructure Layer includes security, data privacy, DUA, and infrastructure components, as described in the following subtopics.

Security, Data Privacy, and Data Usage Agreement

Security is of paramount importance in a BI Environment and is applicable to every component at every level of the architecture. Security not only addresses access to applications and data, but also enables business-rule and role-based views of data.

Safeguarding privacy is closely related to security and is of utmost importance in a BI Environment. Privacy processes and technology focus on defining and managing highly sensitive PII and PHI data.

Security and data privacy adhere to the technologies, processes, and organizational components that meet with CMS and federal government security and privacy regulations, policies, and standards. The CMS PISP provides specific security policies; the current Moderate Level CMSRs published in the CMS ARS provide specific security requirements; and the CMS TRA Network Services, Access Control and Identity Management chapter provides specific engineering guidance for implementing system access controls that partially address the CMSR.

A Data Use Agreement is a legally binding agreement between CMS and an external entity (e.g., contractor, private industry, academic institution, or other federal government or state agency), formed when an external entity requests the use of CMS personally identifiable data that is covered by the Privacy Act of 1974.

CMS BI tools use the CMS Enterprise LDAP Directory for user authentication. UserIDs and passwords are managed through CMS Enterprise Identity Management services.

Role-based authorization is used to manage access to BI applications, reports, analytic functions, tables, views, and procedures based upon BI user classifications, discussed in detail below in BI Implementation, BI User Management.

Infrastructure

Infrastructure provides the technologies, processes, and services that enable the BI Environment to exist and operate. Infrastructure primarily includes Systems and Data Management and Administration; Database Management and Administration; Security Administration and Control; Network Connectivity, Protocols, and Access Middleware; and Hardware and Software Platforms.

Functions, capabilities, services, hardware, and software provided in each infrastructure component are described in the following subtopics.

System Management and Administration

System Management and Administration in the CMS BI Environment provide the following support and services:

  • Support BI server administration via role-based access control
  • Support authentication and authorization within the BI Tool and BI Portals
  • Establish DUAs for all contractor staff
  • Create database schemas to support the BI analytics applications
  • Create user account(s) to access the various reporting databases via the BI tool
  • Create ETL business rules, workflows, and source-to-target mappings to move data from source systems to IDR or other platforms

Hardware and Software Platforms

The CMS BI Environment employs a variety of hardware and software platforms, including the following:

  • Web servers, including failover and load-balancing
  • Application servers, including failover and load-balancing
  • Database servers to support storage of application data and metadata as well as any reporting databases
  • Server operating systems
  • Software application servers
  • Database middleware clients
  • Workload / scripting automation software
  • Web server software
  • BI tool software and associated Software Development Kit (SDK)
  • BI tool web services

Alignment with the CMS TRA Multi-Zone Architecture

The CMS BI Reference Architecture aligns with the CMS Multi-Zone Architecture, as illustrated in Alignment of CMS BI Reference Architecture with CMS TRA Multi-Zone Architecture .

Alignment of CMS BI Reference Architecture with CMS TRA Multi-Zone Architecture (page 24)

BI Component Placement in the CMS TRA Multi-Zone Architecture

CMS Business Intelligence Component Placement depicts the common placement for BI tool components in the CMS TRA Multi-Zone Architecture.

CMS Business Intelligence Component Placement (page 25)

 

Load Balancers

Load balancers are the initial point of contact for new user sessions, ensuring that user session load is dynamically and evenly distributed across all Web servers. Dynamic distribution of load allows scalability, resiliency, and ease of configuration of Web servers. User session load balancers must be placed in the Presentation Zone.

Web Servers

Web servers manage communication between the Web clients and the BI application servers. Web servers will be redundantly configured and load balanced. All Web servers must be placed in the Presentation Zone.

Web servers must be implemented on a CMS-approved Web server platform. Web browsers must communicate with the Web server through HTTPS.

Business Intelligence Servers

Business Intelligence servers must be implemented in the Application Zone. These servers provide the core analytical processing and job management for all reporting, analysis, and monitoring applications. Exceptions should be authorized by the TRB.

BI Servers communicate with the Web servers in the Presentation Zone through Transmission Control Protocol / Internet Protocol (TCP/IP) and with the Database Servers in the Data Zone through standard database access methods — ODBC and JDBC, which are non-compliant with the CMS TRA Multi-Zone Architecture, as stated above in BI Environment, Non-Compliance with CMS TRA Multi-Zone Architecture.

The BI application servers currently implemented in CMS are:

  • MicroStrategy Intelligence Server
  • Business Objects Enterprise Server
  • Cognos Business Intelligence Server
  • SAS Enterprise BI Server

 

Database Servers

Database Servers support the DWs, data marts, and ODSs implemented in the data zone.

Metadata Repositories

The BI tool metadata repository is a set of database tables that store BI information including warehouse schema, server definition, projects, queries, reports, users, and warehouse connectivity information. A server definition is a specification of connectivity information such as the metadata data source name (DSN) and the metadata ID and password for a configured instance of a BI server.

 

Business Intelligence Business Rules 

The business rules provided in this topic serve as the CMS standards and conventions for implementing CMS BI solutions. CMS established the following guidance and standards for developing the Agency’s BI implementations based upon its chosen BI software suites.

BR-BI-1: All Production BI Applications and Systems Must Comply with the CMS TRA and the BI Reference Architecture

BR-BI-2: (Withdrawn after TRA 2016R1): BI applications must run on the Sun Solaris and/or Z-Linux platform, and the server components may reside on one or multiple platforms

BR-BI-3: Authentication, Auditing, and Logging of All BI User Accounts Must Be Managed from the CMS Enterprise LDAP Directory and Enterprise User Administration

BR-BI-4: All Traffic Must Be Encrypted between a BI User’s Browser, Web Services, or Device and the BI Server

BR-BI-5: Role-Based Authorization Must Be Used to Manage Access to BI Applications, Queries, Reports, Analytic Functions, Tables, Views, and Stored Procedures

BR-BI-6: All BI Applications and Systems Must Comply with the Current CMS ARS

BR-BI-7: BI Applications Must Leverage the BI Portal Framework as an Enterprise-Wide Secure Gateway that Enables All BI Tools to Access, Manipulate, and Analyze Data

Business Intelligence Implementation Considerations

This topic provides key considerations in implementing new CMS BI applications. These considerations represent guidelines and best practices in planning, implementing, and deploying a BI environment that maximizes business benefits.

Integrate Operational Data Sources

The ETL process is essential in integrating CMS operational data sources into the BI Environment for analytics. ETL extracts and transforms data based on business rules and data warehousing best practices and loads that data into the target data repository(ies).

ETL programs will be created by one of two methods: creating custom ETL programs with a programming language or using a COTS product such as for example, Informatica PowerCenter. This decision will be based upon the project scope, data volume, size, and cost. It is noteworthy that Informatica provisions the use of custom code embedded in its PowerCenter workflow.

Data Integrity / Quality

Data in the CMS data repositories for use in the BI Environment must be clean, consistent, and complete. Data integrity / quality checks will ensure the data’s cleanliness, consistency, and completeness before it is published to the user community. This process involves cleansing the data by correcting misspellings, resolving domain conflicts, addressing missing elements, or parsing into standard formats.

An integral part of the data quality process is data profiling, which uses specific tools to monitor and measure the level of data quality expected from a specific data source.

BI User Management

This topic describes the user management measures CMS will consider in implementing any new BI application.

Owners and developers of new BI application must define user requirements that include the following functions:

  1. Classifying BI application users for role-based access, including some or all the following user classes:
    1. Standard User – An individual with access to predefined reports or data structures within the authorized BI application (e.g., executives, managers, CMS business partners, and U.S. citizens).
    2. Power User – An individual with standard user access plus the ability to generate ad hoc reports using data within an authorized BI application (e.g., savvy business analysts and BI analysts).
    3. BI Developer – An individual responsible for developing and implementing BI applications (e.g., CMS BI Contractor).
    4. Business Intelligence Analyst – An individual responsible for providing information based on unique and changing needs within each BI application in the BI environment e.g., EADG’s BI consultants and developers, business analysts within the Division of Quality Coordination and Data Distribution).
  2. Defining user access roles and implementing a process for authorizing users
  3. Identifying BI contents such as metadata, queries, and reports that specific user groups may access
  4. Identifying the data that user groups may access from the CMS data repositories

After the BI application security requirements are defined, the following security measures will be applied in the BI Environment.

Authentication

Authentication will be performed by the methods described in the Network Services, Access Control and Identity Management topic. Identity and authentication requirements may vary depending on the sensitivity of the data involved.

Authorization

Access control applies to BI contents in BI applications and to database objects in the CMS data repositories.

The authorization of access to database objects—tables, views, stored procedures, columns, and even rows—will be controlled by the BI application and the RDBMS.

Administration

The Office of Information Technology-designated administrator(s) will manage and control BI application security with an application’s available administrative tools

For BI tools that provide centralized management of user administration, application, and core infrastructure, the administration tool will be placed on a hardened and secure workstation. The OIT-designated administrator will align users, groups, and roles in EUA / LDAP with the BI objects. BI tools that provide web-based administration, only authorized users with valid credentials can log into the central administration console, which may only be accessed from internal CMS networks.

Sizing and Configuration

New BI applications must be properly sized and configured in the BI production environment. EADG will provide guidance in sizing and configuration. This organization will work with new BI applications to determine the optimum hardware configuration based on the BI application requirements. The BI tool vendor will also provide information on best practices when implementing its COTS product.

The following are some (but not all) key factors used to determine sizing and configuration: number of users, complexity of reports, number of ad hoc reports versus cached reports, and the need for data mining or internet access. Analyzing these factors is a starting point in determining the optimal configuration for the business user BI Environment.

Number of Users

The user size of the business intelligence community is a significant factor in determining hardware requirements. The number of users can be categorized as follows:

  • Total users – The number of all user accounts that will be created in the BI Environment.
  • Active users – The number of users who will be logged into the BI system.
  • Concurrent users – The number of users who are likely to have queries and reports processing simultaneously on the BI system.

Of these categories, it is most important to determine the number of concurrent users. The BI system must be able to support the maximum number of concurrent users expected at any given time.

Report Complexity

BI system resource use and processing time depends on the complexity of reports to be compiled. Some factors contributing to report complexity are:

  • Number of result rows required
  • Complexity of metric calculations needed
  • Degree of analytical processing performed

Ad Hoc Reports versus Cached Reports

Standard reports can be scheduled to execute at non-busy times. When scheduled reports execute, BI tools create a report cache. When these standard reports are subsequently executed, the system accesses a cached report and does not execute a report against the DW or data mart.

Because ad hoc queries are executed on the spur of the moment, they cannot be cached in advance. Therefore, ad hoc reports are executed against the DW / data mart. If users execute a high percentage of ad hoc reports, additional system resources are required.

Data Mining

Data mining is an analysis that discovers patterns in sample data sets using specific algorithms such as decision trees, neural networks, and clustering. When algorithms find patterns in small sample sets, data mining further validates the hypothesis with larger data sets in very large DWs. This process may take several hours.

Data mining can consume the greatest amount of BI system resources. Although the best practice is to schedule data mining procedures in non-busy hours, additional system resources must be considered in sizing and configuring BI systems for data mining.

Internet and Extranet Accesses

The workload of BI users in the CMS extranet for internal users, CMS business partners, and contractors is more manageable and predictable than the workload of public users from the Internet. When BI queries and reports are made public, the workload of the public—the number of concurrent users, the complexity of reports, and the number of standard reports—will be difficult to predict. Constant monitoring of the Internet workload is thus required to balance the Internet and extranet workloads.

User Activity Control and Monitor

To proactively administer and monitor usage and system performance, BI administrators must establish user activity controls and monitor activity based on the workload definition rules. BI administrators must establish thresholds such as query processing time as well as output row limitations to monitor for exception conditions.

Future Considerations

Metadata Integration

Metadata management is a key aspect of the BI Environment that relies on the Agency’s strategy to define and maintain the enterprise-wide metadata source. In the current CMS BI Environment, each BI tool and RDBMS creates and manages its own metadata relevant to its function and purpose. CMS will need an overarching metadata strategy to enable more sophisticated enterprise management of data throughout its life cycle, improve data quality, facilitate cross-organizational ease of data sharing, and increase the value of information delivery to business users.

Common Security Controls

A BI portal provides a central facility for adding new BI applications. An additional topic detailing the controls that can be inherited from a BI portal by new applications will be reviewed and added to provide greater coverage of security controls at lower overall cost.

 

REFERENCE INFORMATION

TRA Acronyms

The TRA Acronyms contains a list of acronyms referenced in each section of the CMS TRA.

TermDefinition
3PAOThird-Party Assessment Organization
AAAAuthentication, Authorization, and Accounting
AIPv4 Address Record
a.k.a.Also Known As
AAAddress Allocation
AAAAIPv6 Address Record
AALAuthenticator Assurance Level
ACAccess Control
ACAAffordable Care Act
ACEAccess Control Entitlements
ACFAccess Control Facility
ACLAccess Control List
ACOAccountable Care Organization
ACPAccess Control Product (e.g., ACF/2, RACF, TSS)
ACRArchitecture Change Request
ACTAdaptive Capability Testing
ADActive Directory
ADApplication Development
ADMApplication Development Methodology
ADOApplication Development Organization
AESAdvanced Encryption Standard
AHRQAgency for Healthcare Research and Quality
AJAXAsynchronous JavaScript and XML
ALFApplication Layer Filtering
ALFAApplication Layer Filtering Authorization and Authentication
ALFGApplication Layer Filtering Gateway
ALOMAdvanced Lights Out Manager
AMDAdvanced Micro Device
AMQPAdvanced Message Queuing Protocol
AOAdministrative Officer
AOAuthorizing Official
APFAuthorized Program Facility
APIApplication Programming Interface
APMApplication Performance Monitoring
AQAcquisition
ARIAAccessible Rich Internet Applications
ARMApplication Response Measurement
ASAutonomous System
AS3Simple Storage Service (Amazon)
ASNAutonomous System Number
ASPApplication Service Provider
ASPAAssistant Secretary for Public Affairs
ATAGAuthoring Tool Accessibility Guidelines
ATOAuthorization to Operate
AUAudit and Accountability
AVAnti-Virus
AWSAmazon Web Services
AZApplication Zone
AZAvailability Zone
BAMBusiness Activity Monitoring
batCAVEContinuous Authorization and Verification Engine
BDCBaltimore Data Center
BGPBorder Gateway Protocol
BIBusiness Intelligence
BIABusiness Impact Analysis
BINDBerkeley Internet Name Daemon
BLOBBinary Large Objects
BMPBitmap
BPELBusiness Process Execution Language
BRBusiness Rule
BRMBusiness Reference Model
BSDBerkeley Software Distribution License
BSRBootstrap Router
CACertificate Authority
CAACMS Access Administrator
CAREContinuity Assessment and Record Evaluation
CBOCommunity-Based Organization
CBPCMSNet Business Partner
CBWFQClass-Based Weighted Fair Queuing
CCBChange Control Board
CCBConfiguration Control Board
CCECommon Configuration Enumeration
CCICCMS Cybersecurity Integration Center
CCMCloud Controls Matrix
CCOCall Center Operations
CCSSCommon Configuration Scoring System
CCWChronic Care Warehouse
CDACentral Database Administration
CDMContinuous Diagnostics and Mitigation
CDNContent Delivery Network
CDNContent Distribution Network
CECustomer Edge
CEAChief Enterprise Architect
CEICommon Enterprise Infrastructure
CERTCarnegie Mellon University Computer Emergency Response Team
CFACTSCMS FISMA Controls Tracking System
CFRCode of Federal Regulations
CHDCContractor-Hosted Data Center
CHPIDChannel Path Identifiers
CICloud Infrastructure
CIConfiguration Item
CIContinuous Integration
CIAConfidentiality, Integrity, and Availability
CICSCustomer Information Control System
CIEMCanonical Modeling for Information Exchange Methodology
CIFSCommon Internet File System
CIMCommon Information Model
CIOChief Information Officer
CISCenter for Internet Security
CISOChief Information Security Officer
CLFCommon Log Format
CLOBCharacter Large Objects
CMConfiguration Management
CMCMS Cloud Manager
CMAComputer Matching Agreement
CMaaSContinuous Monitoring as a Service
CMEContinuing Medical Education
CMISContractor Management Information System
CMSCenters for Medicare & Medicaid Services
CMSNetCMS Private Network
CMSRCMS Minimum Security Requirements
CNAMECanonical Name Record
COCentral Office of CMS
COBCoordination of Benefits
COBSCoordination of Benefits Service
COICommunity of Interest
COOPContinuity of Operations Plan
CORSCross-Origin Resource Sharing
CoSClass of Service
CPUCentral Processing Unit
CRChange Request
CRACyber Risk Advisor
CRCCyclic Redundancy Code
CROWNWebConsolidated Renal Operations in a Web-based Network
CSACloud Security Alliance
CSIRCComputer Security Incident Response Center
CSMConfiguration Settings Management
CSPCredential Service Provider
CSPCloud Service Provider
CSRCustomer Service Representative
CSSCascading Style Sheets
CSVComma Separated Variable
CTOChief Technology Officer
CVECommon Vulnerabilities and Exposures
CVSSCommon Vulnerability Scoring System
CWE™Common Weakness Enumeration
CWFCommon Working File
CYCalendar Year
DDelivery
DAData Architecture
DASDDirect Access Storage Device
DBADatabase Administrator
DBidSDMEPOS Bidding System
D-BidsDurable Medical Equipment Billing System
DBMData and Database Management
DBMSDatabase Management System
DCData Center
DCEPData Converter Evaluation Platform
DDESDivision of Data Enterprise Services
DDLData Definition Language
DDoSDirect Denial-of-Service
DDPSDrug Data Processing System
DEADivision of Enterprise Architecture
DESYData Extract Software System
DEVDevelopment
DFMDesign for Maintainability
DFSDigital Forensics Services
DHCPDynamic Host Configuration Protocol
DHSDepartment of Homeland Security
DIIMPDivision of IT Investment Management and Policy
DIMEDirect Internet Message Encapsulation
DISADefense Information Systems Agency
DITDefect and Issue Tracking
DITGDivision of Information Technology Governance
DLMData Life-cycle Management
DMData Mart
DMEDurable Medical Equipment
DMEPOSDurable Medical Equipment Prosthetic, Orthotic, and Supplies
DMLData Modification Language
DMVPNDynamic Multipoint Virtual Private Network
DMZDemilitarized Zone
DNSDomain Name System / Domain Name Service
DNSSECDomain Name System Security Extensions
DoDDepartment of Defense
DOMDocument Object Model
DoSDenial of Service
DPDevice Profiler
dpiDots per Inch
DPLDynamic Program Link
DRDisaster Recovery
DSCPDifferentiated Services Code Point
DSDLDocument Schema Definition Language
DSNData Source Name
DSSData Storage Services
DUAData Use Agreement
DWData Warehouse
DZData Zone
E01Expert Witness
EAEnterprise Architecture
EaaSEnterprise as a Service
EADGEnterprise Architecture and Data Group
eBGPExternal Border Gateway Protocol
EBPExtranet Business Partner
EC2Amazon’s Elastic Compute Cloud
eCHIMPElectronic Change Management Portal
ECMEnterprise Content Management
ECMAEuropean Computer Manufacturers Association
EDEngineering Documentation
EDCEnterprise Data Center
EDEEnterprise Data Environment
EDLEnterprise Data Lake
EDMEnterprise Data Mesh
EDREnterprise Data Repository
EDSREnhanced Dedicated SONET Ring
EDWEnterprise Data Warehouse
EEEnterprise Edition (Java)
EFExpedited Forwarding
EFIEUA Front End Interface
EFTEnterprise File Transfer
EFTEnterprise File Transfer, Electronic File Transfer
EHRElectronic Health Record
EHRDElectronic Health Records Demonstration
EIDEnterprise Identifier
EIDMEnterprise Identity Management
EIGRPEnhanced Interior Gateway Routing Protocol
EINEmployer Identification Number
EITElectronic and Information Technology
EJBEnterprise Java Bean
ELAEnterprise License Agreement
ELDMEnterprise Logical Data Model
EMPIEnterprise Master Person Indexes
EOCEnterprise Operations Centers
EPPEndpoint Protocol
EPSEncapsulated PostScript
ERREnvironmental Readiness Review
ESEnterprise Security
ESBEnterprise Service Bus
ESIMEnterprise Services for Identification Management
ESOCEnterprise Security Operations Center
ESQLEmbedded Structured Query Language
ESSEnterprise Shared Services
ESSGEnterprise Shared Services Group
ETLExtract, Transform, and Load
ETLExtract, Transform, Load
EUAEnterprise User Administration
EUDCEnterprise User Data Catalog
FALFederal Assurance Level
FAQFrequently Asked Questions
FARFederal Acquisition Regulation
FBIFederal Bureau of Investigation
FCFibre Channel
FCIPFibre Channel Over IP
FCoEFibre Channel Over Ethernet
FDCCFederal Desktop Core Configuration
FDCCIFederal Data Center Consolidation Initiative
FedRAMPFederal Risk and Authorization Management Program
FFRDCFederally Funded Research and Development Center
FICONFibre Connection
FIDFraud Investigation Database
FIPSFederal Information Processing Standards
FISMAFederal Information Security Modernization Act
FMATForensic & Malware Analysis Team
FOIAFreedom of Information Act
FOUOFor Official Use Only
FQDNFully Qualified Domain Name
FSSSFederal IT Shared Services
FTFault Tolerance
FTIFederal Tax Information
FTPFile Transfer Protocol
FTP/SFile Transfer Protocol with SSL for Security
FWFirewall
FWAFirewall Administration
G2BGovernment-to-Business
G2CGovernment-to-Citizens
G2GGovernment-to-Government
GAOGovernment Accountability Office
GBGigabyte
GDGroup Director
GDOIGroup Domain of Interpretation
GETVPNGroup Encrypted Transport Virtual Private Network
GFEGovernment Furnished Equipment
GFIGovernment-Furnished Information
GFSGovernment-Furnished Software
GIDGroup Identifier
GIFGraphics Interchange Format
GISGentran Integration Suite (now called IBM Sterling Integration Suite)
GISGeographical Information Systems
GMTGreenwich Mean Time
GNOSCGovernment Network Operations and Security Center
GNUGNUs Not UNIX
GOTSGovernment Off-the-Shelf
GPLGNU General Public License
GPOGroup Policy Object
GPUGraphical Processing Unit
GRCGovernance, Risk and Compliance
GREGeneric Routing Encapsulation
GSAGeneral Services Administration
GSSGeneral Support System
GTLGovernment Task Lead
GUIGraphical User Interface
GWTGoogle Web Toolkit
HAHighly Available
HBSSHost-Based Security Systems
HCAHPSHospital Consumer Assessment of Healthcare Providers and Systems
HEARHHS Enterprise Architecture Repository
HETSHIPAA Eligibility Transaction System
HETS UIHIPAA Eligibility Transaction System User Interface
HHSDepartment of Health and Human Services
HIDSHost-based Intrusion Detection System
HIGLASHealthcare Integrated General Ledger Accounting System
HIPAAHealth Insurance Portability and Accountability Act of 1996
HIPSHost-based Intrusion Prevention System
HITECHHealth Information Technology for Economic and Clinical Health Act
HOLAPHybrid Online Analytical Processing
HOP QDRPHospital Outpatient Quality Data Reporting Program
HPHewlett-Packard
HPMSHealth Plan Management System
HQAHospital Quality Alliance
HRHuman Resources
HSPDHomeland Security Presidential Directive
HSRPHot Standby Routing Protocol
HSTSHTTP Strict Transport Security
HTMLHyperText Markup Language
HTTPHypertext Transfer Protocol
HTTPSSecure Hypertext Transfer Protocol
HTTPSHypertext Transfer Protocol over Secure Sockets Layer
HVAHigh Value Assets
HWHardware
HWAMHardware Asset Management
I/OInput/Output
IAInformation Assurance; Identification and Authentication
IAAInter-Agency Agreement
IaaSInfrastructure as a Service
IACSIndividuals Authorized Access to the CMS Computer Services
IALIdentity Assurance Level
IAMIdentity and Access Management
IANAInternet Assigned Numbers Authority
IATOInterim Authority To Operate
ICDInterface Control Document
ICMPv6Internet Control Message Protocol for IPv6
ICSAInternational Computer Security Association
ICTInformation Communication Technology
IDIdentity
IDIdentifier
IDIdentifier, Identity
ID/IQIndefinite Delivery/Indefinite Quantity
IDEIntegrated Development Environment
IDMIdentity Management System
IDPIntrusion Detection and Prevention
IDQInformatica Data Quality
IDRIntegrated Data Repository
IDSIntrusion Detection System
IEAInformation Exchange Agreement
IEEEInstitute of Electrical and Electronics Engineers
IEMIBM Endpoint Manager
IETFInternet Engineering Task Force
IGPInterior Gateway Protocol
IHSIBM HTTP Server
IIOPInternet Inter-ORB Protocol
IISInternet Information Server
IKEInternet Key Exchange
ILCIntegrated IT Investment & System Life Cycle
IMIdentity Manager (Sun Microsystems product)
IMAPInternet Message Access Protocol
IMTIncident Management Team
IOCIndicators of Compromise
IOSImmediate Office of the Secretary
IoTInternet of Things
IPInternet Protocol
IPAIntegration Partner Agreement
IPMInfrastructure Performance Monitoring
IPMPInternet Protocol Network Multipathing
IPSIntrusion Prevention System
IPSecInternet Protocol Security
IPv4Internet Protocol version 4
IPv6Internet Protocol version 6
IRIncident Report, Incident Response
IRIncident Report
IRFInpatient Rehabilitation Facility
IRRImplementation Readiness Review
IRTIncident Response Team
ISInformation Security
IS&CTIInformation Sharing and Cyber Threat Intel
IS2PInformation System Security and Privacy
IS2P2CMS Information System Security and Privacy Policy
ISAInteragency Security Agreement
ISATAPIntra-Site Automatic Tunnel Addressing Protocol
ISCIInternet Small Computer System Interface
ISCMInformation Security Continuous Monitoring
iSCSIInternet Small Computer Systems Interface
ISISIBM Sterling Integration Suite (formerly Gentran Integration Suite)
ISOInformation Systems Officer
ISOInternational Organization for Standardization
ISPInternet Service Provider
ISPGInformation Security and Privacy Group
ISPGInformation Security & Privacy Group
ISRAInformation Security Risk Analysis
ISSOInformation Systems Security Officer
ISSOInformation System Security Officer
ITInformation Technology
IT PMPerformance Monitoring/Management
ITCAMIBM Tivoli Application Composite Monitor
ITILInformation Technology Infrastructure Library
IUIInductive Use Interface
IV&VIndependent Verification and Validation
J2EEJava 2 Platform Enterprise Edition
Java EEJava Platform, Enterprise Edition
JCLJob Control Language
JCPJava Community Process
JDBCJava Database Connectivity
JMSJava Message Service
JMXJava Management Extensions
JNDIJava Naming and Directory Interface
JPEG/JPGJoint Photographic Experts Group
JPSJava Portlet Specification
JRAJava Resource Adapter
JREJava Runtime Environment
JSJavaScript
JSFJavaServer Faces
JSONJavaScript Object Notation
JSONPJSON with Padding
JSPJava Server Pages
JSRJava Specification Request
KEKKey Encryption Key
KPIKey Performance Indicator
KSKey Server
KSMKeys and Secrets Management
LANLocal Area Network
LASRLightweight Asset Summary Results
LDAPLightweight Directory Access Protocol
LDAPSSecure LDAP, also known as “LDAP over SSL”
LDOMLogical Domain
LDPLabel Distribution Protocol
LGPLGNU Lesser General Public License
LIRLocal Internet Registry
LOA3Level of Assurance 3
LPARLogical Partition
LRECLLogical Record Length
LTCLong-Term Care
LUNSLogical Unit Numbers
MAC (address)Media access control address
MACMedicare Administrative Contractor
MACMedicare Administrative Contractor, Media Access Control
MAC PPOMedicare Administrative Contractor Preferred Provider Organization
MAPIMessaging Application Programming Interface
MARxMedicare Advantage and Prescription Drug System
MASMedicare Appeals System
MBD DWMedicaid Beneficiary Database Data Warehouse
MBESMedicaid Budget &Expenditures System
MBGPMultiprotocol Border Gateway Protocol
MCMetadata Catalog
MCOManaged Care Organization
MD5Message Digest number 5
MDBMessage-Driven Bean
MDCNMedicare Data Communications Network
MDMMaster Data Management
MDMMobile Device Management
MDRMaster Data Repository
MECMultichassis Etherchannel
MEDMulti-Exit Discriminator
MEDPARMedical Provider Analysis and Review
MFAMulti-Factor Authentication
MFTManaged File Transfer
MIBManagement Information Base
MIDASMultidimensional Information and Data Analytics System
MIGMedicare Insured Group
MIIRManagement Information Integrated Repository
MIISMicrosoft Identity Integration Server
MIMEMultipurpose Internet Mail Extension
mIoTMedical Internet of Things
MISManaged Internet Service
MITAMedicaid Information Technology Architecture
MLSMulti-Level Security
MMAMedicare Prescription Drug, Improvement, and Modernization Act of 2003 (Medicare Modernization Act)
MMAMedicare Modernization Act
MMSMultimedia Message Service
MOAMemorandum of Agreement
MOLAPMultidimensional Online Analytical Processing
MOUMemorandum of Understanding
MPEGMoving Picture Experts Group
MPIOMultipath Input/Output
MPLMozilla Public License
MPLSMultiprotocol Label Switching
MQMessage Queuing
MRIMagnetic Resonance Imaging
msMillisecond
MSHTMLMicrosoft Hypertext Markup Language
MSISMedicaid Statistical Information System
MSMQMicrosoft Message Queuing
MTIPSManaged Trusted Internet Provider Service
MTOMMessage Transmission Optimization Mechanism
MTUMaximum Transmission Unit
MVMainframe Virtualization
MXMail Exchange Record
NANetwork Architecture
NACNetwork Asset Control
NAPTRNaming Authority Pointer Record
NARANational Archives and Records Administration
NASNetwork-attached Storage
NASANational Aeronautics and Space Administration
NATNetwork Address Translation
NCHNational Claims History
NCPDPNational Council for Prescription Drug Programs
NDMNetwork Data Mover
NFSNetwork File System
NICNetwork Interface Card
NIDSNetwork Intrusion Detection System
NIDSNetwork-based Intrusion Detection System
NIEMNational Information Exchange Model
NIHNational Institutes of Health
NISTNational Institute of Standards and Technology
NLRNational Level Repository
NMUDNational Medicare Utilization Database
NPDNetwork Protection Device
NPINational Provider Identifier
NPMNode Package Manager
NPPESNational Plan and Provider Enumeration System
NSName Server
NSANational Security Agency
NSEPNetwork Security Endpoint Protection
NTPNetwork Time Protocol
NVNetwork Virtualization
NVDNational Vulnerability Database
O&MOperations and Maintenance
OAGMOffice of Acquisition and Grants Management
OASISOrganization for the Advancement of Structured Information Standards
OAuthOpen standard to authorization OAuth 2.0 Authorization Framework
OCOffice of Communications
OCIOOffice of the Chief Information Officer
ODBCOpen Database Connectivity
ODCOrthogonal Defect Classification
ODSOperational Data Store
OESSOffice of E-Health Standards and Services
OFMOffice of Financial Management
OIDObject Identifier
OIGOffice of the Inspector General
OIGOffice of Inspector General
OITOffice of Information Technology
OLAOperational Level Agreement
OLAPOnline Analytical Processing
OLTPOnline Transaction Processing
OM&MOperations and Maintenance Manual
OMBOffice of Management and Budget
OMB DMOffice of Management and Budget Data Mart
ONE PIOne Program Integrity
OOBOut-of-Band
OPDIVDepartment of Health and Human Services Operating Division
ORROperational Readiness Review
OSOperating System
OSDOpen Source Definition
OSIOpen Systems Interconnection
OSIOpen Source Initiative
OSPFOpen Shortest Path First
OSSOperations Support Systems
OSSOpen Source Software
OTPOne-Time Password
OWASPOpen Web Application Security Project
P2PPoint-to-Point Messaging
PaaSPlatform as a Service
PANProcessor Area Network
PATPort Address Translation
PBPetabyte
PBARPart B Analytics Reports
PBKDFPassword-Based Key Derivation Function
PCPersonal Computer
PCAPPacket Capture
PCIPeripheral Component Interconnect
PCMPrivacy Continuous Monitoring
PDPacking and Deployment
PDAPersonal Digital Assistant
PDFPortable Document Format
PDFPortable Document File
PDOPHP Data Objects
PDRPreliminary Design Review
PEProvider Edge
PECOSProvider Enrollment, Chain, and Ownership System
PHBPer-Hop Behavior
PHIProtected Health Information
PHPPHP: Hypertext Preprocessor (PHP)
PHRPersonal Health Record
PIAPrivacy Impact Assessment
PIDProcess ID
PIIPersonally Identifiable Information
PIM-SMProtocol-Independent Multicast-Sparse Mode
PISPPolicy for the Information Security Program
PKIPublic Key Infrastructure
PL/SQLOracle Procedural Language/Structured Query Language
PL/SQLOracle Procedural Language/Structured Language
PMMPerformance Monitoring and Measurement
PMOProgram Management Office
PNGPortable Network Graphics
POA&MPlan of Action and Milestones
POCPoint of Contact
POMProject Object Model
PoPPoints of Presence
POTSPlain Old Telephone Service
PPAProject Process Agreement
PPIDParent Process ID
PQRIPhysician Quality Reporting Initiative
PR/SMProcessor Resource/Systems Manager
PRRProduction Readiness Review
PSTIBCO Platform Server
PS&RProvider Statistics & Reimbursement Report
PSLProblem Statement Language
PSTNPublic-Switched Telephone Network
PTRPointer Record
PUBPublication
Pub/SubPublication and Subscription Messaging
PVLANPrivate Virtual Local Area Network
PZPresentation Zone
QIESQuality Improvement Evaluation System
QIPSQualityNet Identity Provisioning System
QMQueue Manager
QoSQuality of Service
QTSOQIES Technical Support Office
R/SSOReduced or Single Sign-On
RARouter Advertisement
RARisk Assessment
RACFResource Access Control Facility
RACIResponsible, Accountable, Consulted, Informed
RAIDRedundant Array of Independent Disks
RAMRandom Access Memory
RAMLRESTful API Modeling Language
RAPSRisk Adjustment System
RBACRole-Based Access Control
RBSRole-Based Security
RBTRole-based Training
RCARoot Cause Analysis
RDBMSRelational Database Management System
RDPRemote Desktop Protocol (Microsoft)
RDS-COBRetiree Drug Subsidy - Coordination of Benefits
ResDACCMSResearch Data Assistance Center
RESTRepresentational State Transfer
REXXRestructured Extended Executor
RFCRequest for Comment
RFIRequest for Information
RFPRequest for Proposal
RIARich Internet Application
RIBRouting Information Base
RMRelease Management
RMFRisk Management Framework
RMHRisk Management Handbook
RMIRemote Method Invocation (Java)
RORegional Office(s) of CMS
ROLAPRelational Online Analytical Processing
RPRecommended Practice
RPRelying Party
RPRecommended Practices
RPCRemote Procedure Call
RPORecovery Point Objective
RRBRailroad Retirement Board
RSSRich Site Summary
RSSReally Simple Syndication
RTORecovery Time Objective
RTTRound-Trip Time
S/FTPSecure Shell File Transfer Protocol
S3Simple Storage Service
SASecurity Administration
SASoftware Architecture
SaaSSoftware as a Service
SAESecurity Architecture and Engineering
SAMLSecurity Assertion Markup Language
SANStorage Area Network
SANSSysAdmin, Audit, Network, Security
SASSerial Attached SCSI
SATASerial Attached Technology Adapted
SBSwing Bed
SBISoftware Build and Integration
SCSecurity Configuration; System and Communications Protection
SCSecurity Category
SCSecurity Configuration
SCSoftware Coding
SCASecurity Control Assessment
SCAPSecurity Content Automation Protocol
SCMSoftware Configuration Management
SCSISmall Computer System Interface
SDSoftware Design
SDKSoftware Developer Kit
SDKSoftware Development Kit
SDLCSystem Development Life Cycle
SDMSystem Developer and Maintainer
SDOCSupplier’s Declaration of Conformity
SESecurity
SEI®Software Engineering Institute
SEMGSecurity and Emergency Management Group
SGServices General
SGMLStandard Generalized Markup Language
SHA1Secure Hash Algorithm 1
SHA2Secure Hash Algorithm 2
SHA256SHA2 w/256-bit digest
SIASecurity Impact Analysis
SIDSystem ID
SIEMSecurity Information and Event Management
SLAService Level Agreement
SLAACStateless Auto-configuration
SLOService Level Objective
SLSScalable Login Service
SMSystem Maintenance
SMBSmall Message Block
SMESubject Matter Expert
SMSShort Message Service
SMTPSimple Mail Transfer Protocol
SNASystem Network Architecture
SNMPSimple Network Management Protocol
SOSecurity Operations
SOAService-Oriented Architecture
SOAPSimple Object Access Protocol
SOCSecurity Operations Center
SOCaaSSecurity Operations Center as a Service
SONETSynchronous Optical Networking
SOPSenior Official on Privacy
SORSystem of Record
SORNSystem of Record Notice
SPSpecial Publication
SPASingle Page Applications
SPISensitive Personal Information
SPISecurity Programming Interface
SQSoftware Quality
SQLStructured Query Language
sRGBStandard Red, Green, and Blue
SRMService Reference Model
SRPSingle Responsibility Principle
SSSecure Software
SSASocial Security Administration
SSDSolid-State Drive
SSHSecure Shell
SSLSecure Sockets Layer
SSNSocial Security Number
SSOSingle Sign-On
SSOSystem Security Officer
SSPSystem Security Plan
SSPMOShared Services Project Management Office
ST&ESecurity Test and Evaluation
STARSystem for Tracking Audit & Reimbursement
STIGSecurity Technical Implementation Guide (DISA)
STIXStructured Threat Information eXpression
SUASoftware Usage Analysis
SVServer Virtualization
SVGScalable Vector Graphics
SWSoftware
SWAMSoftware Asset Management
SWCISoftware Configuration Item
T1HVWindows Virtualization
T2HVUNIX Virtualization
TAPTest Anything Protocol
TBTerabyte
TCOTotal Cost of Ownership
TCPTransmission Control Protocol
TEKTraffic Encryption Key
TermDefinition
TICTrusted Internet Connectivity
TICTrusted Internet Connection
TICAPTrusted Internet Connection Access Provider
TIDTarget ID
TIFFTagged Image File Format
TLCTarget Life Cycle
TLSTransport Layer Security
T-MSISTransformed Medicaid Statistical Information System
ToSType of Service
TPTransport Protocol
TPWAThird-Party Websites and Applications
TPWSThird-Party Web Site
TRATechnical Reference Architecture
TRBTechnical Review Board
TSIGTransaction Signature
TSOTime Sharing Option
TSSComputer Associates Top Secret Security Access Control Program (ACP)
TTLTime-to-Live
TTP-HFPPTrusted Third Party – Healthcare Fraud Prevention Partnership
TWSTivoli Workload Scheduler
TXTText Record
UAUniversal Accessibility
UAAGUser Agent Accessibility Guidelines
UATUser Acceptance Testing
UCUnified Communications
UCDUser-Centered Design
UDDIUniversal Description, Discovery, and Integration
UDPUser Datagram Protocol
UGAUnique Global Unicast Address
UIDUser Identifier
URIUniform Resource Identifier
URLUniform Resource Locator
US-CERTUnited States Computer Emergency Response Team
USGUnited States Government
USGCBUS Government Configuration Baseline
UTCUniversal Time Coordinate
UTFUnicode Transformation Format
UXUser Experience
VATVulnerability Assessment Team
VBSVerizon Business Systems
VCVersion Control
VCSVeritas Cluster Server
VCSVersion Control System
VDCVirtual Data Center
VDCVirtual Device Context
VDDVersion Description Document
VDIVirtual Desktop Infrastructure
VDMVirtual Data Mart
VIVMware Infrastructure
VLANVirtual Local Area Network
VMVirtual Machine
VMFSVMware Virtual Machine File System
VPCVirtual Private Cloud
vPCVirtual Port Channel
VPNVirtual Private Network
VRFVirtual Routing and Forwarding
VRRValidation Readiness Review
VSAMVirtual Storage Access Method
VSNVirtual Server Network
VSSVirtual Switching Systems
VULVulnerability Management
W3CWorld Wide Web Consortium
WADLWeb Application Description Language
WAFWeb Application Firewall
WAIWeb Accessibility Initiative
WANWide Area Network
WASWebSphere Application Server
WCAGWeb Content Accessibility Guidelines
WCMWeb Content Management
WCMSWeb Content Management System
WSDLiffWeb Services Description Language
WebDAVWeb Distributed Authoring and Versioning
Wi-FiWireless Fidelity
WINSWindows Internet Name Services
WMIWindows Management Instrumentation
WPSWisconsin Physicians Service
WSWeb Services
WS-BPELWeb Services-Business Process Execution Language
WS-IWeb Services Interoperability
WSRPWeb Services for Remote Portlets
WSSEWeb Services Security Elements
XCCDFExtensible Configuration Checklist Description Format
XHTMLExtensible HyperText Markup Language
XMLExtensible Markup Language
XMLAExtensible Markup Language Administration
XMLAExtensible Markup Language Authentication and Authorization
XMLGExtensible Markup Language General
XMLPExtensible Markup Language Protection
XOPXML-binary Optimized Packaging
XPExtreme Programming
XSDXML Schema Definition
XSLTExtensible Stylesheet Language Transformation
XSSCross-Site Scripting
YUIYahoo! User Interface
ZTAZero Trust Architecture
ZTMMZero Trust Maturity Model
z/VMIBM hypervisor for the virtualization technology platform supporting IBM virtual operating systems
  

 


 

 

REFERENCE INFORMATION

TRA Business Rule Index

The TRA Business Rules Index contains links to Business Rules and Recommended Practices that are referenced in each section of the CMS TRA. Retired, deprecated, and withdrawn rules (when shown) show the text of the last TRA version that contained them. Any rationales and explanations can be found in the archived version in the CMSShare TRA[JD21] [MT22]  page.

Foundation Cybersecurity Network Infrastructure Applications Data

Foundation Business Rules
Rule IDRule
BR-F-1Any Deviations from the CMS TRA Must Be Requested and Approved
BR-F-2The CMS TRA Applies to All CMS Processing Environments
BR-F-3The CMS TRA Defines a Zoned Architecture
BR-F-4Within a CMS Processing Environment, Communication Must Flow Only between Adjacent Zones or within a Single Zone
BR-F-5Any System That Processes CMS Data Must Be Covered by a CMS ATO
BR-F-6Mainframes Must Be Dedicated to CMS
BR-F-7Cost-Effective Reuse of Data Centers with Established TIC, CMSNet, and CCIC Integration
BR-F-8Backup CMS Data
BR-F-9Test CMS Backups on a Documented Schedule
BR-F-10Annual Review and Exercise of Data Center Disaster Recovery Plans
BR-F-11Annual Review and Exercise of Contingency Plans
BR-F-12Role-Based Security AAA Must Be Used for Management and User Roles
BR-F-13Consistent Security Categorization within ARS Security Boundary
BR-F-14Applications with Disparate FIPS-199 Security Categorization Levels Must Not Be Hosted on the Same Server
BR-F-15Ensure Timely Version, Patch, and Configuration Management Practices
BR-F-16All Hosts Must Share a Common Authenticated Time Server
BR-F-17The CMS TRA Applies to Custom-Produced as well as COTS Products and Services
BR-F-18Use .Gov Domain Names for All CMS Internet Traffic
BR-F-19Communication Initiated to Internet-Based Services from ATO’d Environments Must Be Allowlisted
BR-F-20Untrusted Services and Code from Third-Party Websites and Applications
RP-F-21Limit Data in the Application and Presentation Zones
BR-F-22The CMS TRA Defines a Services Framework Architecture
BR-AI-1AI Tools and Services Must Meet Federal AI, Cybersecurity and Privacy Standards in the Handling of Sensitive Data
BR-AI-2High-Impact AI Use Cases Must Meet Minimum Risk Management Practices
BR-AI-3Foreign Entity AI Tools May Only Be Used if Deployed on CMS Infrastructure
BR-AI-4Human Review Must Follow Use of AI Tools to Write CMS Policies
BR-AI-5Do Not Rely on AI for Final Decisions for “High Impact” Cases
BR-AI-6AI-Supported Official Actions Are Subject to Records Retention Requirements
RP-AI-1Clearly define the context of the prompt
RP-AI-2Clearly define the role the AI should adopt
RP-AI-3Break down complex requests into clear, sequential steps
RP-AI-4Provide Guidelines
RP-AI-5Create New Chats or Sessions When Switching Topics
RP-AI-6Use Structured Prompts
RP-AI-7Use Iterative Refinement
RP-AI-8Release and Maintain AI Code as Shareable Open Source Software
RP-AI-9Recommended AI-Assisted Development Methodologies

Cybersecurity

Cybersecurity Business Rules
Rule IDRule
BR-SEC-Gen-1Traffic between and within Zones Must Be Available in Unencrypted Form for Security Analysis Purposes
BR-SEC-Gen-2Software and Hardware Components Must Adhere to a Secure Baseline Configuration
BR-SEC-Gen-3Disable All Unnecessary Features and Capabilities
BR-SEC-Int-3HTTPS on CMS Public-Facing Websites and Services on the Internet
BR-SEC-Gen-4Administrative Access to CMS Services and Devices
BR-SEC-Gen-17Software Assurance Measures
BR-SEC-Gen-18Malicious Code Protection in CMS Processing Environments
BR-SEC-Gen-18aSubmitted Files Must Be Scanned in the Presentation or Application Zone
BR-SEC-Gen-19User or External Inputs Must Be Validated
BR-SEC-Gen-21Malware and Malicious Code Scanning Results Must Be Sent to the Security Zone
BR-SEC-Gen-20XML Firewalls Must Be Used to Authenticate XML Exchanges
BR-SEC-Gen-22All Information Systems Must Have a System Risk Assessment in CFACTS
BR-SEC-Gen-7Vulnerability Management for CMS Networks, Services, and Devices
BR-SEC-Gen-10Periodic and Continuous Network Vulnerability Scanning Is Required
BR-SEC-Gen-11Periodic and Continuous Configuration Compliance Scanning Is Required
BR-Sec-Gen-12Host Intrusion Detection Capabilities on All IT Components
BR-SEC-Gen-15Logs Must Be Securely Collected, Aggregated, and Analyzed
BR-SEC-Gen-23CMS Network Time Protocol Services
BR-SEC-FW-1Separate Network Interfaces for Each Network Segment and Zone
BR-SEC-FW-2Adhere to CMS Security Hardening Guidance
BR-SEC-FW-3External Connections to CMS Zones
BR-SEC-FW-4Access to Services of the Firewall
BR-SEC-FW-5Filtering Traffic between Zones
BR-SEC-FW-6Firewalls Transmit Logs and Notifications to the Security Zone
BR-SEC-FW-9Network Traffic Entering a Zone Must Terminate in That Zone
BR-SEC-FW-10Impede Attempts to Traverse Network Zones
BR-SEC-FW-11Utilize Firewalls from Two or More Different Vendors
BR-SEC-FWA-1Administrative Access to Firewalls
BR-SEC-FWA-2Firewall Implementation
BR-SEC-FWA-3Disable Non-Firewall Functions
BR-SEC-FWA-4Firewall Functional Requirements
BR-SEC-ALFG-1Accept Only Enveloped or Detached Signatures
BR-SEC-ALFG-2Apply Exclusive Canonicalization
BR-SEC-ALFG-3Expand All Non-Character Entries
BR-SEC-ALFG-4Protect Against Web Service Attacks
BR-SEC-ALFG-5Reject Certain Types of Digital Signatures
BR-SEC-ALFA-1Authenticate Request Originators
BR-SEC-ALFA-2Chain of Identity, Authentication, and Authorization
BR-SEC-ALFA-3Prevent Transformations on Signed Data Elements
BR-SEC-ALFP-1Protect Against Attacks on Protocols
BR-SEC-ALFP-2Protocol Headers
BR-SEC-ALFM-1Limit Administrative Access to ALF Services
BR-SEC-ALFM-2ALF Services Must Not Support General-Purpose Computing
BR-SEC-ALFM-3ALF Service Functions
BR-SEC-ALFM-4ALF Service Configuration Must Be Restricted
BR-SEC-ALFM-5ALF Service Configuration Changes Must Be Logged and Audited
BR-CCIC-01Security Authorization of Systems
BR-CCIC-02Assessment of Information Security and Privacy Risks
BR-CCIC-03Onsite Incident Response Team
BR-CCIC-04Local Secure Enclave to Support CCIC Capabilities
BR-CCIC-05Interoperability with CCIC
BR-CCIC-06Timely Response to CCIC Requests
BR-CCIC-07Local Security Information and Event Management Capability
BR-CCIC-08Local Forensic and Malware Analysis Support
BR-CCIC-09Local Information Sharing and Cyber Threat Intelligence Support
BR-CCIC-10Penetration Testing Support
BR-CCIC-11Local Security Architecture and Engineering Support
BR-CCIC-26High Value Assets
BR-CCIC-12CyberScope Data Feeds
BR-CCIC-13Local ISCM / CDM Management Capability
BR-CCIC-14Hardware Asset Management Capability
BR-CCIC-15Software Asset Management Capability
BR-CCIC-16Configuration Settings Management Capability
BR-CCIC-17Vulnerability Management Capability
BR-CCIC-18Perimeter Monitoring Prerequisites
BR-CCIC-19Full Packet Capture Capability
BR-CCIC-20Network Intrusion Detection / Prevention Capability
BR-CCIC-21Malware Detection / Prevention Capability
BR-CCIC-22Network Firewall Perimeter Requirements
BR-CCIC-27Network Data Loss Prevention
BR-CCIC-23Network Security Endpoint Protection Capability
BR-CCIC-24FIPS 140-2 or FIPS 140-3 Validated Encryption Use
BR-CCIC-25Insider Threat Detection
BR-CCIC-28Endpoint Data Loss Prevention
BR-ACID-1Valid Purpose Required to Access CMS Information Systems
BR-ACID-2Known Identity Required to Access CMS Information Systems
BR-ACID-3Single Identity Record for Each Individual Accessing CMS Systems
BR-ACID-4User Identities Must Be Vetted and Managed Using a Common Framework
BR-ACID-5Use Personally Identifiable Information Only When Appropriate
BR-ACID-6Minimize Retention of PII in Identity Life-Cycle Management
BR-ACID-7Collect and Use Social Security Numbers Only When Necessary
RP-ACID-8Use Third-Party Data Sources for Identity Proofing and Credential Management
RP-ACID-9CMS May Delegate Registration and Identity Proofing to Employers
BR-ACID-10EUA Manages UserIDs of CMS Employees and Contractors
BR-ACID-11CMS Business Owners Provide Privilege Administration
BR-ACID-12Local User and System Accounts Must Be Auditable
BR-ACID-13OIT Is Responsible for Identity Management of Users with Credentials Provisioned in the CMS Enterprise Directory

Network Services

Network Services Business Rules
Rule IDRule
BR-WAN-S-0Use Mutual Authentication and Encrypted Tunnels between Data Centers
BR-WAN-S-1The WAN Must Implement FIPS 140-2 or FIPS 140-3 Compliant Encryption
BR-WAN-S-2CMS Business Partners Only Access the Presentation Zone
BR-WAN-S-3Communication between CMS Data Centers Is Only Permitted between Like Zones
BR-WAN-S-4Business Partner Access Restrictions
BR-WAN-S-11Customer Edge Devices Must Be Configured Securely
BR-WAN-S-12Management of CMSNet Is Via a Dedicated Logical Network
BR-WAN-IP-1WAN Services and Devices Will Be Internet Protocol Version 6 (IPv6) Capable
BR-WAN-IP-2The WAN Provider Will Use IP Space Provided by CMS
RP-WAN-IP-3CMS or Data Centers Provide IP Space for the Multi-Zone Environment
R-WAN-IP-5Extranet Business Partners Provide Their Own IP Space
BR-WAN-CM-1WAN-Related Service Requests Will Be Maintained on the CMS SOR
BR-WAN-CM-2CMSNet Services Must Be Certified Annually
BR-WAN-CM-3Ensure Timely Version, Patch, and Configuration Management Practices Relative to the Identification and Release of New Security Features
BR-WAN-M-1IP Multicast Routing Support across the CMS WAN
BR-WAN-O-2Access to the Public Internet Will Comply with HHS TIC Policy
BR-DNS-1Keep Domain Names Flat
BR-DNS-2The Physical Location of Name Servers in CMS Data Centers Must Be Easily Identifiable
BR-DNS-3Minimal DNS Changes When Transitioning from One Environment to Another
BR-DNS-4Use Stealth Primary Name Servers
BR-DNS-5Secure Means Must Be Used for Zone Transfers
BR-DNS-6Provide Segregated DNS Resolution of Internal and External Queries
BR-DNS-7Inter-Data Center Hosting of Secondary DNS Servers Is Required
BR-DNS-8Place Primary and Secondary Name Servers in Each CMS Data Center
BR-DNS-9Every “A” Record Defined Must Have a PTR Record
BR-DNS-12CMS DNS Design Must Minimize Impact to the WAN
BR-DNS-13All DNS Changes Must Be Subject to CMS’s Change Management Procedures
BR-DNS-14DNS servers deployed within ATO(ed) environments must use DNSSEC

Infrastructure Services

Infrastructure Services Business Rules
Rule IDRule
BR-SV-1Apply Separation of Duties to Virtualization Administration
BR-SV-2Provide Hypervisor Root Access Only to Specific Administrative Accounts
BR-SV-3Different Administration Account on Blade Controllers and Hypervisors
BR-SV-4Configure UserIDs and GroupIDs to Be Unique across the Processing Environment
RP-SV-5Maintenance Window Planning
RP-SV-6Consider High-Availability Configuration
BR-SV-7No Co-Hosting on Production and Non-Production Hypervisors
BR-SV-8Do Not Oversubscribe ATO(ed) Environments and Management Zones
BR-SV-9Use Storage Quotas for Virtual Machines
BR-SV-10ATO(ed) Environment VMs and Resource Pools Must Not Be Shared between Zones
RP-SV-11Collect Virtualization Performance Metrics
BR-SV-12Perform Asset Management of Virtual Instances
BR-SV-13Keep Forensic Evidence per CMS Security Rules
RP-SV-14Use VM Configuration Templates
BR-SV-15Production Management Zone VMs May Not Use IP Multipathing
BR-SV-16Originate Administrator Access to Blade Controllers and Hypervisors from the Management Zone
BR-SV-17All Management Traffic Must Originate or Terminate in the Management Zone and Use Only Isolated and Protected Interfaces
BR-SV-18Hypervisor Access Is Permitted Only Via the Management Interface
BR-SV-19Separate Security Segment from All Other Management Zone Segments
BR-SV-20Oversubscription of Non-Production Instances Is Permitted
BR-SV-21No Business Applications May Run on the Hypervisor’s Host OS
BR-SV-22Operate Applications under Application-Specific System Accounts
BR-NV-1Use Highly Available Network Services to Implement Zone Separation
BR-MV-1z/VM Sole Configuration
BR-MV-2z/Linux as Sole Guest for z/VM
BR-MV-3No Nested z/VM Instances
BR-MV-4Multiple Zones for a Single Application
BR-CI-1Choosing a Cloud Deployment and Service Model
BR-CI-3Document the Impact of Cloud Deployment
BR-CI-4Engage TRB Consulting for CMS-Owned Equipment
BR-CI-5Acquisition of New IaaS or PaaS Cloud Service Providers
BR-CI-6Define Data Backup and Contingency Plans in a Cloud
BR-CI-7Cloud Resource Capacity Planning
BR-CI-8Separate Production, Management, and Non-Production Resource Clusters
BR-CI-9Applicability of Multi-Zone Architecture
BR-CI-11Cloud Services Covered by a CMS ATO May Be Used in Lieu of Virtual Network Elements
BR-PMM-1Performance Monitoring Data Is FOUO
BR-PMM-2Control Access to Performance Monitoring Data
BR-PMM-3Monitor Production Environments
RP-PMM-1Coordinate Application Changes with Monitoring Operations
RP-PMM-2Consider Monitoring Lower Environments
RP-PMM-3Provide Performance Data to CMS NOC
RP-PMM-4Conduct Performance Management Planning
RP-PMM-5Use a Trouble Ticketing System to Track Performance Problems
RP-PMM-6At Least One End-to-End Test
RP-PMM-7Identify Unmonitorable Components as a Risk
RP-PMM-8Support Standards-Based Tools
RP-PMM-9Define Services in a Service Catalog
BR-SAAS-1SaaS Clouds Are Defined by NIST SP-800-145
BR-SAAS-2SaaS Must Have a CMS ATO
BR-SAAS-3Ensure CMS Security May Perform Periodic Security Assessments
BR-SAAS-4Plan for Data Archival to Comply with Federal Records Management
BR-SAAS-5Establish SaaS-Specific Contingency Program
BR-SAAS-6Perform Configuration Management
BR-SAAS-7Integrate with CMS CCIC
BR-SAAS-8CMS Data Must Always Reside in the U.S.
RP-SAAS-1Integrate with CMSNet
RP-SAAS-2Comply with CMS Defense-in-Depth Architecture
RP-SAAS-3Continuous Monitoring
RP-SAAS-4Integrate SaaS with CMS Identity Management Systems
BR-KSM-1KSM Auditing Must Be Enabled and Connected to CMS Logging Infrastructure
BR-KSM-2KSM Must Use RBAC with Separate Roles by CMS Application Environment
BR-KSM-3Ensure the KSM Has Sufficient Availability for Your Applications
BR-KSM-4Audit Logs for Credential Leaks
RP-KSM-1Applications Should Use a KSM to Manage Keys and Secrets
RP-KSM-2Configure Applications to Use Dynamic Credentials
RP-KSM-3Consider Credential Injection into VM Images during Startup
BR-MD-1Mobile Devices (GFE and non-GFE) Must Be Authorized to Access CMS Systems and Government Data
BR-MD-2Mobile Devices Must Use Encrypted Communication to Access CMS Data
BR-MD-3Managed Mobile Devices Must Support Remotely Erasing All Stored Data If the Device Is Lost, Stolen, or Compromised
BR-MD-4Managed Mobile Devices Must Support Encryption for Internal and Removable Storage
BR-MD-5Managed Mobile Devices Must Meet CMS Security Requirements
BR-MD-6Only CMS Authorized Applications May Be Installed on CMS-Managed Devices
BR-MD-7User Agreements Must Be In Place for Mobile Devices Accessing CMS Services and Data
BR-IoT-1CMS IoT Platforms Must Comply with CMS Requirements for CMS Processing Environments
RP-IoT-1CMS-Managed IoT Devices Should Comply with NIST SP 1800-15 and the Latest MUD Specifications
BR-DR-1Annual Review of Disaster Recovery Plans
BR-DR-3All CMS FISMA systems must have a plan for DR
BR-DR-4Required Risk Analysis, System BIA, and ISCP
BR-DR-6The BIA is the Primary Determinant of DR Parameters

Application Development

Application Development Business Rules
Rule IDRule
BR-ADM-1Use of the CMS Life Cycle Is Mandatory
BR-ADM-2The Development Methodology and Artifacts Must Be Documented
BR-SA-1Use CMS Shared Services
BR-SA-2Integrate with the CMS Identity Management Services
BR-SA-3No Custom Application Code Is Permitted in the Presentation Zone
BR-SA-4Use CMS-Validated Mediation and Data Access Services to Access Data in the Data Zone
BR-SA-5No Long-Term, Persistent Sensitive Application Data Storage in the Presentation or Application Zones
BR-SA-6Network Communications Must Meet the TRA Rules for Encryption
BR-SA-7Substantive Changes to the Architecture, Products, or Technology of an Existing Application Must Be Documented and Reviewed by the CMS TRB
BR-SA-8Logging Must Be Configurable and Use Common Platform Standards
BR-SA-9Systems Must Define Metrics for IT Health Monitoring
BR-SA-10Applications in CMS Data Centers May Not Use Some Native Email Protocols
RP-SA-11Servers Should Include Instrumentation for Application Performance Monitoring
RP-SA-12Minimize Manual File Copying by Using Integrating File Transfer Automation
RP-SA-13Consider Data Services in the Data Zone to Improve Performance of Database-Intensive Services
BR-SA-14Use of Short Message Service / Multimedia Message Service by CMS Applications
BR-SA-15Protect Sensitive Information in Transit
BR-SA-16Protect Sensitive Information at Rest
BR-SD-1External Configuration Is Mandatory
BR-SD-2Web-Based User Interfaces Must Comply with TRA Guidance
RP-SD-3Configurations Should Be Validated and Checked on Each System Startup
RP-SD-4Consider Dependency Injection to Achieve External Configuration
RP-SD-5External System Dependencies Should Be Stubbed Out for Development and Testing
RP-SD-6Timestamps Logged by the System Must Be in UTC or GMT and Should Be Expressed in ISO-8601 Format
RP-SD-7Software Should Be Designed Based on SOA Principles
RP-SD-8Consider Non-Blocking Service Implementations to Improve Performance and Scalability
BR-SC-1Inventory all Open Source Software Licenses
BR-SC-2All Custom-Written Source Code for a Project Must Conform to an Identified Coding Standard
BR-SC-3All Custom-Written Code for a Project Must Be Shareable
RP-SC-3Do Not Intermingle Code in Different Programming Languages in the Same File
RP-SC-2Capture Code Metrics and Defect Tracking Metrics for Quality Improvement Purposes
RP-SC-5When Using Flat Files for Data Transfer, Include Helpful Metadata in the File
RP-SC-6When Using Flat Files for Data Transfer, Consider Including a Machine-Readable Schema
RP-SC-7Use Decimal Math Types for Financial Calculations
RP-SC-8Consider Synthetic Transactions
RP-SC-9AI-Generated Code Should Undergo Human Review
RP-SC-10AI-Generated Code Should Protect Sensitive Information
RP-SC-11Configure Content Exclusions for AI-Generated Code
RP-SC-12Implement Code Review Safeguards for AI-Generated Code
BR-SQ-1All Custom-Written Software Must Have Associated Automated Unit Tests
BR-SQ-2Run Automated Unit Tests during Full Builds
BR-SQ-3Automated Unit Tests Must Use a Commercially Available Unit Testing Framework or Test Runner
BR-SQ-4All CMS User Interfaces Must Meet Section 508 Accessibility Requirements
BR-SQ-5Manual Code and Design Reviews Are Mandatory
BR-SQ-6De-Identification of Production Data Is Required in Non-Production Environments
RP-SQ-7Code Coverage Analysis Is Highly Encouraged During Unit Testing
RP-SQ-8Use Static Analysis Tools During Build to Catch Common Coding Errors
RP-SQ-9Developers Assist Testers in Generating Test Data
BR-SS-1All Software on CMS Production Servers Must Have Recorded Provenance
BR-SS-2Use NIST SP 800-132-Specified Password-Based Key Derivation Functions (PBKDFs)
BR-SS-3SQL Code Must Use Binding Variables
BR-SS-4Check for Common Security Vulnerabilities
BR-SS-5Use Static Analysis Tools to Catch Common Security Weaknesses
RP-SS-6Use Profiling to Perform Dynamic Code Analysis
BR-SS-7Error Handling Must Not Reveal Information That Could Lead to an Exploit
RP-SS-8Perform Threat Modeling During the Design Phase to Identify Potential System Threats
BR-ED-1Custom-Written Software Must Include Inline Documentation for Public APIs
BR-ED-2The CMS TLC Phase Review Artifacts Must Be Produced
RP-ED-3Engineering Documentation Should Be Versioned Along with Source Code in the Same Repository
RP-SM-1Consider Building Self-Diagnosis Capability into Systems
RP-SM-2Consider Designing Maintenance Capability into Systems
BR-DBM-1Systems Must Meet Federal Record Management Requirements
BR-DBM-2Systems Must Meet Federal Government FOIA Requirements
BR-DBM-3Systems Must Meet CMS Data and Database Management Standards
BR-SCM-1All Source Code Must Be Checked in to Version Control
BR-SCM-2All Code Must Be Baselined Prior to Release into Implementation, Validation, and ATO(ed) Production Environments
RP-SCM-3Apply Database-Oriented Configuration Management Practices
BR-SCM-4Configurations Must Be Checked in to Version Control
BR-DIT-1All CMS Software Development Projects Must Use a Defect Tracking System
BR-DIT-2A Defined Defect Classification Standard Is Mandatory
RP-DIT-3Defects Should Be Correlated to Baselines
BR-SBI-1All Builds Must Occur in Controlled Environments
BR-SBI-2All Production-Deployed Custom Code Must Be Built and Installed from Version-Controlled Source Code
BR-SBI-3Production Builds Must Have Zero Compile Errors
RP-SBI-4Use Explicit Library and Build Dependency Management
RP-SBI-5Consider Instituting Continuous Integration
BR-PD-1Software Must Be Packaged for Deployment
BR-PD-2Software Target Packaging Must Be in Either the Operating System or Language Platform Native Form
BR-PD-3Database Changes Must Include Back-Out Scripts
RP-PD-4The Package Manifest Should Include a List of All Defects Corrected in the Release
RP-PD-5Changes Applied to Databases Should Be Recorded in the Database Itself
RP-PD-6Support A/B Testing of User Interfaces
BR-D-1Developers Do Not Have Unsupervised Administrative Access to Production Servers
BR-D-2All Installation and Back-Out Scripts Must Have Been Tested in Lower Environments Prior to Use in Production
RP-D-3Support Rolling Deployment
RP-D-4Use Feature Flags to Gradually Introduce New Features to Users
RP-D-5Deployment Should Integrate with Monitoring to Coordinate Outages
RP-D-6Support Rollback of Package Installation
RP-D-7Support Automated Startup, Shutdown, and Maintenance Mode Entry / Exit
RP-RM-1Establish and Follow Organizational Standards for Deployment of Custom Software
BR-WS-1Describe All Services
BR-WS-2Services Must Use Standard Invocations
BR-WS-3Services Must Validate Input and Outputs
BR-WS-4Use TRB-Approved Data Zone Mediation and Data Access Services to Access Data in the Data Zone
BR-WS-5Messages Must Include Timestamp and Originator
BR-WS-6Services Must Use Open Data Formats
BR-WS-7Web Services Must Follow CMS Encryption Policy
BR-WS-8SOAP-Based Services Must Comply with WS-* Standards
BR-WS-9Inter-Zone Web Services Must Transverse a Mediated Service
BR-WS-10Web Services Must Be Version Numbered
BR-WS-11Use Certificate-Based Mutual Authentication for Machine-to-Machine Web Services
BR-WS-12Messages Must Pass through All Intermediate Zones
BR-WS-13CMS Public APIs Must Be Published
BR-UX-1Ensure Usability and Accessibility
BR-UX-2Collect Feedback
BR-UX-3Provide Multilingual Capability
BR-UX-4Ensure Privacy and Security
BR-UI-5Skip Navigation
BR-UI-6Provide a Home Page Link
BR-UI-7No Frames
BR-UI-8Keyboard and Mouse
BR-UI-9Phishing and Redirection Prevention
BR-UI-10Cross-Site Request Forgery Prevention
BR-UI-12Authentication
BR-UI-13Session Management
BR-UI-14Protecting PHI
BR-UI-15User Identifiers
BR-UI-16Input Validation
BR-UI-18Need to Disclose
BR-UI-19Color Contrast Ratio
BR-UI-25Scripts Compatibility
BR-OSS-1Products in Use at CMS that Provide the Required Functionality Are Preferred
BR-OSS-2Criteria for Evaluation
BR-OSS-3Total Cost of Ownership
BR-OSS-4License Compatibility for Using OSS
BR-OSS-5Use OSS Built from a Controlled Source
BR-OSS-6Binary Package Management Is Mandatory
BR-OSS-10CMS OSS Code Released as CMS-Managed Code Requires a Governance and Support Model
BR-OSS-11CMS OSS Code Released as Unmanaged Code Must Be Identified as Such
BR-OSS-12CMS-Released OSS Code Must Include Automated Unit Tests, Build Scripts and Be Checked for Software Vulnerabilities
BR-OSS-13CMS-Released OSS Code Must Include Documentation Accessible to the Open Source Community
BR-OSS-14Use the CMS’s External GitHub Repository and Code.gov for CMS-Released OSS Code
BR-OSS-15Publish project metadata in an open source repository to support agency software inventory
RP-OSS-1Provide Ample Documentation with CMS-Released OSS Code
RP-OSS-2Implement the Tools to Support the Community Around a CMS-Released OSS Project
RP-OSS-3Use codejson-generator and automated-codejson-generator to add and update project metadata
BR-P-1Each Portlet Must Include a Deployment Descriptor as Specified in JSR 286
BR-P-2Each Portlet Must Implement Specific Portlet Life-Cycle Management Functions as Specified in JSR 286
BR-P-3Each Portlet Must Log Events via Portlet Container Logging Functions
BR-P-4If Functionality to Customize / Personalize the Portlet User Interface Is Provided, It Must Be Consistent with JSR 286
BR-P-5Each Portlet Must Control Access to Content / Functionality Based on User and Role Information via the Portal
BR-P-6Each Portlet Must Use Portal Services to Authenticate Users
BR-P-7Each Portlet Must Securely Transport Sensitive Content
BR-P-8Inter-Portlet Communication Must Be Performed Only by Either Public Render Parameters or Events as Specified in JSR 286
BR-P-9Remote Portlets Must Follow Web Service Standards
BR-CA-1The CMS Zonal Architecture Must Be Preserved
BR-CA-2Lower Environments Must Be Separated from Production
BR-CA-3The CMS TRA Zonal Hierarchy Will Be Enforced
BR-CA-4Interfaces between the Containers Implemented in Each Zone Must Be Locked Down (Source / Destination)
BR-CA-5Container Traffic to Non-Container or External Destinations Must Follow Existing TRA Rules
RP-CA-1Force Containers to Write to Container-Specific File Systems
RP-CA-2Implement Read-Only File Systems Whenever Possible
RP-CA-3Run Your Containers as Non-Root Whenever Possible
RP-CA-4Create Containers with the Least Privilege Possible
RP-CA-6Use a Security Mechanism for Mandatory Access Controls
RP-CA-7Use Cgroups
RP-CA-8Use a Secure Computing Mode Profile
RP-CA-9Harden All Containers / Components
RP-CA-10Validate All Third-Party Containerized Applications before Implementation
RP-CA-11Maintain the Immutability of Your Containers
BR-OR-1Container Images Must Be Hardened
BR-OR-2Use Security Monitoring on Containers
BR-OR-3The Deployment Infrastructure for Containers Must Be Hardened and Monitored
BR-OR-4Containers Use Must Respect the Multi-Zone Architecture
BR-OR-5Libraries of Containers Must Be Maintained in CMS-Only Stores
BR-OR-6Required Orchestration Capabilities
RP-OR-7Prefer Container Orchestration Tools That Allow for Container Motion
RP-OR-8Rolling Upgrades
BR-LD-1Avoid Using Lambda for CMS Sensitive Data
BR-LD-2Integrate with CMS Enterprise Security
BR-LD-3Configure Lambda to Control the Cost / Budget
BR-LD-4Do Not Use the Local Volatile Storage of Lambda Containers for Persistent Data
BR-LD-5Use Lambda Only for Stateless Transactions
RP-LD-1Avoid Complex Application Programming or Workflows in Lambda
BR-CM-1Projects Must Produce a Configuration Management Plan
BR-CM-2Projects Must Identify Items to Be Placed under Configuration Control
BR-CM-3Significant Changes to Configuration Items of a System or Component Managed by a CCB Requires the Approval of That CCB
BR-CM-4Projects Must Maintain Accurate and Reliable CI Information
BR-CM-5Projects Must Conduct Periodic Audits of CM Activities and Products
BR-CM-6Projects Must Maintain Configuration Baselines
BR-CM-7Follow the CMS ARS Hierarchy for Security Configuration

Data

Data-Related Business Rules
Rule IDRule
BR-DM-1CMS TRA Compliance
BR-DM-2Data storage is to be separated from compute
BR-DM-3Data assets are not to be copied or moved
BR-DM-4Shared data assets are to be registered in Snowflake and a user-facing data catalog where available
BR-DM-5The data mesh does not share raw data or unstructured data. All data in the EDM is fully structured and immediately consumable
BR-DM-6Data sets are to remain within the data owner’s security boundary
BR-DM-7Data owners are required to curate their data assets and manage freshness and usability
BR-DM-8Data consumers bring their own compute resources
BR-DM-9The data owner is responsible for determining the users, groups, roles, and policies that govern data access
BR-DG-1All CMS enterprise data must be stored within a CMS authorization boundary
RP-DG-2Any CMS enterprise data sharing beyond CMS authorization boundaries should “share-in-place” where feasible, for example, a workspace or a remote API , avoiding file export or replication
BR-EFT-2Limited Protocols Are Permitted in the CMS EFT System
BR-EFT-3Use Mailboxes between Business Partners Only
BR-EFT-4Registration Is Required When Using the CMS EFT System
BR-EFT-5Internal Integrity Validation Is an Application Responsibility
BR-EFT-6File Encryption Is an Application Responsibility
BR-EFT-7Secured Transmission Is Required
BR-EFT-8IP-Based File Transfer Protocols Only
BR-EFT-10Encrypt Files Residing in EFT Mailboxes
BR-EFT-11CMS Data Files May Only Be Transferred to the Data Zone
BR-EFT-12CMS Data in ATO’d Environments May Not Be Transferred to Non-ATO’d Environments
BR-EFT-13CMS Data May Not Be Transferred Outside of CMS Processing Environments without a Prior Agreement
RP-EFT-1Transfer Files Using Check-Pointing Protocols When Performance Impacts Are a Concern [Formerly BR-EFT-9]
BR-DSS-1Do Not Use UDP for Data Storage Services or File Transfer Services
RP-DSS-2Use a Trust Relationship or Federation for Authenticated Access to CMS Files
RP-DSS-3Avoid Unnecessarily Separating Storage Clients and Storage Services into Different Zones
RP-DSS-4External Storage Services Are External to Any CMS Zone
RP-DSS-5A Storage Service Data Store Should Not Be Accessible from More Than One Zone
BR-URL-1Authentication Is Required to Access a Non-Public CMS File Referenced by a URL
BR-URL-2Authentication and Authorization Are Required to Upload a File to a CMS Location Referenced by a URL
BR-URL-3Logs Must Identify the Users Who Access a CMS File Referenced by a URL from the External Networks or CMSNet
BR-URL-4Logs Must Identify the Users Who Upload a File to a CMS Location Referenced by a URL from the External Networks or CMSNet
BR-URL-5URLs May Not Include Passwords or Decryption Keys
RP-URL-6Use Signed URLs to Indicate the Source and Creation Time of a CMS File
RP-URL-7Use Time Limits and/or Click Limits to Control How Long a CMS File Is Available to Unauthenticated External Users for Download through a URL
BR-AWS-1Encrypt Amazon S3 Data at Rest
BR-AWS-2The S3 Storage Must Be Attached / Accessible from Not More Than One Zone
BR-AWS-3Do Not Allow Cross-Zone Access to S3
BR-AWS-4External Content in S3 Requires a Data User Agreement Process
BR-AWS-5Users Must Be Authenticated by AWS S3 Using a CMS-Managed IAM UserID When Accessing Data in CMS S3 Buckets
BR-AWS-6External Data Upload to S3 Is Not Permitted
BR-AWS-7No Direct Internet Access to an Application’s Main S3 Storage
BR-AWS-8Remove Data from S3 When They Are No Longer Needed
BR-AWS-9Remove Unneeded S3 Buckets
RP-AWS-1Use Amazon AWS Security / Access Features for S3
RP-AWS-2Use Short-Lived URLs for External Access to S3
BR-BI-1All Production BI Applications and Systems Must Comply with the CMS TRA and the BI Reference Architecture
BR-BI-3Authentication, Auditing, and Logging of All BI User Accounts Must Be Managed from the CMS Enterprise LDAP Directory and Enterprise User Administration
BR-BI-4All Traffic Must Be Encrypted between a BI User’s Browser, Web Services, or Device and the BI Server
BR-BI-5Role-Based Authorization Must Be Used to Manage Access to BI Applications, Queries, Reports, Analytic Functions, Tables, Views, and Stored Procedures
BR-BI-6All BI Applications and Systems Must Comply with the Current CMS ARS
BR-BI-7BI Applications Must Leverage the BI Portal Framework as an Enterprise-Wide Secure Gateway that Enables All BI Tools to Access, Manipulate, and Analyze Data

 

TRA References

The TRA References contains references in each section of the CMS TRA.

Foundation

Network Services

CMS Network Services

Security Services

OMB Memorandum M-17-06, Policies for Federal Agency Public Websites and Digital Services, November 8, 2016, rescinded by M-23-22

OMB Memorandum M-23-22, Delivering a Digital-First Public Experience, September 22, 2023

CCIC Integration

Wide Area Network Services

Access Control and Identity Management

Domain Name System Services

Infrastructure Services

Virtualization

Cloud IaaS and PaaS Infrastructure

IT Performance Management

File Transfer

Internet of Things (IoT)

Disaster Recovery

Application Development

Centers for Medicare & Medicaid Services (CMS) Publications

Executive Branch Guidance

Defense Information Systems Agency (DISA) Guides

National Institute of Standards and Technology (NIST)

Additional References

  • Managing the Software Process, Watts Humphrey, Addison-Wesley, 1990.
  • CWE™/SANS TOP 25 Most Dangerous Software Errors, SANS Institute
  • SAFECode, “Fundamental Practices for Secure Software Development”, 3rd Edition, March 2018
  • “Common Weakness Enumeration (CWE™)”, The MITRE Corporation
  • Viega and Messier, Secure Programming Cookbook for C and C++: Recipes for Cryptography, Authentication, Input Validation & More, O’Reilly, 2003.
  • Martin, R. C. Clean code: a handbook of agile software craftsmanship. Pearson Education, 2008.
  • Beck, K. Test-driven development: by example. Addison-Wesley Professional, 2003.
  • Gilb, T., Graham, D., & Finzi, S. (1993). Software inspection. Addison-Wesley Longman Publishing Co., Inc., 1993.
  • Meszaros, G. xUnit test patterns: Refactoring test code. Pearson Education, 2007.
  • Capers Jones, “Software Engineering Best Practices: Lessons from Successful Projects in the Top Companies”, McGraw-Hill, 2010. 
  • Watts S. Humphrey, Managing the Software Process, Addison-Wesley, 1990.
  • John Ousterhout, “Why threads are a bad idea (for most purposes)”, September 28, 1995

CMS Standards

Best Commercial Practices

  • Duvall, Paul, Matyas, Steve, and Glover, Andrew, Continuous Integration: Improving Software Quality and reducing Risk, Addison Wesley, 2007.
  • Humble, Jez and Farley, David, Continuous Delivery: Reliable Software Releases Through Build, Test, and Deployment Automation, Addison Wesley, 2011, ISBN 978 0 321 60191.9.

Web Services and Web APIs

CMS Standards

NIST Special Publications

Industry Security Standards

REST Standards

REST References

SOAP Standards

Web-based UI Services

Open Source Software

Portlet Services

Business Intelligence

  • CMS Business Intelligence Strategy, Version 1.5, CMS, December 9, 2008.
  • Cognos ReportNet Guidelines, CMS, March 3, 2006.
  • MicroStrategy 8 Guidelines, CMS, March 16, 2006.
  • CMS Integrated Data Strategy, Draft, CMS / Office of E-Health Standards and Services (OESS), August 2007.
  • CMS Acceptable Risk Safeguards (ARS)
  • CMS MicroStrategy System Security Plan, Version 1.0, Draft, April 18, 2010.

Containers and Microservices

Input Validation

Configuration Management

Zero Trust

Internal

External

OMB

NIST

CISA

 

TRA Glossary

The TRA Glossary contains terms referenced in each section, along with their definition or explanation.

Foundation

TermDefinition
Artificial Intelligence
Sensitive Personally Identifiable Information (SPII)
  • Information that could be used to identify an individual, especially federal employees or members of the public. Examples include Social Security numbers, Alien Registration Numbers, Passport Details.
  • Human genomic, phenotypic data, or genetic sequencing information that can be linked to an identifiable individual.
  • Research data on genetic engineering, gene editing (e.g., CRISPR), or synthetic biology that may present dual-use or biosecurity concerns.
  • Illustrative Data Elements/Examples:
    • Social Security number (full or truncated)
    • Driver’s-license or state-ID number
    • Bank-account or credit-card number with name
    • Passport number
    • Full date of birth linked to name
  • Governing Authority:
    • Privacy Act of 1974, 5 U.S.C. § 552a
    • OMB Circular A-130
Protected Health Information (PHI)
  • Health information that is protected by the Health Insurance Portability and Accountability Act of 1996 (HIPAA) Privacy, Security, and Breach Notification Rules (commonly known as the “HIPAA Rules”). PHI includes most individually identifiable health information maintained or transmitted by HIPAA covered entities (health plans, health care clearinghouses, and most health care providers) or their business associates in any form or media. Examples include information in medical, billing, payment, and claims records. PHI does not include information that has been de-identified in the manner specified in the HIPAA Privacy Rule.
  • Illustrative Data Elements/Examples:
    • Medical record or patient chart numbers
    • Laboratory or imaging results tied to an individual
    • Health-plan beneficiary or enrollee ID
    • Appointment schedules containing patient identifiers
  • Governing Authority:
    • Health Insurance Portability and Accountability Act of 1996 (HIPAA), 42 U.S.C. § 1320d et seq.
    • HIPAA Privacy & Security Rules, 45 C.F.R. §§ 160, 164-
Classified Information
  • National security information (Top Secret, Secret, Confidential).
  • Certain CUI categories such as law enforcement sensitive data, export control information, critical infrastructure, or sensitive security information (SSI).
  • Illustrative Data Elements/Examples:
    • Documents or data marked Top Secret, Secret, Confidential
    • Restricted data under the Atomic Energy Act
    • Sensitive Compartmented Information (SCI)
  • Governing Authority
    • Exec. Order 13526 (as amended)
    • 32 C.F.R. Parts 2001–2004-
Export Controlled Data
  • Information that is commercially sensitive or has national security implications, such as genetic engineering methods, select agent research, or strong encryption technologies.
  • Biomedical, pharmaceutical, information technology, or public health research involving technologies, materials, or data with potential dual-use applications.
  • Illustrative Data Elements/Examples:
    • Missile-guidance source code (ITAR Cat. IV)
    • 3-D-print files for firearm components
    • Design specs for dual-use satellite sensors (EAR 9E515)
  • Governing Authority:
    • International Traffic in Arms Regulations (ITAR), 22 C.F.R. Parts 120–130
    • Export Administration Regulations (EAR), 15 C.F.R. Parts 730–774-
Confidential Commercial Information or Trade Secret Data
  • Proprietary or confidential commercial information, such as trade secrets, formulas, or manufacturing processes.
  • Financial or economic data that could influence markets if disclosed, including pre-decisional budget details or pricing strategies.
  • Information with potential national security implications, including technology development plans or sensitive intellectual property.
  • Illustrative Data Elements/Examples:
    • Strategic pricing models or market analyses
    • Proprietary chemical formulas
    • Non-public manufacturing processes
  • Governing Authority:
    • Trade Secrets Act, 18 U.S.C. § 1905
    • FOIA Exemption 4, 5 U.S.C. § 552(b)(4)
Zero Trust
Zero TrustA collection of concepts and ideas designed to minimize uncertainty in enforcing accurate, least privilege per-request access decisions in information systems and services in the face of a network viewed as compromised. Zero trust is a cybersecurity paradigm focused on resource protection and the premise that trust is never granted implicitly but must be continually evaluated.
Zero Trust ArchitectureAn enterprise’s cybersecurity plan that uses zero trust concepts and encompasses component relationships, workflow planning, and access policies. Therefore, a zero trust enterprise is the network infrastructure (physical and virtual) and operational policies that are in place for an enterprise as a product of a ZTA plan.
Zero Trust Maturity ModelThe Zero Trust Maturity Model represents a gradient of implementation across five distinct pillars, in which minor advancements can be made over time toward optimization. The pillars include Identity, Devices, Networks, Applications and Workloads, and Data. Each pillar includes general details regarding the following cross-cutting capabilities: Visibility and Analytics, Automation and Orchestration, and Governance. The path to zero trust is an incremental process that may take years to implement.

Infrastructure Services

TermDefinition
Virtualization
Non-ProductionDevelopment, Test, and Integration environments; IT resources intended to develop, updating, and maintain the current and future Production environments. Non-Production environments are Non-ATO(ed) environments as long as they do not contain CMS sensitive data.
ProductionProduction environments contain IT resources intended to operate and maintain CMS data and systems. Production environments are required to have a CMS ATO.
Application Performance Management
Application HealthGenerically refers to any metrics that measure the availability and performance of an application or that measure a resource that may immediately or eventually affect the availability and performance of an application.
Acceptable Performance StandardContractually agreed-upon range of a performance measure. Performance that is outside this range may be subject to penalties or incentives.
Application Performance Monitoring (APM)Monitoring that focuses on an application’s business services and transactions rather than the servers and network components that support an application.
Application Response Measurement (ARM)A protocol used by applications or middleware to pass performance information about transactions to management software.
Business Service MetricsMeasures of the availability or performance of a business service as provided by an application.
Degradation “Red Line”Range of a performance measure necessary to meet some worst-case requirement. May reflect a dependency by some other service or business operation performance goal. The point at which performance is so degraded that users or business operations are significantly impacted. This is the point at which a Level 1 incident ticket is generated and the CMS Enterprise Operations Center (EOC) is alerted.
Event CorrelationProcess of correlating multiple events using decision logic to determine a measurement of a multiple-event process or to detect a pattern indicating a performance issue.
Event EnrichmentInformation and analysis added to an event’s description defined by a set of Event Situation Rules. For example, a description of the business impact of that event or a note that the event may be avoided by adjusting cache size.
Event ManagementActivity of monitoring and taking action on events.
Event Situation RulesPredefined rules consisting of metrics, thresholds, conditions, and actions. When the conditions of a situation have been met, an event occurs and predefined actions are triggered such as displaying an event indicator on the IT PM portal dashboard.
Event ThresholdsDefine a range of performance for a given metric. An event occurs whenever measured performance exceeds an event threshold, and exceeding event thresholds may trigger some predefined action(s). More than one threshold for a given metric may be defined, representing Logged, Yellow, and Red condition events. Unless used for SLA compliance monitoring, thresholds may be adjusted at any time for a variety of diagnostic, performance management, or business reasons.
IT Operational Level Agreement (OLA)Defines the interdependent relationships among the internal support groups of an organization working to support an SLA and the responsibilities of each internal support group toward other support groups.
Maximum DegradationThe worst recorded value of a performance measure.
Performance Management PlanIdentifies specific application monitoring requirements, business service metrics, Key Performance Indicators (KPI), and thresholds for alerts and SLA violations. Also identifies contacts to be notified in case of performance degradation issues.
IT Performance Management (IT PM)System that provides both heartbeat monitoring and performance analytics of data center operations, including service delivery reports by the data center operations contractors. Delivers data for managing application performance SLAs and provides the service-level performance reports to CMS through high-level dashboards.
Reporting ThresholdsSimilar to Event Thresholds but used for performance reporting summaries and traffic lights on dashboards. Often a percentage of the SLA Acceptable Performance Standard.
Service Level Agreement (SLA)A negotiated agreement between a service provider and the customer that defines services, priorities, responsibilities, guarantees, and warranties by specifying levels of availability, serviceability, performance, operation, or other service attributes.
Service Level Objective (SLO)The acceptable range of performance outside of which penalties apply. Same as Acceptable Performance Standard.
Stress TestingA form of system testing in which a system or application is subjected to a heavy workload.
Target RangeRange of a performance measure that is expected to be the norm. Condition “Green.”
Application HealthGenerically refers to any metrics that measure the availability and performance of an application or that measure a resource that may immediately or eventually affect the availability and performance of an application.
Reporting ThresholdsSimilar to Event Thresholds but used for performance reporting summaries and traffic lights on dashboards. Often a percentage of the SLA Acceptable Performance Standard.
Mobile Device Management
Jailbreaking (iOS)On Apple devices running iOS and iOS-based operating systems, jailbreaking is the use of a privilege escalation exploit to remove software restrictions. This is similar to "rooting" an Android device.
Rooting (Android)Rooting is the process by which users of Android devices can attain privileged control (known as root access) over various subsystems of the device, usually smartphones.

Application Development

TermDefinition
Application Development
IdempotentAn idempotent operation is defined as one for which the side effects of N > 0 identical requests is the same as for a single request. Please refer tosection 9.1.2 of RFC 2616.
Release Management
Baseline

A known configuration of files at a given time that is the minimum of what can be released.

A baseline consists of a set of CIs and their parts (requirements, source code, data, build scripts, configuration files, test scripts, etc.) at a point in time. Baselines change over time through approved CRs.

Break-FixThe correction of defects identified in the production environment.
Change Request (CR)The Change Request (CR) for a CI consists of a set of requirements (for waterfall or iterative projects) or a set of User Stories (for Agile projects). The CR includes an impact analysis as well as additional supporting material as needed to justify changing one or more CIs.
Configuration Item (CI)The minimum testable unit of software is a Configuration Item, which may be a component of infrastructure, an enterprise service, a mainframe business application, a Service-Oriented Architecture (SOA) web application, or an SOA service. A CI contains files that may or may not be considered CIs.
Distributed SystemA system implemented on a distributed operating system platform, such as UNIX / Linux / zLinux or Microsoft Windows.
EnvironmentAn environment is a separated set of systems with a designated purpose. Typically, there are four CMS environments: Development, Test, Implementation, and Production.
InfrastructureAny hardware, software, or networking that is not a Centers for Medicare & Medicaid Service (CMS) business application (e.g., switches, servers, and firewalls).
Non-ProductionDevelopment, Validation, and Implementation environments. Also referred to as “lower environments”
Pre-ProductionA term often synonymous with the “Implementation” environment. CMS discourages use of this term.
PromotionThe process of approving and deploying a baseline through a succession of environments. A baseline that fails testing can be demoted to a prior test environment, if the project process allows for this.
Root Cause Analysis (RCA)Documented analysis of an issue or outage. The analysis must answer all questions related to an issue or outage. The focus in performing an RCA is not on employee error, but rather, on all supporting systems to reduce the risk of incident recurrence.
Server-level AccessRights provided to make changes to an application, such as reading log files or removing unnecessary files from application-specific folders.
User TrainingEducation provided to the end user on system usage and functionality.
Portlet Services
ServletA Java Server component responsible for processing requests and responding , typically with HTML. While not tied to any specific protocol, servlets are most often associated with HTTP.
Web ServicesA distributed processing system that involves the exchange of XML data over Internet protocols, typically HTTP/S.
Web-Based UI Services
Publicly AccessibleOnline resources and services available over HTTP or HTTPS over the public internet that are maintained in whole or in part by the Federal Government and operated by an agency, contractor, or other organization on behalf of the agency. They present government information or provide services to the public or a specific user group and support the performance of an agency’s mission. This definition includes all web interactions, whether a visitor is logged-in or anonymous.

Data

TermDefinition
Data Management
Active Data WarehouseWarehouse featuring high-performance transaction processing which supplies data for online processing and updating.
Data LakehouseA modern data architecture that creates a single platform by combining the key benefits of data lakes (large repositories of raw data in its original form) and data warehouses (organized sets of structured data). Specifically, data lake houses enable organizations to use low-cost storage to store large amounts of raw data while providing structure and data management functions.
Enterprise Data MeshA CMS-wide connectivity facility that enables CMS enterprise data to be “shared in place” for consumption within and external to CMS.
Enterprise Data WarehouseWarehouse comprised of any CMS data warehouse or data mart built for decision support.
Data Use Agreement(DUA)A legally binding agreement between CMS and an external entity (e.g., contractor, private industry, academic institution, federal government agency, or state agency), when an external entity requests use of CMS personally identifiable data covered by the Privacy Act of 1974. The agreement delineates the confidentiality requirements of the Privacy Act, security safeguards, and CMS’s data use policies and procedures. The DUA serves as both a means of informing data users of these requirements and a means of obtaining their agreement to abide by these requirements. The DUA also serves as a control mechanism by which CMS can track the location of data and the reason for the data’s release. A DUA requires that a System of Records (SOR) be in effect, which allows for the data to be disclosed.
Information Exchange Agreement (IEA)A legally binding agreement between CMS and an external entity for Business Owners and Privacy Advisors working together to determine the terms of sharing PII with other federal or state agencies.
Interconnection Security Agreement (ISA)A legally binding agreement between CMS and an external entity that defines the relationship between CMS information systems and external systems.
Computer Matching Agreement (CMA)A CMA is created when CMS records are matched with records from another Federal or State agency and the results of such match may have an adverse impact on an individual in relation to a Federal benefit program.
Data Storage Services
Online Storage

Online storage is in constant use in the data center and performs real-time data transactions for applications. Online storage consists of disk drive-based storage that resides in or is attached (direct or fabric) to a server. Direct-attached storage allows only that server attached to the storage to access the storage. Fabric-attached storage enables all servers attached to the fabric to share the available storage resources, such as in a Network Attached Storage (NAS) or Storage Area Network (SAN) configuration. (The “fabric” is the hardware that connects workstations and servers to storage devices in a SAN.)

Online storage devices should allow high-speed access to the storage while at the same time providing data protection and security. High-speed access to the storage is achieved with the high-speed Input/Output (I/O) for the network, system bus, and disk drive interfaces.

Offline Storage

Commonly referred to as “archive” or “backup” storage. Offline storage typically consists of a tape drive or low-end disk drive (virtual tape). Offline storage backs up the data stored on both the online and online archive storage devices. Offline storage is designed store data for long periods.

Because data is archived, offline storage appliances should focus on data accuracy, protection, and security.

Data ArchivingData Archiving (or data migration) is driven by proactive and efficient Information Lifecycle Management. This process involves relocating static, inactive and rarely accessed data from the production environment to a secure, reliable and more cost-effective storage/archival location. Data will be migrated or archived to different storage classes based on such factors as age of data, frequency of use, probability of access, size, usage profile, etc. Data can also be restored from the archival storage location to the production environment as needed.
Direct Access Storage Disk (DASD)Storage that is attached directly to a computer (typically a mainframe computer). This differentiates DASD from storage that is attached to a network, such as Network Attached Storage and Storage Area Networks.
Network Attached Storage (NAS)

Describes a complete storage system that is attached to a traditional data network.

NAS clusters or grids enable scaling of capacity and/or performance into the multi-petabyte (PB) range with bandwidth in the 10’s of GB per second at up to 1,000,000 iops.

NAS consists of one or more file servers. These file servers serve either Windows clients in the form of Common Internet File System (CIFS) or UNIX/Linux clients in the form of Network File System (NFS). Storage on these file servers is provided to the remote system over Ethernet.

In most cases, NAS is less expensive to purchase and less complex to operate than a SAN; however, a SAN can provide better performance and a larger range of configuration options.

Storage Area Network (SAN)

A Storage Area Network is a network specifically dedicated to the task of transporting data for storage and retrieval. SAN architectures are alternatives to storing data on disks directly attached to servers or storing data on Network Attached Storage devices that are connected through general purpose networks.

A SAN is used to access Block Level Data. The remote system accesses a Logical Unit (LUN) as though it were a local disk. The remote system can put a file system on the storage or, in the case of data bases, use it as a raw device.

Storage Area Networks are traditionally connected over Fibre Channel (FC) networks. SANs have also been built using SCSI (Small Computer System Interface) technology. An Ethernet network that is dedicated solely to storage purposes would also qualify as a SAN. Internet Small Computer Systems Interface (iSCSI) is a SCSI variant that encapsulates SCSI data in Transport Control Protocol (TCP) packets and transits them over Internet Protocol (IP) networks.

SAN Connectivity Option and Protocols include Fibre Channel (FC), Fibre Channel over IP (FCIP), Fibre Channel over Ethernet (FCoE), and SCSI over Ethernet. Fibre Channel is the preferred connectivity protocol. Each SAN should consist of two (2) independent fabrics (connections) for high availability.

Disk Storage Virtualization

Disk Storage Virtualization carves up physical disks into smaller chunks that are then used to build traditional RAID (Redundant Array of Independent Disks) constructs. Storage Virtualization abstracts the concept of “Disk” to “Logical Units Numbers” or LUNS. LUNS may be treated exactly the same as Physical Disks. Each LUN is assembled from parts of one or more physical disk and may be arranged as RAID stripes or other logical organization. LUNs may be “Thin Provisioned”.

The key advantage of this technology is that the storage administrator does not have to think about what business applications might be sharing a given set of disks in a RAID set because virtually every RAID set uses some portion of every disk in the array. Disk virtualization also makes it possible to move individual chunks between different tiers of storage within a single array. If the data are then referenced often, the storage administrator can move it dynamically back to the higher performance storage.

The main business driver for virtualization is consolidating multiple storage and services on a single environment. Virtualization reduces costs by sharing hardware, infrastructure, and administration. The benefits of virtualization include:

  • Increased hardware utilization
  • Greater flexibility in resource allocation
  • Reduced power requirements
  • Fewer management costs
  • Lower cost of ownership
  • Administrative and resource boundaries between applications on a system.

There are three (3) Virtualization levels:

  • Fabric Level Virtualization: A SAN fabric enables any-server-to-any-storage device connectivity through the use of protocols such as Fibre Channel switching technology. A “fabric” is the hardware that connects workstations and servers to storage devices in a SAN.
  • SAN fabric is zoned to allow the virtualization appliances to see the storage subsystems and for the servers to see the virtualization appliances. Servers would not be able to directly see or operate on the storage subsystems.
  • Storage Subsystem Level Virtualization: RAID subsystems are an example of virtualization at the storage level.
  • Server Level Virtualization: Abstraction at the server level is by means of the logical volume management of the operating systems on the servers.
Redundant Array of Inexpensive Disks (RAID)

Used to create LUNs that span multiple physical disks. RAID can enhance the reliability of storage compared to single disks, and also decouples the storage from the details of the physical disks.

There are several forms of RAID.

  • RAID 0 = Stripe across several disks. This offers the advantage of creating LUNS larger than the physical disk. It also has the advantage of greater performance than a single disk. The disadvantage of RAID 0 is that a single disk failure corrupts the LUN and it must be rebuilt from backup.
  • RAID 1 = Mirror between pairs of disks. In RAID 1, the same data is stored on two disks. If one disk fails, the other disk still has all of the data. RAID 1 has the disadvantage of requiring twice the amount of physical storage.
  • RAID 5 = Stripe with parity. In RAID 5, the data is striped as in RAID 0, but another disk for parity data is added. In this way, if a single disk fails, the data is not corrupted because it can be recreated from the remaining data and the parity drive. RAID 5 has the disadvantage of slower write speeds due to the parity calculation.
  • RAID 6 = Striped set with dual distributed parity. RAID 6 provides fault tolerance from two drive failures; the array continues to operate with up to two failed drives. This makes larger RAID groups more practical, especially for high-availability systems. This becomes increasingly important because large-capacity drives lengthen the time needed to recover from the failure of a single drive. Single parity RAID levels are vulnerable to data loss until the failed drive is rebuilt: the larger the drive, the longer the rebuild will take. Dual parity gives time to rebuild the array without putting the data at risk if a (single) additional drive were to fail before completing the rebuild.
  • A hybrid RAID called RAID10 combines the performance of RAID0 with the reliability of RAID5.
Thin Provisioning

Thin provisioning allows the storage administrator to over-commit storage on a per-volume basis as long as the amount of data that is actually written does not exhaust the free space in the array or in a particular pool of storage from which the thin-provisioned volumes are backed.

With Thin Provisioning, one can create LUNS that exceed the total amount of physical storage present. The storage appliance only allocates as much physical storage as is used. For instance, the creation of a 500GB LUN only uses 400GB. With thin provisioning, only 400GB of physical storage would be allocated to the LUN. Additional storage would only be added to the LUN if the amount of storage used grew beyond 400GB.

Thin Provisioning should be approached carefully. Heavily over-committing of volumes on a relatively small number of disks can lead to poor performance. The combination of disk spindle virtualization and thin provisioning can mitigate performance concerns.

De-DuplicationDe-duplication reduces the physical storage that is required by identifying duplicate files or blocks. These duplicates are then referenced by a pointer to an original and the space associated with the duplicate is available for other writes. This technology is most common in backup or archiving applications. It may fit into other areas such as server virtualization, where there are many identical blocks or files across multiple operating system instances.
File Transfer
Connect:DirectAn IBM Sterling multi-platform file transfer application capable of transferring files as well as executing job scripts. Connect:Direct also describes the proprietary protocol used to transfer files securely from system to system. Also known as Network Data Mover (NDM)
File Transfer Protocol (FTP)A TCP/IP-based protocol, originally developed for Unix. FTP is a popular but un-secure protocol for transferring files. Alternatives at CMS include S/FTP or FTP/S.
IBM Sterling Business Integration Suite (ISIS)An IBM file transfer management suite used at CMS, capable of communicating with S/FTP, FTP/S, and Connect:Direct protocols.
TIBCO Platform ServerA managed file transfer application used at CMS, able to communicate with S/FTP, FTP/S, and Connect:Direct protocols.
Trigger ScriptsA script executed once the Sweeps system has determined that it needs to be further processed. Scripts can be applications, Job Control Language (JCL), Restructured Extended Executor (REXX), etc.
Analytics and Business Intelligence
Ad Hoc QueriesQueries formulated by users on the spur of the moment and therefore unpredictable in nature.
Dashboards and ScorecardsDashboards are a subset of reporting and include the ability to publish formal, web-based reports with intuitive displays of information, including dials, gauges, sliders, check boxes, and traffic lights. They are designed to deliver historical, current, and predictive information typically represented by key performance indicators (KPI), and they use visual cues to focus user attention on important conditions, trends, and exceptions. Scorecards take the dashboard metrics a step further by applying them to a strategy map that aligns KPIs with a strategic objective. A scorecard implies the use of a performance management methodology such as Balanced Scorecard, Six Sigma, or Capacity Maturity Models.
Data MartA specific subset of a data warehouse for specific analyses needed by a specific group of users.
Data MiningEnables users to conduct exploratory and predictive analytics to extract meaning from seemingly unrelated data by using parameters to search for relationships and patterns.
Data Quality AssuranceProcess of testing data for consistency, completeness, and fitness for publishing to the user community.
Data VisualizationProvides the ability to display numerous aspects of data more efficiently by using interactive pictures and charts instead of rows and columns. The main goal of data visualization is to communicate information clearly and efficiently through graphical means.
Data Warehouse (DW)A collection of data serving as the basis for informational processing and created for decision support. A DW is subject-oriented, integrated, non-volatile, and time-variant.
Enterprise User Administration (EUA)

A CMS system used to manage enterprise userIDs and passwords.

Administrators enter access requests using an EUA workflow system and these requests are forwarded to approvers. Upon approval, the system automatically grants the accesses.

Geographic Information System (GIS)A type of location intelligence revealing trends and patterns using spatial and geographic relationships in the data. For instance, Medicare data at the national, state, zip code, or congressional district level is displayed in a graphical map format with data roll-up and drill-down capabilities integrated into the map itself.
Hybrid Online Analytical Processing (HOLAP)Analytical processing that provides the best of both worlds, using ROLAP and MOLAP techniques. Some HOLAP tools provide capability to drill through from aggregated data stored in multi-dimensional cubes to detailed data stored in relational tables.
Lightweight Directory Access Protocol (LDAP)An application protocol for querying and modifying the data of directory services implemented in Internet Protocol (IP) networks.
Multidimensional Online Analytical Processing (MOLAP)A form of OLAP that works with multidimensional array storage rather than a relational database. MOLAP requires the pre-computation and storage of information in multi-dimensional cubes for slicing and dicing. The data in the cubes is usually aggregated.
Online Analytical Processing (OLAP)While standard and ad hoc query tools are typically used to answer questions like “What happened?” and “When and where did it happen?” OLAP tools are used to answer questions like “Why did it happen?” and to perform “What if?” analysis. Also known as “slicing and dicing” analysis, OLAP allows power users to see facts (numerical, typically additive numbers like claims, payment amounts, and account balances) almost instantaneously regrouped, re-aggregated, and resorted according to any specified dimension (descriptive elements like time, region, claim type, or diagnosis). OLAP methodology can be implemented in ROLAP, MOLAP, or HOLAP depending upon data storage requirements.
Predictive ModelingAnswers questions about what is likely to happen next. Using various statistical models, these tools attempt to predict the likelihood of attaining certain metrics in the future, given various existing and future conditions.
Relational Online Analytical Processing (ROLAP)Works directly with relational databases and stores database and dimension tables as relational tables. This methodology relies on manipulating the data stored in the relational database to produce the appearance of traditional OLAP’s slicing and dicing functionality.
Role-Based Access Control (RBAC)A security approach that restricts system access to authorized users. RBAC is used in BI server administration.
Role-Based AuthorizationProcess in which users are granted access rights based on their primary role (group to which a BI user belongs). In the CMS BI Environment, roles are divided into the following user classifications, depending upon the BI application: standard users, power users, BI developers, and BI analysts.
Standard/Ad Hoc ReportingProvides the ability to create formatted and interactive reports with highly refined scheduling capabilities for offline batch processing. The ad hoc query capability enables users to ask their own questions of a set of data, without relying on IT staff to create a report. Most BI tools have a robust semantic layer enabling users to navigate available data sources. In addition, these tools offer query governance, security, and auditing capabilities to ensure that queries perform well, making only the appropriate data available to users based on user role and access rights.