Autonomous Space Mission Operations: Integrating Fault Management and MBSE for NASA’s HelioSwarm
Fully autonomous space mission operations depend on the ability to detect, diagnose, and mitigate failures without human intervention. NASA is addressing this challenge by integrating fault management (FM) with model-based systems engineering (MBSE). This model-based approach was demonstrated through failure modes and effects analysis (FMEA) and fault tree generation using early design information from NASA’s HelioSwarm mission.
Why Autonomous Fault Management Matters in Space Missions
As NASA advances space exploration through the Artemis program and future deep-space science missions, system autonomy, reliability, and resilience are becoming essential technologies. Autonomous missions require fault management software capable of detecting in-space problems and automatically applying corrective actions without relying on human operators.
Designing an autonomous mission requires a multidisciplinary approach that connects fault management with the model-based systems engineering processes used during mission development. Integrating these disciplines ensures that resilient and fault-tolerant systems are designed, modeled, tested, and incorporated from the earliest stages of the project.
To support this need, NASA awarded a Phase II Small Business Innovation Research (SBIR) contract to Qualtech Systems, Inc. (QSI). The project focused on developing fault management capabilities and enhancing QSI’s commercially available TEAMS® toolset. TEAMS® evolved from earlier NASA-sponsored SBIR commercialization efforts and was enhanced to support HelioSwarm and other NASA heliophysics missions.
Connecting System Health Management with Systems Engineering
A central objective of the project was to connect system health management (SHM) and fault management with the systems engineering process. SHM and FM include the methods and technologies used to prevent failures, detect faults when they occur, diagnose their causes, and mitigate their effects so that mission objectives can still be achieved.
Systems engineering coordinates, integrates, and verifies system elements throughout the design lifecycle. It is essential for defining requirements, developing specifications, and performing verification and validation (V&V). NASA frequently applies a model-based approach to systems engineering, using a systems modeling language such as SysML as a framework for representing system architecture, requirements, behavior, and performance.
Although SHM and FM are closely related to systems engineering, they are often addressed only after the nominal system design has been completed. This approach can make fault management reactive rather than preventive. It may also overlook opportunities to eliminate failure causes or improve system resilience during the design phase.
In addition, systems engineering and SHM/FM activities are frequently managed by separate subject matter experts using different knowledge repositories, modeling techniques, and analysis processes. These disconnected workflows can produce inconsistent results, duplicate effort, and inefficiencies throughout the mission lifecycle.
Integrating Fault Management into MBSE
The NASA-funded QSI team developed an approach that integrates system health management and fault management directly into the MBSE process from the beginning of a project. This method enables engineers to evaluate fault management designs in an operational context by simulating component-level physical and functional failures and demonstrating how different SHM/FM strategies mitigate their effects.
The integrated methodology also supports architecture trade studies. Mission designers can compare alternative fault management concepts during the design phase and select the solutions that provide the best balance of reliability, diagnostic coverage, availability, performance, cost, and complexity.
As part of the SBIR effort, QSI collaborated with the SysML v2 Submission Team (SST). The SST includes end users, software vendors, academic researchers, and government representatives involved in developing the SysML v2 specification. QSI contributed fault management concepts and modeling standards to SysML v2 and demonstrated how SysML v2 models can be translated into fault-space models using the QSI toolset.
FMECA and Fault Tree Analysis for Spacecraft Design
This capability allows systems engineers to analyze fault management characteristics within system designs modeled in SysML v2. By examining the causes and consequences of failures, the QSI toolset supports failure modes, effects, and criticality analysis (FMECA) as well as fault tree analysis (FTA).
These analyses help engineers quantify risks, improve system diagnostics, increase spacecraft availability, and identify opportunities to enhance system resilience. The toolset can also recommend design improvements, such as optimal sensor locations, based on the results of the analysis. Recommendations are delivered in an industry-standard format so they can be reviewed by mission teams and incorporated into the overall system design.
During the SBIR project, the QSI toolset was enhanced to operate within an MBSE framework and support the creation, evaluation, and selection of fault management concepts. Testing fault management strategies early in the design process allows mission teams to incorporate the appropriate sensors, diagnostic functions, redundancy, and mitigation methods before development is complete.
These capabilities can reduce total development costs by improving communication and coordination among mission team members, identifying design issues earlier, and reducing technical, cost, and schedule risks.
NASA’s HelioSwarm Mission
NASA’s HelioSwarm mission is designed to advance understanding of turbulence in the solar wind and the Sun-Earth system. The mission will use a constellation, or “swarm,” of eight small satellites operating with a single hub spacecraft. Together, the spacecraft will make the first simultaneous, multiscale measurements of magnetic-field fluctuations and proton flux in the dynamic cislunar environment.
Plasma turbulence transfers energy across multiple scales, from fluid-scale motion to the kinetic behavior of individual particles. Because turbulence cannot be fully understood through measurements taken at a single point or scale, HelioSwarm’s spacecraft will operate at distances ranging from tens to thousands of kilometers apart.
This distributed spacecraft configuration will allow scientists to reconstruct the three-dimensional structure and dynamics of turbulent space plasma. The resulting observations may reveal how energy moves through the solar wind and improve scientific understanding of fundamental plasma processes near Earth, around the Sun, and throughout the universe.
Plasma turbulence occurs when energy contained in fluctuating magnetic fields and plasma motion cascades from larger spatial scales to smaller ones. As the cascade reaches kinetic scales associated with dissipation, energy is transferred into particle heat. Turbulence is therefore a fundamental thermodynamic process in cosmic plasmas and remains one of the most important unsolved problems in classical physics.
Applying MBSE and Fault Analysis to HelioSwarm
The QSI team created a SysML v2 design model of the HelioSwarm subsystems and top-level mission requirements. This model helped demonstrate how mission objectives flow down into system and subsystem designs.
The enhanced QSI toolset was then used to convert the HelioSwarm SysML v2 model into a fault management model. The model represented the primary subsystems of the hub spacecraft and eight node satellites, including command and data processing, power, attitude control, propulsion, thermal management, isolation hardware, payload sensors, and communication systems connecting ground and space assets.
Mission designers used the toolset to generate FMECA and fault tree analyses and convert the results into standardized SysML reports. The fault management analysis also generated design recommendations, including potential sensor placement improvements, which could be fed back into the system design.
This integrated process supports the development of small spacecraft constellations with built-in redundancy and improved fault tolerance. Such capabilities can enhance scientific observations and support other NASA objectives, including future mission operations around the Moon and lunar surface activities.
Future Applications for Autonomous Spacecraft
The technology developed through this SBIR project could provide significant value for future NASA missions that require autonomous operations, high availability, and increased resilience. The QSI TEAMS® toolset was baselined for vehicle systems management functions associated with NASA’s Gateway project and remains applicable to future crewed and uncrewed spacecraft.
By integrating fault mitigation strategies into the earliest stages of system design, engineers can gain deeper insight into overall system resilience and identify improvements before hardware and software development begin. This approach can help reduce redesign costs, improve mission reliability, and increase confidence in autonomous spacecraft operations.
The technology also has applications beyond NASA. Comprehensive fault management analysis and architectural trade studies are important for complex, high-value systems, including military aircraft, surface ships, submarines, and modern ground combat vehicles. Additional applications may include commercial spacecraft, civil aviation, maritime systems, transportation networks, and power generation and distribution equipment.
For more information about this initiative, visit the related NASA TechPort project entries: Project 154400, Project 154638, Project 158021, and Project 157599.
Project leaders: Dr. Sudipto Ghoshal and Mr. Deepak Haste, Qualtech Systems, Inc.
Sponsoring organization: NASA Ames Research Center
Source: science.nasa.gov


