Category: Safety Analysis

  • System Requirements Hazard Analysis

    System Requirements Hazard Analysis

    In this 45-minute session, I’m looking at System Requirements Hazard Analysis, or SRHA, which is Task 203 in the Mil-Std-882E standard. I will explore Task 203’s aim, description, scope, and contracting requirements.  SRHA is an important and complex task, which must be done on several levels to succeed.  This video explains the issues and discusses how to perform SRHA well.

    This is the seven-minute demo video, the full version is 40 minutes’ long.

    Topics: System Requirements Hazard Analysis

    • Task 202 Purpose;
    • Task Description:
      • Determine Requirements;
      • Incorporate Requirements; and
      • Assess the compliance of the System.
    • Contracting;
    • Section 4.2 (of the standard); and
    • Commentary.

    Transcript

    Introduction

    Hello and welcome to the Safety Artisan, where you will find professional, pragmatic and impartial advice on all things system, safety and related.

    System Requirements Hazard Analysis

    Today, we’re talking about system requirements hazard analysis. And this is part of our series on Mil. Standard 882E, and this one is Task 203. And it’s a very widely used system safety engineering standard. Its influence is found in many places, not just in military procurement programs.

    Topics for this Session

    We’re looking at this task, which is very important, possibly the most important task of all, as we’ll see. I’m talking about the purpose of the task, which is word-for-word from the task description itself.

    We’re talking about in the task description, the three aims of this task, which is to determine or work out requirements, incorporate them, and then assess the compliance of the system with those requirements, because, of course, it may not be a simple read-across. We’ve got six slides on that. That’s most of the task.

    Then we’ve just got one slide on contracting, which if you’ve seen any of the others in this series, will seem very familiar. We’ve got a bit of a chat about Section 4.2 from the standard and some commentary, and the reason for that will become clear. Let’s crack on!

    System Requirements Hazard Analysis

    Task 203.1, the purpose of Task 203 is to perform and document a System Requirements Hazard Analysis or SRHA. And as we’ve already said, the purpose of this is to determine the design requirements. We’re going to focus on design rather than buying stuff off the shelf – we’ll talk about the implications of that a little bit later.

    Design requirements to eliminate or reduce hazards and risks, incorporate those requirements, into a says, into the documentation, but what it should say is incorporate risk reduction measures into the system itself and then document it.

    Finally, to assess compliance of the system with these requirements. Then it says the SRHA address addresses all life-cycle phases, so not just meant for you to think about certain phases of the program. What are the requirements through life for the system? And in all modes. Whether it’s in operation, whether it’s in maintenance or refit, whether it’s being repaired or disposed of, whatever it might be.

    Task Description #1

    The first of six slides is the task description. I’m using more than one colour because there’s some quite a lot of important points packed quite tightly together in this description.

    We’re assuming that the contractor performs and documents this SRHA. The customer needs to do a lot of work here before ever gets near a contractor. More on that later. We need to determine system design requirements to eliminate hazards or reduce associated risks.

    Two things here. By identifying applicable policies, regulations, standards, etc. More on that later. And analyzing identified hazards. So, requirements to perform the analysis as well as to simply just state ‘We want a system to do this and not to do that’. So, we need to put some requirements to say ‘Here’s what we want analyzed maybe to what degree? And why.’ is always helpful.

    Task Description #2

    Breaking those breaking those two requirements down.

    Part a. We identify applicable requirements by reviewing our military and industry standards and specs, and historical documentation of systems that are similar or with a system that we’re replacing, perhaps. It’s assumed that the US Department of Defense is the customer, the ultimate customer. So, the ultimate customer’s requirements, including whatever they’ve said about standard ways of mitigating certain common risks.

    The system performance spec, that’s your functional performance spec or whatever you want to call it. Other system design requirements and documents – a bit of a catchall there. And applicable federal, military, state, and local regulations.

    This is a US standard. It’s a federated state, much like Australia and lots of modern states, even the UK. There are variations in law across England, Wales, Scotland and Ireland. They’re not great, but they do exist.

    And in the US and Australia, those differences are greater. And it says applicable executive orders. Executive orders, they’re not law, but they are what the executive arm of the U.S. government has issued, and international agreements. There are a lot of words in there – have a look at the different statements that are in white, blue, and yellow.

    Basically, from international agreements right down to whatever requirements may be applicable, they all need to be looked at and accounted for. So, there’s a huge amount of work there for someone to do. I’ll come back to who that someone should be later.

    End: System Requirements Hazard Analysis

    You can find a free pdf of the System Safety Engineering Standard, Mil-Std-882E, here.

    Meet the Author

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • Failure Mode Effects Analysis

    Failure Mode Effects Analysis

    TL;DR: This article on Failure Mode and Effects Analysis explains this powerful and commonly-used family of techniques. You can access this webinar (and all the others) here.

    I have used FMEA and related techniques on many programs, and it can produce powerful results quickly and cheaply. Recently, I’ve seen some criticism of FMEA on social media. However, I’m convinced that this is only clickbait. The secret of success is to understand what a technique is good for – and not – and to apply it well. It’s as simple as that!

    This article covers:

    • A description of the technique, including its purpose;
    • When it might be used;
    • Advantages, disadvantages and limitations;
    • Sources of additional information;
    • A simple example of an FMEA/FMECA; and
    • Additional comments.

    I’ve added some ‘top tips’ of my own based on my personal experience in the industry.

    Top Tip

    In this article, I have used material from a UK Ministry of Defence guide, reproduced under the terms of the UK’s Open Government Licence. I have rewritten the very dull source material to make it more readable!

    A Description of the Technique, Including Its Purpose

    Failure modes and effects analysis (FMEA) was one of the first systematic techniques for failure analysis. It was developed in the United States military (Military Procedure MIL-P-1629, titled ‘Procedures for Performing a Failure Modes, Effects and Criticality Analysis’, November 9, 1949) as a reliability evaluation technique to determine the effect of system and equipment failures. Failures were classified according to their impact on mission success and personnel, equipment, and safety. In the 1960s, it was used by the aerospace industry and NASA during the Apollo program. More and more industries – notably the automotive industry – have seen the benefits to be gained by using FMEAs to complement their design processes.

    This qualitative technique helps identify failure potential in a design or process i.e., to foresee failure before it actually happens. This is done by defining the system that is under consideration to ensure system boundaries are established, and then by following a procedure, which helps to identify design features or process operations that could fail. The procedure requires the following essential questions to be asked:

    • How can each component fail?
    • What might cause these modes of failure?
    • What could the effects be if these failures did occur?
    • How serious are these failure modes?
    • How is each failure mode detected?
    • What are the safeguards in place to protect against accidents resulting from the failure mode?

    As always with safety analyses, the more precisely you can answer these questions (above), the better the results you will get.

    Top Tip

    As an aid in structuring the analysis and ensuring a systematic approach, results are recorded in a tabular format. Several different forms are in use, and the form design can be tailor-made to suit the particular requirements of a study. Examples of forms can be found in several standards (links below).

    Make the form support the flow of the process, left-to-right, then top-down!

    Top Tip

    The FMEA analysis can be extended if necessary by characterizing the likelihood, severity, and resulting levels of risk of failures. FMEAs that incorporate this criticality analysis (CA) are known as FMECAs. A FMECA is an analytical quantitative technique that ranks failure modes according to their probability and consequences (i.e., the resulting effect of the failure mode on the system, mission, and personnel). This technique is referred to as a “bottom-up approach” as it starts by identifying the potential failure modes of a component and analyzing their effects on the whole system. It can be quite complex depending on how the user drives the technique.

    We should note that the FMECA does not provide a model by which system reliability can be quantified. Hence, if the objective is to estimate the probability of events, a technique that results in a logic model of the failure mechanisms must be employed, typically a fault tree and/or an event tree.

    Reliability Block Diagrams, or for repairable systems, Markov Chains can also be used.

    Top Tip

    A FMEA or FMECA can be conducted on either a component or a functional level. A functional FMEA/FMECA only covers hardware aspects, but a functional FMEA/FMECA can cover all aspects of a system. For either approach, the general principle remains the same.

    When it Might be Used

    FMEA is applicable for any well-defined system, but is primarily used for reviews of mechanical and electrical systems. It can be used in many situations, for example, to assess the design of a product in terms of what could go wrong in manufacturing and in-service as a result of the weakness in the design. We can also use it to analyze failures in the manufacturing process itself and during service. It is effective for collecting information needed to troubleshoot system problems and improving maintenance and reliability of plant and equipment (defining and optimizing), as it focuses directly and individually on equipment failure modes.

    It’s fair to say that you need a design, on which to perform a FMEA. Pre-design you could use Functional Failure Analysis (FFA) instead.

    Top Tip

    The FMECA technique is best suited for detailed analysis of system hardware, and should preferably be carried out by the designer in parallel with system development. This will not only speed up the analysis itself, but also force the design team to think systematically about the failure characteristics of the system. The primary use of the FMECA is in verifying that single-component failures cannot cause a catastrophic system failure.

    There are a number of areas today in which the use of FMECA has become mandatory to demonstrate system reliability. Examples of such requirements are in the classification of Dynamically Positioned (DP) vessels and in a number of US military applications for which MIL-STD documents apply.

    Advantages, Disadvantages, and Limitations

    Advantages

    • It is widely used and well-understood, and easy to understand and interpret
    • It can be performed by a single analyst or more if required
    • Qualitative data about the causes and effects can be incorporated into the analysis
    • It is systematic and comprehensive, and should identify hazards with an electrical or mechanical basis
    • The level of detail incorporated can be varied to suit the analysis
    • It identifies safety-critical equipment where a single failure would be critical for the system
    • Even though the technique can be quite time-consuming it can lead to a thorough understanding of the system being considered

    Disadvantages

    • The technique adopts a bottom-up approach, and if conducting a component-level FMEA or FMECA, this can be boring and repetitive
    • The benefit gained is dependent upon the experience of the analyst or the group.
    • It requires a hierarchical system drawing as the basis for the analysis, which the analyst usually has to develop before the FMEA process can start
    • It is optimised for mechanical and electrical equipment, and does not apply easily to Human Factor Integration, procedures, or process equipment
    • It is difficult for the technique to cover multiple failures, as equipment failures are generally analysed one by one; therefore, important combinations of equipment failures may be overlooked
    • Most accidents have a significant human or external influence contribution, and these are not a usual failure mode with FMEA
    • More than one FMEA may be required for a system with multiple modes of operation
    • Due to its wide use, there can be a temptation to read across data from ARM or ILS projects where, for example, the fault-tree technique has been used. As a consequence, the safety perspective can be lost as human error has been excluded and the focus has been solely on determining faults and not on more far-reaching safety issues
    • Perhaps the worst drawback of the technique is that all component failures are examined and documented, including those that do not have any significant consequences.
    • For large systems, especially those with a fair degree of redundancy built into them, the amount of unnecessary documentation is a major disadvantage. Hence, the FMECA should primarily be used by designers of reasonably simple systems. It should, however, be noted that the concept of the FMECA form can be quite useful in other contexts, e.g., when reviewing an operation rather than a hardware system. Then the use of a form similar to the FMECA can provide a useful way of documenting the analysis. Suitable columns in the form could, for example, include: operation, deviation, consequence, correcting or reversing action, etc.

    ARM = Availability, Reliability, Maintainability
    ILS = Integrated Logistic Support (or logistics engineering
    )

    Top Tip

    Sources of Additional Information, such as Standards, Textbooks, and Websites

    BS 5760: Part 5 Reliability of Systems, Equipment and Components: Part 5 Guide to Failure Modes, Effects and Criticality Analysis.

    HSE Website – Marine Risk Assessment, Offshore Technology Report 2001/063

    IEC 60812:2018 Failure modes and effects analysis (FMEA and FMECA)

    As always, Understand your Standard (what it was designed to do) to get the best out of it!

    Top Tip

    A Simple Example of an FMEA/FMECA

    An example extract from an FMEA of a ballast system is shown below. This can be found in the HSE Marine Risk Assessment Report. The column headings are based on the US Military Standard Mil-Std 1629A, but with modifications to suit the particular application. For example, the failure mode and cause columns are combined. The criticality of each failure is ranked as minor, incipient, degraded, or critical.

    An example of an FMEA Output Table

    To properly understand these results you need to know how a Sea Chest works (see context here). Otherwise the example just shows what kind of output a FMEA can produce.

    Top Tip

    Additional comments

    Failure Modes and Effects and Criticality Analysis (FMECA) is an analytical QRA technique, used by ARM and ILS systems engineers, most commonly and effectively at the late design, test and manufacture stage of a project. It requires the breakdown of the system into individual components and the identification of possible failure modes or malfunctions of each component, (such as too much flow through a valve). Referred to as a bottom-up approach, it starts by identifying the potential failure modes of a component and analyzing their potential effects on the whole system. Numerical levels can be assigned to the likelihood of the failure and the severity or consequence of the failure.

    Note: It is important to recognize that FMEA/FMECA Standards have different approaches to criticality. Failure mode severity classes 1 – 5 for Standards MIL1629A and ARP926A go from Class 1 being the most severe (e.g. loss of life) to Class 5 being less severe (i.e. no effect), whereas BS 5760 deals with criticality in the opposite direction where Class 5 is the most severe.

    Note that FMECA for ARM/ILS looks at availability or mission criticality, not safety criticality.  A FMECA for safety will have a different focus.

    Top Tip

    Software:

    • Isograph;
    • Reliability Work Bench;
    • Reliasoft;
    • Microsoft Excel.

    These are not recommendations!

    FMEA/FMECA tables for complex systems can run to hundreds of pages, so good tool support is essential.

    Top Tip

    Failure Mode Effects Analysis: Have You Used This Technique?

    Back to the Safety Assessment topic page.

    Meet the Author

    Learn safety engineering with me, an industry professional with 25 years of experience. I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • Five Ways to Identify Hazards

    Five Ways to Identify Hazards

    In my webinar ‘Five Ways to Identify Hazards’ I look at a mix of techniques. We need these diverse techniques to assure us (give justified confidence) that we have identified the full range of hazards associated with a system.

    To do this I draw on my 25 years of experience (see ‘Meet the Author‘, below) and relevant standards. Here’s the introduction to the webinar.

    Five Ways to Identify Hazards: Video Introduction

    Webinar: ‘Five Ways to Identify Hazards’

    Four Things to Remember

    For hazard identification, we need to be aware of four things.

    What we’re doing is we are imagining what could go wrong. And I want to emphasize, first of all, imagination. We need to be open to what could happen. That’s the mindset that we need, and we’re looking at what could go wrong, not what will go wrong. Think about possibilities, not certainties.

    The second thing is that it’s very easy to dive down a rabbit hole and get into mega detail about one particular thing and spend lots of time, waste lots of time doing that. That’s not what we need to be doing. We need a broad approach. We need to go wide and think about as many different possible hazards as we can. Don’t dive deep that will come later, the deep analysis will come later.

    Another aspect of that point is we’re talking about hazard identification. We’re just here to identify hazards. We’re not here to try to assess them yet.

    Yet another mistake that people make is to try and jump straight to fixing the hazard. Many of us watching will be engineers. We love fixing problems. We like to solve problems, but we’re not here to solve the problem yet. We’re only here to identify it. So we’re going to avoid the temptation to jump in and try and come up with a solution. That’s not what we’re doing with hazard identification.

    So those are four things to bear in mind.

    Five Ways to Identify Hazards

    Let’s move on. So I’ve said that this was entitled five ways to identify hazards.

    There are, of course, many ways to identify hazards, but I just thought I’d pick on these five because there was a nice broad range of things and things that I can show you how to do straight away.

    Those are the five things that we’ve got and we’ll have a slide on each one of those. First, we can ask the workers or end users or their representatives. Secondly, we can inspect the workplace, we can look around for hazards. And maybe we’ve got a real workplace that we can look at or maybe we’ve just got a representation, we can do both.

    We can use a hazard identification checklist, we can survey historical data. So all the squiggly lines at the bottom of the screen, there’s an example of some historical data and we can conduct a number of analyses on that.

    But the analysis I picked on (Number 5) is Functional Failure Analysis and we’ll see why in just a moment. So those are the five things that we will cover in the next hour. We’ll also have time for a Question and Answer session and then a worked example of how to do a simple Functional Failure Analysis…

    There’s More!

    This is just one of many webinars in my Safety Engineering Academy. You can see summaries of them all in this blog post.

    Meet the Author

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • Exploring Causal Analysis: Techniques and Insights

    Exploring Causal Analysis: Techniques and Insights

    In this post, ‘Exploring Causal Analysis: Techniques and Insights’, I provide a quick summary of my recent webinar. You can see a short video introduction below, or access the full webinar at my Safety Engineering Academy.

    Introduction:

    Causal analysis is a vital aspect of system safety engineering, offering insights into the root causes of issues and guiding effective problem-solving strategies. In this webinar, we delve into various causal analysis techniques and discuss their practical applications in diverse domains.

    Section 1: Introduction to Causal Analysis

    Causal analysis involves understanding the sequence of events leading to an outcome and identifying the underlying factors contributing to it. We explore the fundamentals of causal analysis and its significance in safety engineering.

    Section 2: Popular Causal Analysis Techniques

    We examine eight popular causal analysis methods, including Pareto charts, Failure Mode and Effects Analysis (FMEA), 5 Whys, Ishikawa diagrams, Fault Tree Analysis (FTA), 8D reporting, DMAIC, and Scatter Diagrams. Each technique is analyzed for its strengths, limitations, and practical utility.

    Section 3: Deeper Dive into Selected Techniques

    We take a closer look at selected causal analysis techniques, exploring their application in real-world scenarios. Examples include using Pareto charts to identify dominant causes, leveraging FMEA for failure mode analysis, and utilizing Fault Tree Analysis for assessing complex system failures.

    Section 4: Insights and Reflections

    Drawing from years of experience in system safety engineering across diverse domains and international contexts, we share insights and reflections on the effectiveness of different causal analysis techniques. We emphasize the importance of choosing the right technique based on the specific objectives and available data.

    Section 5: Resources and Next Steps

    We provide attendees with valuable resources for further exploration of causal analysis techniques, including links to webinars, online courses, and templates. Additionally, we offer a glimpse into ongoing research and developments in the field of safety engineering.

    Conclusion:

    Causal analysis is a dynamic and evolving field that plays a crucial role in ensuring system reliability and safety. By employing a range of techniques and approaches, safety practitioners can gain deeper insights into the root causes of issues and implement effective risk management strategies.

    Bonus: Q&A Sessions

    Here are the two Q&A Sessions from the Webinar:

    Causal Analysis – Q&A Session 1

    Exploring Causal Analysis: Techniques and Insights – Get more here

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • Functional Hazard Analysis with Mil-Std-882E

    Functional Hazard Analysis with Mil-Std-882E

    In this video, I look at Functional Hazard Analysis with Mil-Std-882E (FHA, which is Task 208 in Mil-Std-882E). FHA analyses software, complex electronic hardware, and human interactions. I explore the aim, description, and contracting requirements of this Task, and provide extensive commentary on it. (I refer to other lessons for special techniques for software safety and Human Factors.)

    This video, and the related webinar ‘Identify & Analyze Functional Hazards’, deal with an important topic. Programmable electronics and software now run so much of our modern world. They control many safety-related products and services. If they go wrong, they can hurt people.

    I’ve been working with software-intensive systems since 1994. Functional hazards are often misunderstood or overlooked, as they are hidden. However, the accidents that they can cause are very real. If you want to expand your analysis skills beyond just physical hazards, I will show you how.

    This is the seven-minute demo; the full version is 40 minutes long.

    Functional Hazard Analysis: Context

    So how do we analyze software safety?

    Before we even start, we need to identify those system functions that may impact safety. We can do this by performing a Functional Failure Analysis (FFA) of all system requirements that might credibly lead to human harm.

    An FFA looks at functional requirements (the system should do ‘this’ or ‘that’) and examines what could go wrong:

    • Does the function work when needed?
    • Does the function work when not required?
    • Does the function work incorrectly? (There may be more than one version of this.)

    (A variation of this technique is explained here.)

    If the function could lead to a hazard then it is marked for further analysis. This is where we apply the FHA, Task 208.

    Functional Hazard Analysis: The Lesson

    Topics: Functional Hazard Analysis

    • Task 208 Purpose;
    • Task Description;
    • Update & Reporting
    • Contracting; and
    • Commentary.

    Transcript: Functional Hazard Analysis

    Introduction

    Hello, everyone, and welcome to the Safety Artisan; Home of Safety Engineering Training. I’m Simon and today we’re going to be looking at how you analyze the safety of functions of complex hardware and software. We’ll see what that’s all about in just a second.

    Functional Hazard Analysis

    I’m just going to get to the right page. This, as you can see, functional hazard analysis is Task 208 in Mil. Standard 882E.

    Topics for this Session

    What we’ve got for today: we have three slides on the purpose of functional hazard analysis, and these are all taken from the standard. We’ve got six slides of task description. That’s the text from the standard plus we’ve got two tables that show you how it’s done from another part of the standard, not from Task 208. Then we’ve got update and recording, another two slides. Contracting, two slides. And five slides of commentary, which again include a couple of tables to illustrate what we’re talking about.

    Functional Purpose HA #1

    What we’re going to talk about is, as I say, functional hazard analysis. So, first of all, what’s the purpose of it? In classic 882 style, Task 208 is to perform this functional hazard analysis on a system or subsystem or more than one. Again, as with all the other tasks, we use it to identify and classify system functions and the safety consequences of functional failure or malfunction. In other words, hazards.

    Now, I should point out at this stage that the standard is focused on malfunctions of the system. In the real world, lots of software-intensive systems cause accidents that have killed people, even when they’re functioning as intended. That’s one of the shortcomings of this Military Standard – it focuses on failure. But even if something performs as specified, either:

    • The specification might be wrong, or
    • The system might do something that the human operator does not expects.

    Mil-Std-882E just doesn’t recognize that. So, it’s not very good in that respect. However, bearing that in mind, let’s carry on with looking at the task.

    Functional HA Purpose #2

    We’re going to look at these consequences in terms of severity – severity only, we’ll come back to that – to identify what they call safety-critical functions, safety-critical items, safety-related functions, and safety-related items. And a quick word on that, I hate the term ‘safety-critical’ because it suggests a sort of binary “Either it’s safety-critical. Yes. Or it’s not safety-critical. No.” And lots of people take that to mean if it’s “safety-critical, no,” then it’s got nothing to do with safety. They don’t recognize that there’s a sliding scale between maximum safety criticality and none whatsoever. And that’s led to a lot of bad thinking and bad behavior over the years where people do everything they can to pretend that something isn’t safety-related by saying, “Oh, it’s not safety-critical, therefore we don’t have to do anything.” And that kind of laziness kills people.

    Anyway, moving on. So, we’ve got these SCFs, SCIs, SRFs, SRIs and they’re supposed to be allocated or mapped to a system design architecture. The presumption in this – the assumption in this task is that we’re doing early – We’ll see that later – and that system design, system architecture, is still up for grabs. We can still influence it.

    COTS and MOTS Software

    Often that is not the case these days. This standard was written many years ago when the military used to buy loads of bespoke equipment and have it all developed from new. That doesn’t happen anymore so much in the military and it certainly doesn’t happen in many other walks of life – But we’ll talk about how you deal with the realities later.

    And they’re allocating these functions and these items of interest to hardware, software, and human interfaces. And I should point out, when we’re talking about all that, all these things are complex. Software is complex, human is complex, and we’re talking about complex hardware. So, we’re talking about components where you can’t just say, “Oh, it’s got a reliability of X, and that’s how often it goes wrong” because those types of simple components are only really subject to random failure, that’s not what we’re talking about here.

    We’re talking about complex stuff where we’re talking about systematic failure dominating over random, simple hardware failure. So, that’s the focus of this task and what we’re talking about. That’s not explained in the standard, but that’s what’s going on.

    Functional HA Purpose #3

    Now, our third slide is on purpose; so, we use the FHA to identify the consequences of malfunction, functional failure, or lack of function. As I said just now, we need to do this as early as possible in the systems engineering process to enable us to influence the design. Of course, this is assuming that there is a system engineering process – that’s not always the case. We’ll talk about that at the end as well.

    Also, we’re going to identify and document these functions and items and allocate and it says to partition them in the software design architecture. When we say partition, that’s jargon for separating them into independent functions. We’ll see the value of that later on. Then we’re going to identify requirements and constraints to put on the design team to say, “To achieve this allocation in this partitioning, this is what you must do and this is what you must not do”. So again, the assumption is we’re doing this early. There’s a significant amount of bespoke design yet to be done….

    Then What?

    Once the FFA has identified the required ‘Level or Rigor’, we need to translate that into a suitable software development standard. This might be:

    • RTCA DO-178C (also know as ED-12C) for civil aviation;
    • The US Joint Software System Safety Engineering Handbook (JSSEH) for military systems;
    • IEC 61508 (functional safety) for the process industry;
    • CENELEC-50128 for the rail industry; and
    • ISO 26262 for automotive applications.

    Such standards use Safety Integrity Levels (SILs) or Development Assurance Levels (DALs) to enforce appropriate Levels of Rigor. You can learn about those in my course, Principles of Safe Software Development.

    Meet the Author

    My name’s Simon Di Nucci. I’m a practicing system safety engineer, and I have been, for the last 25 years; I’ve worked in all kinds of domains, aircraft, ships, submarines, sensors, and command and control systems, and some work on rail air traffic management systems, and lots of software safety. So, I’ve done a lot of different things!

  • How to do Preliminary Hazard Analysis with Mil-Std-882E

    How to do Preliminary Hazard Analysis with Mil-Std-882E

    In this 45-minute session, I look at how to do a Preliminary Hazard Analysis with Mil-Std-882E. Preliminary Hazard Analysis, or PHA, is Task 202 in the Standard.

    I explore Task 202’s aim, description, scope, and contracting requirements. There’s value-adding commentary, and I explain the issues with PHA – how to do it well and avoid the pitfalls.

    Now, I have worked in System Safety since 1996, and I think that PHA is one of the three tasks that EVERY project needs to do. The other two are:

    I look at these three tasks together in my course ‘Foundations of Safety Assessment’. This is one of five linked courses on Mil-Std-882E. They will teach you how to get the maximum benefits from your System Safety Program.

    This is the seven-minute-long demo video. The full video is 45 minutes long.

    Topics: How to do Preliminary Hazard Analysis

    • Task 202 Purpose;
    • Task Description;
    • Recording & Scope;
    • Risk Assessment (Tables I, II & III);
    • Risk Mitigation (order of preference);
    • Contracting; and
    • Commentary.

    Transcript: How to do Preliminary Hazard Analysis

    Hello and welcome to the Safety Artisan, where you’ll find professional, pragmatic, and impartial safety training resources. So, we’ll get straight on to our session, which is on the 8th of February 2020.

    Preliminary Hazard Analysis

    Now we’re going to talk today about Preliminary Hazard Analysis (PHA). This is Task 202 in Military Standard 882E, which is a system safety engineering standard. It’s very widely used mostly on military equipment, but it does turn up elsewhere.  This standard is of wide interest to people and Task 202 is the second of the analysis tasks. It’s one of the first things that you will do on a systems safety program and therefore one of the most informative. This session forms part of a series of lessons that I’m doing on Mil-Std-882E.

    Topics for This Session

    What are we going to cover in this session? Quite a lot! The purpose of the task, a task description, recording, and scope. How we do risk assessments against Tables 1, 2, and 3. These tables describe severities, likelihoods, and the overall risk matrix.  We will talk about all three, about risk mitigation and using the order of preference for risk mitigation, a little bit of contracting, and then a short commentary from myself. I’m providing commentary all the way through. So, let’s crack on.

    Task 202 Purpose

    The purpose of Task 202, as it says, is to perform and document a preliminary hazard analysis, or PHA for short, to identify hazards, assess the initial risks, and identify potential mitigation measures. We’re going to talk about all of that.

    Task Description

    First, the task description is quite long here. And as you can see, I’ve highlighted some stuff that I particularly want to talk about.

    It says “the contractor” [does this or that], but it doesn’t matter who is doing the analysis, and, the customer needs to do something to inform themselves, otherwise they won’t understand what they’re doing.  Whoever does it needs to perform and document a PHA. It’s about determining initial risk assessments. There’s going to be more work, more detailed work done later. But for now, we’re doing an initial risk assessment of identified hazards. And those hazards will be associated with the design or the functions we propose to introduce. That’s very important. We don’t need a design to do this. We can get in early when we have user requirements, functional requirements, and that kind of thing.

    Doing this work will help us make better requirements for the system. So, we need to evaluate those hazards for severity and probability. It says based on the best available data. And of course, early in a program, that’s another big issue. We’ll talk about that more later. It says to include mishap data as well, if accessible: American term mishap, means an accident, but we’re avoiding any kind of suggestion about whether it is accidental or deliberate.  It might be stupidity, deliberate, or whatever. It’s a mishap. It’s an undesirable event.

    We look for accessible data from similar systems, legacy systems, and other lessons learned. I’ve talked about that a little bit in the Task 201 lesson, and there’s more on that today under commentary. We need to look at provisions, and alternatives, meaning design provisions and design alternatives to reduce risks and add mitigation measures to eliminate hazards. If we can all reduce associated risk, we need to include all of that. What’s the task description? That’s a good overview of the task and what we need to talk about.

    Reading & Scope

    First, recording and scope, as always, with these tasks, we’ve got to document the results of the PHA in a hazard tracking system. Now, a word on terminology; we might call a hazard tracking system; we might call it a hazard log; we might call it a risk register. It doesn’t matter what it’s called. The key point is it’s a tracking system. It’s a live document, as people say, it’s a spreadsheet or a database, something like that. It’s something relatively easy to update and change.

    Also, we can track changes through the safety program once we do more analysis because things will change. We should expect to get some results and refine them and change them as time goes on. Very important point.

    That’s it for the Demo…

    End: How to do Preliminary Hazard Analysis

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

    You can find a free pdf of the System Safety Engineering Standard, Mil-Std-882E, here.

  • Preliminary Hazard Identification & Analysis Guide

    Preliminary Hazard Identification & Analysis Guide

    Get your free Preliminary Hazard Identification & Analysis, PHIA Guide here!

    Introduction

    Hazard Identification is sometimes defined as: “The process of identifying and listing the hazards and accidents associated with a system.”

    Hazard Analysis is sometimes defined as: “The process of describing in detail the hazards and accidents associated with a system and defining accident sequences.”

    Preliminary Hazard Identification and Analysis (PHIA) helps you determine the scope of safety activities and requirements. You can identify the main hazards likely to arise from the capability and functionality being provided. Perform it as early as possible in the project life cycle. Thus, you will provide important early input to setting Safety requirements and refining the Project Safety Plan.

    PHIA seeks to answer, at an early stage of the project, the question: “What Hazards and Accidents might affect this system and how could they happen?”

    Aim

    The PHIA aims to identify, as early as possible, the main Hazards and Accidents that may arise during the life of the system. It provides input to:

    1. Scoping the subsequent Safety activities required in any Safety Plan. A successful PHIA will help to gauge the proportionate effort that is likely to be required to produce an effective Safety Case, proportionate to risks.
    2. Selecting or eliminating options for subsequent assessment.
    3. Setting the initial Safety requirements and criteria.
    4. Subsequent Hazard Analyses.
    5. Initiate Hazard Log.

    Description

    Perform a PHIA as early as possible to obtain maximum benefit. Use it to understand what the Hazards and Accidents are, why, and how they might be realized. A PHIA is an important part of Risk Management, project planning, and requirements definition. It helps you to identify the main system hazards and helps target where a more thorough analysis should be undertaken.

    Usually, PHIA is based on a structured brainstorming exercise, supported by hazard checklists. A structured approach helps to minimize the possibility of missing an important hazard. It also demonstrates that a comprehensive approach has been applied.

    Get Your PHIA Guide as part of the FREE Learning Bundle

    Front cover of PHIA Guide
    Subscribe to The Safety Artisan Mailing List and get your Free Gift!

    Find more on basic safety topics at Start Here.

    Meet the Author

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • Foundations of Safety Assessment

    Foundations of Safety Assessment

    In this post on the Foundations of Safety Assessment, I’m going to look at the (few) things that we need to do in every System Safety Program.

    Because we don’t always need to do everything. We don’t always need to throw everything at the problem. Some systems are simpler than others, and they don’t need the ‘whole nine yards’ in order to get a decent result. With that knowledge, we’re going to be able to design an analysis program for different applications or for different systems.

    As an example, I’m going to use Military Standard 882E (Mil-Std-882E). Under that standard we would use these Tasks:

    • Task 201 – Preliminary Hazard Identification;
    • Task 202 – Preliminary Hazard Analysis; and
    • Task 203 – System Requirements Hazard Analysis.

    (You will also find related material in my posts on Safety Analysis Techniques Overview and tailoring your Risk Analysis Program.)

    Foundations of Safety Assessment – The Big Picture

    I promised you we were going to look at the overview of the sequence.

    And I think this is what pulls it all together and explains it powerfully. So the background to this is we’ve got, an accident or mishap sequence. Whatever you want to call it and we start with causes on the left and causes lead two a hazard, and then a has it can lead to multiple consequences.

    Bowtie diagram showing five types of hazard analysis.
    Bowtie showing the Foundations of System Safety

    That is what the bowtie here is representing. It’s showing that multiple causes can lead to a single hazard, and a single hazard can lead to multiple consequences.

    Don’t worry too much about the bow tie. I’m not pushing that in particular, it’s a useful technique, but it’s not the only one. We’ll come onto that – that’s the background.

    This is the accident sequence we’re trying to discover and understand. I’m going to talk a lot about discovery and understanding.

    Preliminary Hazard Identification

    Typically, we will start by trying to identify hazards. There are techniques out there that will help us identify hazards associated with the system being used in a specific application, or purpose, in a specific operating environment.

    Always bear in mind those three questions about the context, that help us to do this. What’s the system? What are we using it for? and in what environment?

    And if we change any of those things, then probably the hazards will change. But we start off with preliminary hazard identification, which is intended to identify hazards. There’s a big, big arrow pointing at hazards, but also, inevitably, it will identify causes and consequences as well, because it’s not always clear. What is the hazard when you start? talking of discovery, we’re going to discover some stuff.

    We may finally classify what we’re talking about later. we’re trying to discover hazards. In reality, we’re going to discover lots of stuff, but mainly we hope hazards, that’s stage one.

    System Requirements Hazard Analysis

    Now, then we’re actually going to step outside of the accident sequence itself. We’re going to do some requirements analysis, and the requirements analysis has to come after the PHIA because some safety requirements are driven by the presence of certain hazards.

    If you’ve got a noise hazard somebody’s hearing might be affected, then regulations in multiple countries are going to require you to do certain things to monitor the noise. Let’s say or monitor the effect that it’s having on workers and put in place a program to handle that. The presence of certain hazards will drive certain requirements for safety controls or risk controls.

    Then there are the broader requirements. Analysis of what the law requires, what the regulations require, codes of practice, etc. We’ll get onto that, and one of the things that requirements analysis is going to do is give us an initial stab of what we’ve got to have – certain controls because we’re required to. That’s a little bit of an aside in terms of the sequence, but it’s very, very important.

    Preliminary Hazard Analysis

    Thirdly, and, fourthly, once we’ve discovered some hazards, we’re going to need to understand what might cause those hazards and therefore how likely is the hazard to exist in particular circumstances, and then also think about the consequences that might arise from a hazard. And once we’ve explored those, we will be in a position to actually capture the risk.

     Because we will have some view on likelihood. And we would also have some view on the severity of consequences from considering the consequences. We’ll come onto that later.

    Looking at Controls

    Finally, having done all those other things, we will be in a position to take a much more systematic look at controls and say, we’ve got these causes. We’ve got these hazards. We’ve got these potential consequences.  What do I need to do to control this risk and prevent this accident sequence from playing out?

    What I need to put in place to interrupt the accident sequence, and I’ve put the controls. The dashed lines indicate that we’ve got barriers to that accident sequence, and they are dashed because no control is perfect. (Other than gravity. But of course, if you turn your vehicle upside down, then gravity is working against you, so even gravity isn’t foolproof.)

    No control is 100% effective. We need to just accept that and deal with that, and understand. There is your overview of the sequence, and I’ve spent a bit of time talking about that because it is absolutely fundamental to everything you’re going to do.

    Well, That’s a Brief Summary of the Foundations of Safety Assessment

    If you have any questions, please leave a comment below.

  • Preliminary Hazard Identification with Mil-Std-882E

    Preliminary Hazard Identification with Mil-Std-882E

    Want to know how to perform Preliminary Hazard Identification with Mil-Std-882E? (This is Task 201 under the standard.)

    This is the first step in safety assessment.  We look at three classic complementary techniques to identify hazards and their pros and cons.  This includes all the content from Task 201, and also practical insights from my 25 years of experience with Mil-Std-882. 

    You Will Learn to:

    • Conduct Preliminary Hazard Identification using diverse techniques for best results;
    • Define what Preliminary Hazard Identification is and does;
    • Record Preliminary Hazard Identification results correctly;
    • Contract for Preliminary Hazard Identification successfully; and
    • Apply it early enough to make a difference.
    This is the seven-minute-long demo video.

    Topics: Preliminary Hazard Identification with Mil-Std-882E

    • Task 201 Purpose & Task Description;
    • Historical Review;
    • Recording Results;
    • Contracting; and
    • Commentary:
      • Historical Data;
      • Hazard Checklists; and
      • Analysis Techniques.

    Transcript: Preliminary Hazard Identification

    Hello, everyone, and welcome to the Safety Artisan, where you will find instructional materials that are professional, pragmatic, and impartial because we don’t have anything to sell, and we don’t have an axe to grind. Let’s look at what we’re doing today, which is Preliminary Hazard Identification. We are looking at one of the first actual analysis tasks in Mil-Std-882E, which is a systems safety engineering standard from the US government, and it’s typically used on military systems, but it does turn up elsewhere.

    Preliminary Hazard ID is Task 201

    I’m recording this on the 2nd of February 2020, however, the Mil-Std has been in existence since May 2012 and it is still current, it looks like it is sticking around for quite a while, and this lesson isn’t likely to go out of date anytime soon.

    Topics for this session

    What we’re going to cover is, quoting from the task, first of all, we’re going to look at the purpose and the task description, where the task talks quite a lot about historical review (I think we’ve got three slides of that), recording results, putting stuff in contracts and then I’m adding some commentary of my own. I will be commenting all the way through, that’s the value add, that’s why I’m doing this, but then there’s some specific extra information that I think you will find helpful, should you need to implement Task 201. In this session, we’ve moved up one level from awareness and we are now looking at practice, at being equipped to actually perform safety jobs, to do safety tasks.

    Preliminary Hazard Identification (T201)

    The purpose of Task 201 is to compile a list of potential hazards early in development. two things to note here: it is only a list, it’s very preliminary. I’ll keep coming back to that, this is important. Remember, this is the very first thing we do that’s an analytical task. There are planning tasks in the 100 series, but actually, some of them depend on you doing Task 201 because you can’t work out how are you going to manage something until you’ve got some idea of what you’re dealing with. We’ll come back to that in later lessons.

    It is a list of potential hazards that we’re after, and we’re trying to do it early in development. And I really can’t overemphasize how important it is to do these things early in development, because we need to do some work early on in order to set expectations, in order to set budgets, in order to set requirements and to basically get a grip, get some scope on what we think we might be doing for the rest of the program. this is a really important task and it should be done as early as possible, and it’s okay to do it several times. Because it’s an early task it should be quick, it should be fairly cheap. We should be doing it just as soon as we can when we’re at the conceptual stage when we don’t even have a proper set of requirements and then we redo it thereafter maybe. And maybe different organizations will do it for themselves and pass the information on to others. And we’ll talk about that later as well.

    Task Description

    This is the task description. It says the contractor shall – actually forget about who’s supposed to do it, lots of people could and should be doing this as part of their project management or program management risk reduction because as I said, this is fundamental to what we’re doing for the rest of the safety program and indeed maybe the whole project itself. So, what we need to do is “examine the system shortly after the material solution analysis begins and compile a Preliminary Hazard List (PHL) identifying potential hazards inherent in the concept”. That’s what the standard actually says.

    A couple of things to note here. Saying that you start doing it after material solution analysis has begun might be read as implying you don’t do it until after you finish doing the requirements, and I think that’s wrong, I think that’s far too late. To my mind, that is not the correct interpretation. Indeed, if we look at the last four words in the definition, it says we’re “identifying potential hazards inherent in the concept”. That, I think, gives us the correct steer. we’ve got a concept, maybe not even a full set of requirements, what are the hazards associated with that concept, with that scope? And I think that’s a good way to look at it.

    Historical Review

    This task places a great deal of emphasis on the review of historical documentation, and specifically on reviewing documentation with similar and legacy systems. an old system, a legacy system that we are maybe replacing with this system but there might be other legacy systems around. We need to look at those systems. The assumption is that we actually have some data from similar and legacy systems. And that’s a key weakness really with this, is that we’re assuming that we can get hold of that data. But I’ll talk about the issues with that when I get to my commentary at the end.

    We need to look at the following…

    End: Preliminary Hazard Identification with Mil-Std-882E

    You can find a free pdf of the System Safety Engineering Standard, Mil-Std-882E, here.

    Meet the Author

    Learn safety engineering with me, an industry professional with 25 years of experience, I have:

    •Worked on aircraft, ships, submarines, ATMS, trains, and software;

    •Tiny programs to some of the biggest (Eurofighter, Future Submarine);

    •In the UK and Australia, on US and European programs;

    •Taught safety to hundreds of people in the classroom, and thousands online;

    •Presented on safety topics at several international conferences.

  • System Safety Engineering Process

    System Safety Engineering Process

    The System Safety Engineering Process – what it is and how to do it.

    This is the full-length (50-minute) session on the System Safety Process, which is called up in the general requirements of Mil-Std-882E. I cover the Applicability of Mil-Std-882E tasks, the General Requirements, the Process with eight elements, and the Application of process theory to the real world. 

    You Will Learn to:

    • Know the system safety process iaw Mil-Std-882E;
    • List and order the eight elements;
    • Understand how they are applied;
    • Skilfully apply system safety using realistic processes; and
    • Feel more confident dealing with this and other standards.
    System Safety Process – this is the free demo.

    Topics: System Safety Engineering Process

    • Applicability of Mil-Std-882E tasks;
    • General requirements;
    • Process with eight elements; and
    • Application of process theory to the real world

    Transcript: Preliminary Hazard Identification

    CLICK HERE for the Transcript

    System Safety Process

    Hi, everyone, and welcome to the Safety Artisan. I’m Simon, your host. Today I’m going to be using my experience with System Safety Engineering to talk you through the process that we need to follow to achieve success. Because to use a corny saying, ‘Safety doesn’t happen by accident’. Safety is what we call an emergent property. And to get it, we need to decide what we mean by safety, decide what our goals are, and then work out how we’re going to get there. It’s a planned systematic activity. Especially if we’re going to deal with very complex projects or situations. Times where there is a requirement to make that understanding and that planning explicit. Where the requirement becomes the difference between success and failure. Anyway, that’s enough of that. Let’s get on and look at the session.

    Military Standard 882E, Section 4 General Requirements

    Today we’re talking about System Safety Process. To help us do that, we’re going to be looking at a particular standard – the general requirements of that standard. And those are from Section Four of Military Standard 882E. But don’t get hung up on which standard it is. That’s not the point here. It’s a means to an end. I’ll talk about other standards and how we perform system safety engineering in different domains.

    Learning Objectives

    Our learning objectives for today are here. In this session, you will learn, or you’ll know, the system safety process in accordance with that Mil. Standard. You will be able to list and order the eight elements of the process. You will understand how to apply the eight elements. And you will be able to apply system safety with some skill using realistic processes. We’re going to spend quite a bit of time talking about how it’s actually done vs. how it appears on a sheet of paper. Also known as how it appears written in a standard. So, we’re going to talk about doing it in the real world. At the end of all that, you will be able to feel more confident dealing with multiple different standards.

    The focus is not on this military standard, but on understanding the process. The fundamentals of what we’re trying to achieve and why. Then you will be able to extrapolate those principles to other standards. And that should help you to understand whatever it is you’re dealing with. It doesn’t have to be Mil. Standard 882E.

    Contents of this Session

    We’ve got four sets of contents in the session. First of all, I’m going to talk about the applicability of Military Standard 882E. From the standard itself and the tasks (you’ll see why that’s important) to understanding what you’re supposed to do. Then other standards later on. I’m going to talk about those general requirements that the standard places on us to do the work. A big part of that is looking at a process following the eight elements. And finally, we will apply that theory of how the process should work to the real world. And that will include learning some real-world lessons. You should find these useful for all standards and all circumstances.

    So, it just remains for me to say thank you very much for listening. You can find a free pdf of the System Safety Engineering Standard, Mil-Std-882E, here.