Root Cause Analysis: The Missing Tool in the Operational Risk Management Toolbox?

A Review of the Operational Risk Management (ORM) Landscape

When the Bank of International Settlements (BIS) published the Basel II Accord in 2001 and mandated that market participants implement these recommendations by 1 January 2007, a majority of the affected players (large international banks with a total domestic annual turnover of over $250bn or an international turnover of over $10bn – referred to as the core banks) found the deadlines unreasonable. Amongst the accord provisions, those related to ORM appeared to pose the greatest challenge, given that most banks had very little infrastructure to systematically measure and manage operational risk. To make matters worse, many national regulators insisted that banks even adopt the advanced approaches, which would necessitate a loss database spanning a minimum of three years.

Despite the initial resistance, most of the core banks had initiated ORM programs by 2002. Their initial focus was on collecting loss events and building reasonably comprehensive loss databases. Once the data collection infrastructure was in place, these banks started tackling the other parts of the puzzle – analytics, reporting, workflows, etc. Though core banks made significant progress, opt-in banks (banking organizations not subject to Basel II on a mandatory basis but that choose to apply those approaches voluntarily) also managed to lay down their respective building blocks. In fact, many financial institutions (brokerage firms, asset managers, insurance companies, etc) that did not fall within the ambit of the Accord, voluntarily implemented the Accord. This market realized that the medium-to-long-term benefits of a strong ORM infrastructure far out-weighed the implementation costs.

Given the recent level of participant enthusiasm, the Basel deadline no longer appears so formidable – leading many to believe that most of the core and opt-in banks are within striking distance. Even though the goal is in sight, some distance still remains to be covered. This is especially true of the risk mitigation part of the ORM puzzle, which includes the establishment of controls and mechanisms of risk transfer.

An Overview of RCA and its Role in ORM

RCA is a problem solving technique that does exactly what the title suggests – it provides a structured way to analyze a problem and identify the associated root cause (as opposed to symptoms). RCA has been, and continues to remain, a frequently used technique in the manufacturing industry. Though it is a relatively unheard term within financial services, this industry agnostic technique is as applicable to financial processes as it is to manufacturing processes. In particular, the probing characteristics of RCA make it eminently suitable to go under the skin of operational problems and subsequently lead to the design of a better control infrastructure.

When first introduced to RCA, reactions typically vary from, “That’s trivial, everyone does it” to “Oh… yet another three letter acronym”. In fact, many people are initially mistaken by the apparent simplicity of the technique and fail to appreciate its usefulness. However, a few RCA exercises later, it often becomes evident that the most apparent causes are rarely the root causes – emphasizing the benefit of employing a structured framework to reveal the root of a problem. For example, when an associate is blamed for a data entry error, the root cause may have a more fundamental source such as overwork, lack of training, poor alignment of processing accuracy with incentives, etc. One rule of thumb for identifying a root cause is that it has to be something actionable – in other words, if a control does not address the vulnerability, then it cannot be classified as a root cause.

Though conceptually simple, implementing a RCA program is not without its challenges. One of the most common roadblocks is the problem of plenty, as there are simply too many tools to choose from. Some of the more common aids include histograms, pareto charts (used to graphically summarise and display the relative importance of differences between groups of data), scatter charts, relations diagrams, affinity diagrams, and cause-effect diagrams. Add these to your list of personal favorites like barrier analysis, change analysis, MORT (management oversight and risk tree), kepner-tregoe method (a decision-making model), etc, and the inventory of options simply become unmanageable. To add to the problem, most of the techniques are usually applied in an isolated fashion and rarely, as sole initiatives, manage to identify the root cause.

The RCA Framework

The framework comprises four distinct phases – gathering information about the loss event, using RCA tools and techniques to analyze the event and identify the root cause, recommending controls to resolve identified vulnerabilities, and reporting. Each of the four phases are detailed below.

Loss Data Collection

This phase involves collecting details about the loss event. It is recommended that, at a minimum, the following data elements are collected. If an organization already has a loss database, most of these data elements are likely to be present and can be directly sourced from the database.

  • Summary information about the event – description of the event, business unit where it occurred, occurrence and reporting dates, etc.
  • Impact details – type of impact (financial, legal, reputational, etc), amount of the loss, type of the loss (direct loss, opportunity cost, cost to fix, near miss), etc.
  • Sequence of events – should capture the exact sequence of events that led to the loss, along with information on date/time and people involved.
  • Action taken – this section would capture an organization’s reaction to the loss event and steps taken to mitigate the possibility of future occurrences.
Analysis

This phase constitutes the heart of the RCA program and comprises multiple steps – each of these steps, along with recommended tools, are described below.

  • Generate event flowthis step involves constructing a visual depiction, usually in the form of a flowchart, the actual sequence of events. Particular attention should be paid to the steps which deviated from the expected path.
  • Use pareto charts to analyze similar events from the pastthis step involves analyzing causes that led to similar events in the past. This enables the RCA analyst to focus on the more probable causes. As is evident, a precondition for this step is a sufficiently deep repository of causes – hence, this step cannot be performed in the early days of a RCA program.
  • Brainstorm on possible causes – at this stage, the RCA analyst, in conjunction with the people (typically line managers and process associates) most familiar with the exception, get together to identify a possible list of causes. These causes are then grouped under four dimensions – people, process, systems or external events. This output, commonly, is referred to as the affinity diagram. At the end of this stage, the probable list of causes includes the ones that were generated in this step as well as in the previous step.
  • Scrutinize each causeeach of the causes identified in the previous two steps are analyzed in detail. Essentially, a series of questions are asked to discover the cause. The questioning is continued to the point at which the cause is found to be actionable and can be controlled by the organization. To aid the investigation process, the fishbone diagram (also referred to as the cause-effect diagram or the ishikawa diagram) could be used.
  • An illustration might clarify the point. For a particular operational error, the most apparent cause could be ‘associate failed to attach source documentation while processing a transaction’. On delving deeper, the action could be attributed to ‘disregard for procedures’, which in turn could be explained by the fact that error rates are not factored in associates’ incentive plans.
  • Identify the root causeall the lowest level actionable causes (terminal causes identified in the previous step) would need to be ranked. The ranking should, ideally, be done by multiple people familiar with the event (this will tend to normalize individual biases). The cause with the maximum score would be selected as the root cause. In certain situations, more than one cause might be identified as the root cause.
  • Add the root cause to a ’cause repository’the last step of the analysis phase involves feeding back the results of the analysis to a repository. Over time, the repository will deepen and facilitate step two.
Recommendation

This is a crucial step, especially from a management perspective. Based on the root cause identified in the analysis phase, appropriate action plans (short term) and recommendations (medium-long term) need to be presented. These should ideally be ranked to facilitate implementation – based on parameters such as cost, impact, ease of implementation, etc. Each recommendation should be assigned to one or more persons, along with implementation plan and schedule, and tracked till completion.

Reporting

Unlike the preceding phases, which are associated with a single event, reporting aggregates multiple events. The reporting frequency should, ideally, be set at a monthly or a quarterly basis as it would ensure that the RCA reporting cycle is synchronized with other organizational reporting initiatives such as quality control, internal audit, etc. The framework suggests some indicative reports which may be altered depending on organizational necessities.

  • Management dashboard – This would provide a snapshot of the RCA program and would typically comprise entities like ‘most affected business lines’, ‘most frequent exception categories’, ‘most prevalent root causes’, etc.
  • Trend report – would be used to display loss trends – both in terms of frequency as well as severity of losses.
  • Drill down analysis – this report would allow users to slice and dice the base data across multiple dimensions such as business line, exception category, root cause, etc. These are truly on-demand reports and the formats can be dictated by individual users. Contrary to popular perception, this kind of analysis can be performed using everyday office tools like pivot charts.
  • Aging report – This report would be used to track the implementation of both short-term and long-term recommendations. The information could be presented both graphically as well as depicted through a tabular structure.

What are the Technology Implications?

At minimum, a basic RCA exercise may be implemented without a significant investment in specialized software tools. In fact, one or more office tools, such as MSWord, Excel, or Visio may be sufficient for implementing a simple RCA program.

However, in the long run, assuming RCA gains acceptability in even a moderately-complex organization, these office-based tools will suffer from a lack of scalability. From a cost-benefit standpoint, a custom solution probably makes more sense. In this case, the framework could be implemented as a stand-alone component or be integrated within the larger context of an organization’s operational risk management platform.

Conclusion

RCA is an enabler to search beyond the obvious symptoms and reveal the root causes behind operational challenges. Once these causes are understood, they can then be leveraged to design more effective controls and risk indicators. We expect that an RCA program will not only lead to a more resilient operational risk management infrastructure, but also pay for itself – given the benefits of reduced operational losses (resulting in a lower capital charge) that should more than compensate for implementation costs. These factors make RCA an indispensable tool in the operational risk manager’s toolbox – and we strongly suggest that ORM initiatives leverage its potential to better defend against operational deficiencies.

Whitepapers & Resources

2021 Transaction Banking Services Survey
Banking

2021 Transaction Banking Services Survey

5y
CGI Transaction Banking Survey 2020

CGI Transaction Banking Survey 2020

6y
TIS Sanction Screening Survey Report
Payments

TIS Sanction Screening Survey Report

7y
Enhancing your strategic position: Digitalization in Treasury
Payments

Enhancing your strategic position: Digitalization in Treasury

7y
Netting: An Immersive Guide to Global Reconciliation

Netting: An Immersive Guide to Global Reconciliation

7y