Operational risk in financial institutions is often difficult to pin down. A bank or financial institution that does not have proven processes in place, and tested plans around a potential disaster or major outage, faces risk every day. In fact, any financial institution that does not revisit this type of planning on a regular basis could be in danger of going out of business. At the same time, any loss of data or critical information on accounts poses huge risk. Ensuring that these critical banking systems are continuously available is the major focus of most IT groups inside a bank – no matter how large or small.
Defining the Problem
A major challenge around business continuity and disaster recovery planning has been finding an agreed and accepted definition of the events that might constitute ‘risk’ and how key they are to the functioning of the bank’s core activities.
Typically it is described as the potential for loss arising from inadequate or failed internal processes, people and systems or external events, which cover a number of risk categories such as legal risks, people risks, information technology risks, compliance risks and so on. And there is now also a growing trend towards using the Basel II definition:
‘The risk of loss resulting from inadequate or failed internal processes, people, and systems or from external events’.
Managing operational risk should be at the heart of any business continuity plan and many of the leading financial institutions are making good progress towards establishing an operational risk framework: a framework that provides their board and executive team with assurances that an appropriate level of control is in place to identify and mitigate the operational risks that pervade the business processes of any organisation today.
Data Security
Data security and recovery may be the most important aspect of a business continuity plan. In fact, most organisations who lead the charge on this effort today don’t even like to think about ‘recovery’ but more often categorise it as ‘disaster tolerance’ or aim for ‘continuous access’ of all critical systems and data. Most IT groups today take into account the wide range of reasons why an application or system might experience downtime. Even a localised power outage lasting just three or four hours could cause a major disruption to the bank and its local customers but continuing business as usual would really depend on what processes and contingency is in place to deal with such an event. One option is having data centres located in different geographic locations that can take care of outages like this occurring: the power may be out in central London, but a customer in that region who needs to draw cash from an ATM could successfully do so if the bank’s secondary data centre is a couple of hundred miles north in Leeds, for example. Of course when the power goes out in Leeds, customers draw from the systems based in London. But that’s another story. In an industry that’s paying more and more attention to the management of operational risk, it makes good business sense to have a plan in place that protects the critical transactional systems and that data and a smarter organisation lays out a proven and tested process for keeping business up and running.
Fire Risk
Evidence that firms do not invest enough time and resources into business continuity preparations can be seen in disaster survival reports and statistics.
Fires permanently closed 44% of all businesses affected in the February 1993 World Trade Center bombing and 42% (150 businesses out of the 350 affected), failed to survive the catastrophic event in the long term. The World Trade Center complex was protected by an extensive fire detection and voice evacuation paging system that was considerably upgraded after the 1993 bombing.
Conversely, the increase in preparation in the years following that attack, meant that the firms affected by the 11 September attacks in 2001, who had well-developed and tested business continuity plans in place, were back in business within a matter of days despite the increased scale of the disaster.
Data and information protection is at the heart of this type of planning and that applies to nearly all regulatory compliance. The way information is processed, stored and archived defines the business process activities of almost every organisation and how well this is followed and complied with determines an organization that follows compliance and regulations. It is no surprise, therefore, that all the regulatory initiatives currently affecting businesses around the world place great emphasis on how data is managed.
Data Transactions
High quality, accurate and consistent data is required for reporting under a wide range of regulations – from the general International Financial Reporting Standard (IFRS) and US Sarbanes-Oxley Act to specific financial sector regulations such as the Basel II capital accord and the European Union’s Market in Financial Instruments Directive (MiFID).
Businesses increasingly must be able to access, analyse, act on, and verify data transactions faster than ever, often in real time – and all without system interruption or downtime.
Top of mind issues for the IT team is focused around the following assumptions:
- Exponential growth in transactional data volumes.
- Store, manage and account for such data transactions.
- The need for flexibility and work within heterogeneous, cross-platform, and distributed environments.
A wide variety of technology solutions have been adopted to protect critical systems and their data including, mirroring, physical replication, and in many cases traditional tape back-up and storage. Many of these approaches have their shortcomings and in the case of tape back-up, a significant amount of system downtime could ensue in the event of a major outage. Locating and uploading the latest tape back-up could take hours and even days under some circumstances.
More importantly, many of these approaches could yield significant data loss that is detrimental to any financial institution. Today’s progressive organisation is deploying an active/active configuration for their most critical systems that means two systems are processing the same transactions, at the same time, so that if any one system goes down, the second takes over and there is no impact to the end-user. The banking customer is not even aware which system they are accessing and they are happily performing their bank transactions as long as everything is up and running with accurate, up to date information.
Live Standby
Another popular approach to providing a high level of availability is known as ‘live standby’ whereby if any outage occurs, you immediately fail-over to a secondary system which is immediately available to support users with the most current data committed at the source up to its point of failure. This is usually a software implementation that captures transactional data from the primary system and applies it to the backup database in real time.
Further, the live standby approach provides for bi-directional data movement so that once the primary system is ready to be brought back online, any new data processed by the backup system is applied to the primary. With this solution, the primary and secondary systems are synchronised at all times. Additionally, the investment in this secondary system can also be leveraged to support other business activities – such as, real-time reporting or testing and development.
Hackers
At the infrastructure level, security is another issue. Hackers and information fraud pose very real threats, and data theft is not as random as it might seem. It’s perpetrated against businesses with the weakest security and the most valuable information. More than 80% of all money loss comes from 20% of fraudulent transactions.
Most financial institutions have deployed specific software solutions that are designed to specifically look for fraudulent activity occurring against any one customer. Credit card theft is a serious matter and by closely examining the transaction behaviour, the bank can apply certain business rules and prevent further card usage.
The underlying technology that is required to do this is real-time data integration across the various banking systems. If you can look for transaction history that is streaming every second of the day, you have a better chance of detecting the fraudulent or suspect activity.
Operational risk is a fact of life and putting plans and contingencies in place to deal with the unforeseen as well as the planned outage events is the best one can do. Theft and fraud is also not going away any time soon, but banking organisations can put more aggressive software and processes in place today to mitigate risk.
Just recently when I took my summer vacation across Europe, I was faced with a situation where my bank card did not work in any store or retail outlet and, additionally, I was only able to get cash out of certain ATMs. It was only after I called my bank to complain about this nuisance that I discovered they were detecting some fraudulent activity and could not understand why I was drawing money and making transactions in a new city across Europe each day. A quick call later, and it was all fixed and I could continue on my holiday knowing that my bank was doing the right thing to protect me as a customer.