In today’s digital and fast-paced financial markets, it is more important than ever for trade workflows to function efficiently. Trading systems and processes must also be robust and resilient so that firms can operate effectively, comply with regulatory requirements, maintain orderly markets and, ultimately, ensure healthy revenues and avoid reputational damage. Business teams must be confident that their trades, systems, and processes are monitored comprehensively so they can mitigate and avert any problems. In this article, GreySpark and ITRS explore why the business should ensure comprehensive real-time monitoring is undertaken by their IT teams.

By GreySpark’s Jennie Brotherston, Senior Specialist and Rachel Lindstrom, Senior Manager

The EU’s Digital Operational Resilience Act (DORA) is due to come into force on 17 January 2025, imposing new requirements on financial institutions that underscore the need to withstand, respond to, and recover from technological disruptions and cyber threats. Financial firms’ digital infrastructures continue to increase in complexity, in part because some elements of the trade process are outsourced to third parties, and in part because trading depends on a wide variety of technologies. Consequently, it is essential that firms monitor all critical systems and processes to detect, respond to, and recover from any emerging issues that might threaten or interrupt their ability to operate.

The smooth functioning of systems and processes is critical not only for an individual firm’s success, but also for the functioning and integrity of the wider market. Efficient markets rely heavily on timely and accurate data — particularly around prices — and on market participants having reliable access to appropriate platforms and communications. Delays or failures that occur in trading, either centrally or with individual participants, can cause delays for other participants and have a knock-on effect on the whole market. This can impact the ability to trade with the full range of available counterparties or sometimes even the ability to trade at all. The faster a market participant can identify an issue, the easier it is to prevent it from affecting the wider market. Globally, regulations including Dodd Frank, MiFID, MAR and DORA all require market participants to have tight control over their trading infrastructure or they may face large fines.

Alongside the collective need for efficiency from market participants, individual organisations also incur a notable cost for system or process inefficiency. Delays in the trading process can result in direct costs and lost revenue, while investigating and resolving any issues places demand on resources. The reputation of a firm with clients, counterparties and other third parties — including exchanges, custodians, regulators, and more — can also be negatively affected by trading failures, which amplifies the impact on revenue and costs. In an industry where trading margins are often slim, this can be the difference between profit and loss. Safeguarding the efficiency of the end-to-end trading process is a vital part of ensuring the business is viable, profitable and sustainable.

What Can Go Wrong

Between the initiation of an RFQ or order and the final trade settlement or billing, there is a wide range of points at which issues could arise, cause delays or prevent successful trade execution. Even seemingly small events can have a significant impact on the successful execution of a trade.

There is also a myriad of systems that can cause the same visible outcome. For example, a stale price may be due to a failure at the exchange, an issue arising with a market data vendor or in a bank’s own data centre, or from a failed service within a trading application. In this situation, being able to see the impact alone is not enough. It is important to have real-time context and a granular understanding of exactly what is happening — not only the manifestation of the problem, but also identification of the events or issues that have caused it to happen.

Figure 1: Key Processes / Systems Involved in the Trade Lifecycle
Source: GreySpark analysis

Figure 1 shows, at a high level, a non-exhaustive illustration of the many systems that may be involved when making a trade. Delays could occur at any number of points in the trade lifecycle and include:

  • Technical and infrastructure failures, such as issues with the FIX engine or connectivity problems. This can lead to incorrect or delayed information being passed by or within systems, which quickly results in delay, interruption and failures in trading.
  • Application failures, which can negatively impact any stage of the trading process.
  • Failures in post-trade processes, resulting in incorrect, incomplete or delayed regulatory reporting or incorrect billing and reconciliation, which then risks financial, operational and reputational impact.
  • Outages at venues or in the external market, significantly affecting an organisation’s ability to conduct business unless they can react swiftly and trade using alternative platforms or markets.

Figure 1 also illustrates that not only sales and trading staff, but also middle- and back-office staff who are involved from booking through to settlement, will be impacted by delays or failures in the process. Issues are likely to impact more than the trade or process segment in which they arise, creating a ripple effect on other parts of the process that may continue after the issue has been resolved.

Support teams must have visibility of the entire trade process if they are to be able to determine where an error has occurred and how to remediate it quickly. When problems arise in the trading process, support teams will often be under the most pressure to rapidly identify and remediate the issue.

What Can Be Done to Avoid It

It is not always possible to prevent issues from arising, but the business can empower support teams to head off issues. Support teams should be enabled to identify issues before they become significant and alert the business in advance. If there are events or problems that might impact trades or the trading process, they should be able to take mitigating or evasive action — before they learn of the problem from clients or the wider market. Where possible, monitoring and alerting should take place as near as possible to real time to facilitate this.

Even when monitoring is in place, some outages may still occur where monitoring teams are unable to immediately identify the cause, yet it is imperative that the outage is resolved quickly to maintain an orderly market. If monitoring teams or systems can alert the business of impending latency data feed problems, for example, it may be possible for them to take evasive action by switching to an alternative provider, exchange or backup feed as needed. Similarly, if a venue were to ‘go down’, the ability to spot this immediately would allow traders to direct their trades elsewhere — potentially in advance of the market. The commodified nature of trading means that market participants can often switch to another provider with relative ease.

An effective monitoring process would ideally incorporate a comprehensive data flow dashboard, so that the monitoring team is able to identify bottlenecks in processes. This lets them see mismatches in the speed of data flow before and after a certain point and quickly identify, investigate and divert resources as necessary. It would also include the ability to detect and alert on anomalies in data quality, such as stale pricing or inconsistency of trade attributes between systems. Achieving this real-time identification and rapid resolution of issues requires robust operational processes, as well as a holistic view of all workflow elements and data. Users must be able to quickly ascertain dependencies and interactions, so they can understand how the problems cascade through systems and processes.

Traditionally, system logs were relied on for monitoring and can still be useful for investigative purposes — for root cause analysis after the event, for example — but they are also backward looking and can only provide information on each system in isolation. In a more complex environment, having a full set of logs, metrics, and telemetry for all systems in a process is necessary to track data and trades flows, since data passes from one system or process element to the next. With these metrics and data visualised in real time, flows may be monitored from end to end for completeness and latency. The tracking and timing of the individual and collective elements of trades through the entire process facilitates the monitoring of hand-offs between systems.

Additionally, incorporating the data itself into the monitoring process will allow qualitative checks, so that accuracy as well as performance can be reviewed. There is great value in retaining clean, normalised historic monitoring data so that, if things do go wrong, it is possible to analyse the cause and the warning signs to avoid that specific issue in future. Historic monitoring data can also be used in ‘backtesting’ scenarios and to conduct scenario analysis. Used in combination with current data to model ‘what ifs’, organisations can make robust plans to maintain business continuity under various scenarios, ranging from market crises to more minor day-to-day errors and issues that might arise.

No matter how much backtesting is done, the reality is there will always be a risk of unforeseen events occurring. In a complex infrastructure, it is vanishingly likely that every possible eventuality or scenario can be mitigated, so it is imperative that monitoring of infrastructure, applications, transactions, and data be undertaken in real time and viewed in a comprehensive dashboard. This oversight can add further value to the firm by enabling, for example, the documentation of risk assessments, audit trails and reporting, all of which are DORA requirements. As a result, financial firms can truly achieve operational resilience and ensure an orderly market.

ITRS Geneos empowers financial services with real-time IT monitoring and alerting. It enables them to gain a competitive advantage from IT performance and drives resilient operations, delivering intelligent insights that reduce the business impact of outages and application issues in production. ITRS Geneos supports the most diverse and interconnected IT estates at scale without impacting system performance. It monitors transactions, processes, applications, and infrastructure across on-premises, cloud, and highly dynamic containerized environments, giving financial services full control over IT ecosystems and ensuring minimal latency 24×7.