Data presents the capital markets with the opportunity to optimise trading and make superior data-backed business decisions. As digitisation of capital markets has progressed, the amount of electronically available data has grown, but also the types of data used has expanded rapidly. As financial firms seek new sources of alpha, alternative data is emerging as an important means by which to improve the competitiveness of investment firms. According to Builtin, roughly half of investment firms already use alternative data to forge investment strategies.

By GreySpark’s Rachel Lindstrom, CMI Research practice

Alternative data refers to the type of data; it is something other than market data and reference data. It does not necessarily come from traditional or official sources, such as press releases, quarterly earnings reports or broker pricing data feeds. Rather, it comes from specialist data vendors that use enhanced methods to obtain different and new data sets. For example, consumer transaction data, ESG data and sentiment data are all types of alternative data. Sentiment data can be sourced from a data vendor which collects billions of data points using machine learning from a social media platform to detect investor sentiment towards a specific area of the market. According to one report, the alternative data market was worth USD 6.61 bn in 2023 and is predicted to grow at a compound annual growth rate of 57.7% between 2024 and 2032. Although it can provide a broader picture of the world, alternative data is often unstructured and requires deeper processing and analysis to extract its full value. For example, while much data can be accessed via a greater number and variety of data channels, the fact is that sometimes these channels are unorthodox and could potentially compromise data quality, control and management.

Data providers play a crucial role in vetting, cleaning, aggregating and distributing alternative data to clients. However, this segment of the data market is still relatively immature, and there remain questions around the quality of the datasets and their value to financial services sector analysts. In 2025, it is vital for financial firms to optimise the use of alternative data as, while it is increasingly being used in investment decision making, data costs are being scrutinised to ensure it is purchased wisely.

Understanding the 2025 Alternative Data Landscape

In line with the rise of alternative data, new data platforms are emerging from which to manage it. In the capital markets, Data-as-a-Service (DaaS), data marketplaces and alternative trading venues are three of the most impactful data trends to have emerged over the last decade.

DaaS models have grown in importance for financial firms due, in part, to their ability to give access to richer pools of alternative data and to improve data management efficiency. Using a DaaS, banks can access data directly from data vendors via DaaS plug-ins, reducing compute and storage requirements of on-premises servers. Given the benefits of the DaaS model, it is unsurprising that financial firms are increasingly embracing them, and that key market data providers and exchanges also now offer clients DaaS plug-in APIs and centralised data distribution models.

In the DaaS model, data from multiple third-party data providers is ingested by the DaaS, whatever the format or protocol, and the DaaS delivers it via the cloud to its financial firm clients (see Figure 1). Although the firm can be exposed to additional latencies when retrieving data from the cloud, the impact of this is generally insignificant. The ingestion challenges presented by the diversity of the data delivery mechanisms, file formats and data schemas are absorbed by the DaaS, which ensures that the financial firm receives data in the format that they want.

These benefits are multiplied if the DaaS ingests third-party data from multiple providers and delivers via one data mechanism to multiple clients – in this case, the DaaS model morphs into a DaaS Utility model. DaaS can therefore provide significant efficiencies for procurement teams who would otherwise have to manage a complex data management operation.

The use of data marketplaces over the last ten years has notably increased. Data providers allow a data marketplace to sell its data for a fee and revenues are passed through from purchaser to data provider. The upside for the purchaser is access to specific data in a standardised or normalised format, which suits its needs, creating a win-win for all three parties. Before the financial crisis, data marketplaces focused on intraday and historical market data, but since then more data marketplaces offer close to real-time alternative data.

Click to download PDF

Data-as-a-Service is a cloud-based software tool used for working with data, such as managing data in a data warehouse or analysing data with
business intelligence. Like all “as a service” (aaS) technology, DaaS builds on the concept that its data product can be provided to the user on demand, regardless of geographic or organisational separation between provider and consumer.

Data Marketplaces bring together data providers/vendors and data consumers, facilitating the buying and selling of data. They enable participants to access data more easily, saving time and resources compared to collecting data from scratch.

Alternative Trading Venues in this context are a specific type of data marketplaces where various types of non-traditional or alternative datasets are bought and sold. As highlighted, alternative data typically includes consumer behaviour and sentiment rather than traditional market data.

Figure 1: Data-as-a-Service (DaaS) Delivery Approach
Source: GreySpark analysis

(Click image to enlarge)

Firms can, of course, access alternative data by directly contracting with and connecting to a third-party data provider, rather than opting for the DaaS or data marketplace approach. This is typically the best option for obtaining niche alternative data sets that are not typically widely available in the public domain. This can include sentiment and geolocation data, for example. Third-party data providers collect content from several external sources and apply their own proprietary algorithms to the data to extract usable insights before selling them on to the consumer.

Broadly speaking, despite having the option to purchase data across a multitude of channels, financial firms still control the data that enters their systems. The large amount of data they can purchase in various formats, however, makes ensuring data quality and sound data management challenging. In many cases, a lack of data monitoring causes duplications, extended latencies, and ultimately higher costs for the firm as it troubleshoots the consequential issues.

Despite the trope that ‘Data is King’, data within financial institutions is often badly managed, squandering the opportunity to use it to generate a competitive advantage. Data management is almost never a simple task, and with traditional and non-traditional data coming from an ever-wider array of sources, it can be a huge challenge to track and manage it effectively across the data value chain. For instance, where there is no centralised data inventory system, the financial services firm will invariably pay for duplicates of data. Each data contract proffered by a third-party data provider can be quite different in form from those of other providers as there is no mandated standard, which means that cross comparison is far from simple.

Figure 2: Three-stage approach to Data Management
Source: GreySpark analysis

(Click image to enlarge)

Firms can mitigate these challenges with a three-stage approach (see Figure 2), which will enable them to manage their data efficiently, regardless of the method they chose to obtain it and the type of data:

1. Sourcing – the process of identifying, obtaining and managing data from various internal and external sources for analysis and decision making. In sourcing data, firms must ensure it complies with legal and organisational policies, with managing permissions, sensitive data protections and data privacy policies in place. The source of the data is identified, which can include databases, data warehouses, data lakes, APIs, web scraping, third-party vendors, DaaS providers and public data sets. The data is then collected, which involves extracting data in various formats such as structured and unstructured data. In addition, the sources of the data and the method of collection are documented.

2. Ingestion – the collection and importing of data from various sources into a centralised repository for processing, analysis and storage. The data is integrated into a firm’s databases. The data is parsed and converted into a format that can be understood and processed by the destination system. This may involve handling different data formats such as .JSON, .XML or .CSV. Data is then transferred to the target repository, which could be a data warehouse, data lake, database or other internal storage systems. Data ingestion can occur in batch mode, where data is collected and processed in large chunks at scheduled intervals, or in real-time / streaming mode, where data is ingested and processed continuously as it arrives. It is essential that the data ingestion process is optimised to handle large volumes of data efficiently and to scale with increasing data loads. The data is then checked for errors and validated, ensuring accuracy, completeness, and reliability. This may involve cleansing or enrichment activities and can be completed through automation or human intervention, or a combination of the two.

3. Handling – the processes and techniques used to manage, manipulate and utilise data throughout its lifecycle, ensuring its quality, security and accessibility. It encompasses a wide range of activities and practices aimed at maintaining the integrity and usability of data. The process for managing data from multiple vendors is similar to the process outlined above. Data is integrated into the firm’s databases, before it is processed, cleansed and converted into a common format. The firms can then apply analytical techniques and data science methods in order to understand the data while also applying mechanisms to retrieve data. Data processing must be in line with regulation and any organisational standards. Ongoing maintenance and management of the data is an important requirement. Copies of the data are made to prevent loss and ensure recoverability in case of failures or disasters. Policies and procedures for data management are established to ensure data quality, consistency and compliance with legal and regulatory requirements. It is essential for firms to manage the long-term storage of data that is no longer actively used for historical, legal or regulatory purposes. This also involves securely disposing of data that is no longer needed.

Alternative data is taking on an increasingly important role when it comes to trading decisions, with increasingly digitised capital markets having greater access to a plethora of new data channels. However, utilising a wider array of data sources could potentially compromise data quality, control, and management across the industry unless properly managed. Understanding this, as well as the different stages of the data lifecycle, can help firms to optimise their market data strategies and generate long-term value to the organisation.