AI surge linked to declining data quality for hedge funds

About one-third of buyers say dataset quality has worsened over two years as data vendors increase use of AI, Neudata research finds.

Hedge funds report that rapid adoption of artificial intelligence by data vendors has lowered the quality of some external datasets, Neudata research shows, with about one-third of buyers saying data quality has deteriorated over the past two years.

Alternative data-information on consumer behavior, corporate activity, web traffic and other non-traditional indicators-is used by funds to seek trading signals ahead of the market. The spread of AI tools has reduced costs and technical barriers for collecting, processing and packaging such information, and the number of firms selling packaged datasets to investors has grown.

Buyers say problems occur when AI is used to generate, classify or interpret data with limited human oversight. Vendors increasingly deploy machine learning and large language models to create products or to justify cutting or redeploying staff who formerly checked and cleaned data. Industry participants say that can create gaps in quality control and make it harder for clients to assess how reliable a dataset is.

Daniel Entrup, co-founder of data product firm AggKnowledge, reported an increase in basic data errors since AI adoption accelerated and noted that even small mistakes can have large effects in trading models. A single flawed observation can distort a model or produce a misleading signal that is incorporated into a strategy.

A hedge fund manager summarized the downstream risk in practical terms: “Once you have a failed signal, it’s polluted data.”

Daryl Smith, head of research at Neudata, warned that funds can face problems when vendors cannot adequately explain how a dataset or report was produced. Some managers prefer to take raw information from providers and process it through their own infrastructure rather than rely on a vendor’s proprietary AI model.

Concerns extend to provenance, reproducibility and regulatory exposure. Using third-party datasets that involve personal data or sit in legal gray areas can create compliance and privacy risks if vendors cannot document collection and processing methods.

Large managers with in-house data science teams are more likely to ingest raw feeds and build their own pipelines to retain control over how raw inputs are transformed into trading signals. Vendors unable to demonstrate human oversight, explainable processing steps or tested model performance face increased scrutiny from clients that require traceability and reproducibility.

Articles by this author