Why Quality Data is Essential for AI Success

This article is by Barley Laing, the UK Managing Director at Melissa

AI has huge potential to help insurers improve decision making at speed, and lower costs, but many are struggling to generate value from it.

A key reason why was identified in a 2025 Lloyds Market Association survey, where 49 per cent of respondents cited data quality and availability as a barrier to AI adoption.

In fact, recent research by MIT has found that 95 per cent of organisations are seeing no measurable return from their generative AI pilots. Only a small minority, around five per cent, are extracting real value at scale.

While there are often a number of factors for these failures the standout reason, as highlighted by the Lloyds research, remains a lack of quality data for AI to work with, rather than the technology itself.

Data quality challenges and ongoing data decay

Currently, 94 per cent of organisations estimate that they have data quality issues and are operating with inaccurate, duplicate or incomplete data.

A leading cause of poor data quality is ongoing data decay – something that’s having a significant impact on the effective implementation of AI.

Data decays rapidly with customer contact data lacking regular intervention degrading at 25 per cent a year as people move home, die and get divorced. Furthermore, 20 per cent of addresses inputted online contain errors, which include spelling mistakes, wrong house numbers and incorrect postcodes. This can lead to data teams spending more than half of their time on data preparation – cleaning and structuring data so it’s ready for analysis – rather than on insight generation.

Access to clean data is critical for insurers attempting to train, deploy, scale and determine the return on investment (ROI) from their AI initiatives. Inaccurate data can lead to unreliable automation, ineffective personalisation, poor recommendations, and ultimately, a loss of customer trust.

Strengthen data verification processes

To avoid the problem of inaccurate contact data it’s essential to have verification processes in place at the point of data capture, and when cleaning held data in batch. This usually requires straightforward, cost-effective improvements to the data quality process.

Utilise address lookup or autocomplete

A good place to begin is to use an address autocomplete or lookup service at the customer onboarding stage. These provide accurate address data in real-time when onboarding new customers by automatically suggesting a correctly formatted address as they begin entering their details. By adopting this technology the number of keystrokes required when typing an address is reduced by up to 81 per cent. This helps to streamline and speed up the onboarding process, reducing the likelihood of users abandoning their applications or purchases. Importantly, this first point of contact verification approach can also be applied to email addresses and phone numbers, enabling these valuable contact channels to be verified in real time.

Deduplicate data

Duplicate rates of 10 – 30 per cent are not uncommon on insurer’s customer databases. It is a significant issue that often arises when two departments merge their data, after a new business acquisition, and with errors in contact data collection occurring at different touchpoints.

Duplication can not only confuse AI applications, but also increase costs in terms of time and money, particularly with printed communications – something that’s potentially very damaging to the sender’s reputation.

A standout way to prevent duplication is to use an advanced fuzzy matching tool to deduplicate data. Employing such a service enables insurers to merge and remove duplicate or inconsistent records, to create a ‘single user record’ that provides an optimised single customer view (SCV) from which AI can derive more accurate insights.

Data cleaning / suppression

Another important part of the data cleaning process is to undertake data suppression, or cleansing. This necessitates using technology that highlights people who have moved or are no longer at the address on file. It’s a vital part of the data cleaning process, and therefore in supporting efforts with AI. Alongside removing incorrect addresses these services can also include deceased flagging which helps to prevent mail and other communications from being sent to those who have passed away and potentially causing distress to their friends and relatives. By employing suppression strategies insurers can save money, protect their reputations, avoid fraud and support their AI efforts.

Enhance data to support AI success

It is imperative to enrich customer data with demographic, firmographic, geographic, social media and property attributes, as well as add missing email and phone information. Along with supporting their underwriting efforts obtaining such data will assist insurers in maximising their wider AI efforts in analytics, personalisation and omnichannel marketing. This way a more complete data picture will emerge which will support the delivery of increasingly accurate predictions, because giving AI additional signals to identify patterns, predict customer needs and assess likely outcomes results in improved recommendations and communications.

The new Centre of Excellence will create more than 200 roles in data analysis, software development, and quality assurance.

Make sure data is machine readable

AI agents need to be able to draw from a foundation of high quality API or machine-readable data. Because AI systems require data they can access, interpret and process quickly at scale without the need for manual intervention.

It is structured, well formatted machine-readable data that reduces ambiguity, which in turn enables AI models to interpret information more accurately. At the same time, decision making is accelerated, as AI can analyse and act on machine-readable data far more quickly than humans. Integration is also enhanced, as machine-readable data can flow more seamlessly between different systems, applications and AI tools.

Therefore, using well labelled data is strongly recommended to ensure accuracy across mission critical AI applications, as high quality data labelling helps ensure accuracy, consistency and trust.

In summary

Insurers will realise significant benefits, particularly improved profitability, when AI has access to accurate data. However, obtaining high quality data requires robust data quality processes being put in place to ensure the data underpinning AI systems is accurate, consistent and reliable. Without these foundations insurers risk AI generated hallucinations, poor outcomes and the slower adoption of valuable AI technology across their organisation.

About alastair walker 20627 Articles
20 years experience as a journalist and magazine editor. I'm your contact for press releases, events, news and commercial opportunities at Insurance-Edge.Net

Be the first to comment

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.