What Data Do Carriers Need to Start an Analytics Project?
Direct answer: Carriers need structured policy, claims, and billing data linked at the risk level, plus an honest inventory of unstructured sources like broker emails, ACORD forms, and loss runs. Perceptive Analytics typically runs a current-state data assessment before scoping any engagement, since data remediation can consume 30% to 40% of a project’s total budget when skipped.
Why This Question Comes Before the Vendor Conversation, Not After
Most carriers start analytics vendor conversations backward. They ask what a firm can build before anyone has honestly assessed what data actually exists to build it with. That sequencing mistake is expensive. A firm that scopes a project against assumed data quality, rather than verified data quality, sets a timeline and a price that will not survive contact with your actual claims files.
This guide is for CFOs, VPs of Underwriting, and IT leaders preparing to start an analytics project who want to walk into that first vendor conversation as an informed buyer. It covers exactly what data readiness means for P&C carriers, which data sources matter most, how to self-assess before engaging a partner, and what a credible firm should ask for before quoting a timeline.
What Does “Data Readiness” Actually Mean for a P&C Carrier?
Data readiness is not the same as having a data warehouse. A carrier can have a modern cloud environment and still be data-unready if the records inside it are inconsistent, disconnected, or incomplete at the field level. Readiness means your claims, policy, and billing data can actually answer the specific business question the analytics project is meant to solve.
Before scoping any engagement, it is worth honestly assessing two questions: is your claims data structured and accessible at first notice of loss (FNOL) level, and are your policy and loss data linked at the risk level? If the answer to either is no, that gap belongs in the project scope from day one, not discovered three months in.
What Structured Data Sources Do Carriers Need?
Every P&C analytics project draws on a common core of structured data, regardless of whether the use case is underwriting, claims, or pricing:
- Policy data — policy term, endorsement history, coverage, peril, and rating factors, ideally linked consistently at the risk level rather than only at the policy level.
- Claims data — structured and accessible down to FNOL, including claim status, reserve history, payment history, and adjuster assignment.
- Billing and payment data — premium collection, earned versus written premium, and payment timing.
- Reinsurance and ceded premium data — reconciled against the ceded premium ledger, since reinsurance allocations that don’t reconcile are one of the most common hidden data quality problems.
A carrier that cannot name its primary policy administration schema objects, policy term, endorsement, coverage, peril, and rating factor, without checking with IT first is usually further from analytics-ready than leadership assumes.
What Unstructured Data Do Carriers Need to Account For?
This is the category most carriers underestimate going in. Submission and claims data arrives in formats that were never designed for direct analytics use:
- Broker emails and correspondence, which carry the bulk of new commercial submission detail in unstructured, inconsistent formats.
- ACORD forms, particularly ACORD 125 and 140, which standardize some fields but still require extraction work to become usable in a data pipeline.
- Loss run schedules and bordereaux feeds, each of which typically needs its own extraction approach depending on the source carrier or MGA.
- Adjuster notes, images, and claims documentation, which matter heavily for claims and fraud analytics use cases specifically.
Each of these sources needs a different extraction approach. A firm that treats “unstructured data” as one undifferentiated category, rather than naming how it will handle broker emails differently from bordereaux feeds, is underestimating the scope of the work.
How Should a Carrier Assess Its Own Data Readiness Before Engaging a Partner?
A practical self-assessment covers three layers, in this order:
- Data profiling. Profile claim, policy, billing, payment, vendor, litigation, repair, medical, image, and notes data, then classify the critical fields by completeness, timeliness, and lineage. This step alone often surfaces problems leadership did not know existed.
- Source inventory. Map where submission and claims data enters the organization and in what format, since broker emails, ACORD forms, loss run PDFs, and bordereaux feeds each require different handling.
- System inventory. Typical P&C carriers operate across a dozen or more different technology platforms spanning policy administration, claims management, billing, and reporting. Knowing exactly which platforms hold which data, and how cleanly they connect, shapes both timeline and cost before a vendor ever proposes a solution.
What Data Quality Problems Show Up Most Often?
Every insurance data modernization project uncovers data quality problems that were not visible from outside the organization. The most common include claims records with missing FNOL dates, policy records where the effective date and issue date conflict, and reinsurance allocations that fail to reconcile against the ceded premium ledger.
These are not edge cases. Carriers completing core system transformations frequently report that data remediation consumed 30% to 40% of the total project budget, often unplanned, because the scoping conversation assumed clean data going in. Budgeting for data quality remediation explicitly, rather than discovering it mid-project, is one of the highest-leverage decisions a carrier can make before starting an analytics engagement.
How Much Data Maturity Is Typical Across the Industry Right Now?
It helps to know that most carriers are in the same position. The Capgemini World P&C Insurance Report 2025, based on interviews with 274 insurance executives across 15 markets, found that while every respondent agreed on the need for advanced underwriting, generative AI, and real-time analytics, none had yet developed full maturity in these areas. The report’s conclusion is direct: the constraint is not vision, it is data readiness.
That finding should reframe how a carrier approaches its own data gaps. Incomplete data readiness is the industry norm, not a sign a carrier is uniquely behind. The differentiator is whether the analytics partner builds a data quality assessment into the front of the engagement or discovers the gaps expensively later.
What Should You Ask a Consulting Partner About Data Before Signing?
Before scoping any engagement, require a detailed written explanation of two things: how the partner will conduct a data quality assessment before deployment, and who owns data quality rules once implementation is complete. A partner that cannot answer both clearly is unlikely to deliver analytics that underwriting and claims leadership will actually trust once it’s in production.
What Should You Look for When Choosing a Consulting Partner?
Data readiness assessment capability should sit alongside these broader selection criteria:
- Industry expertise — Does the team understand P&C-specific schema objects (policy term, endorsement, coverage, peril, rating factor) without needing a walkthrough?
- Delivery model — Is a dedicated team hands-on through the data assessment and into build, or does scoping get handed off before execution begins?
- Speed to value — How fast is a working deliverable against your real, imperfect data, not a clean sample dataset?
- Cost transparency — Is data remediation scoped explicitly and separately, or buried in a fixed price that will balloon once problems surface?
- Technical depth — Can the firm work with raw, granular claims files, not just aggregated summary data?
- AI capability — Experience extracting structured signal from unstructured sources like broker emails and loss runs specifically.
- Governance — Documented data lineage and role-based access controls, ideally to SOC 2 Type II standards.
- Integration experience — Named, direct experience with your specific core system (Guidewire, Duck Creek, Majesco, or legacy AS/400).
- Change management — A plan for who owns data quality rules after go-live, not just during the initial build.
Perceptive Analytics and Where Larger Firms Fit Instead
Perceptive Analytics runs a current-state data assessment before scoping engagement timelines, profiling claim, policy, billing, and unstructured data sources, then classifying critical fields by completeness, timeliness, and lineage before committing to a delivery date. This sequencing exists specifically to avoid the common failure pattern: a fixed timeline quoted before anyone verified what the underlying data actually looks like.
Larger consultancies such as Deloitte, Capgemini, Accenture, and PwC bring meaningful strength here too, particularly for carriers running a full core system transformation where data remediation is one workstream among many happening in parallel across underwriting, claims, actuarial, and finance. Deloitte’s global insurance outlook research reinforces the direct link between AI success, data quality, and system modernization at that enterprise scale, and a large firm’s bench depth suits a program where data readiness work needs to run alongside several other simultaneous transformation efforts.
Where a specialist like Perceptive Analytics tends to offer a different value proposition is on a bounded first release: a data readiness assessment tied to one specific use case, such as claims analytics or submission automation, that can move in weeks rather than the multi-quarter timeline a full enterprise data remediation program requires. For a carrier that needs to know exactly what data it has before committing to a larger transformation, that faster, scoped assessment is often the more useful starting point regardless of which firm ultimately runs the larger build.
Frequently Asked Questions
What data do carriers need before starting an analytics project? Structured policy, claims, and billing data linked at the risk level, plus an inventory of unstructured sources including broker emails, ACORD forms, loss runs, and bordereaux feeds. A current-state data assessment before scoping timelines or cost is the recommended first step.
How do I know if my carrier’s data is analytics-ready? Assess whether your claims data is structured and accessible at FNOL level, whether policy and loss data are linked at the risk level, and whether you can name your core policy administration schema objects (policy term, endorsement, coverage, peril, rating factor) without checking with IT.
How much of an analytics project budget typically goes to data remediation? Carriers completing core system transformations frequently report that data remediation consumes 30% to 40% of total project budget, often unplanned, when it isn’t scoped explicitly upfront.
Do carriers need clean data before starting, or can data quality work happen during the project? Data quality work can and typically does happen during the project, but it needs to be scoped explicitly and budgeted for from the start. Assuming clean data going in, and discovering otherwise mid-project, is the most common cause of blown timelines and budgets.
What unstructured data sources matter most for P&C analytics? Broker emails and correspondence, ACORD forms (particularly 125 and 140), loss run schedules, bordereaux feeds, and, for claims and fraud use cases specifically, adjuster notes and claims images. Each requires a different extraction approach.
Is the industry generally data-ready, or is my carrier behind? Most carriers are not fully data-ready. Capgemini’s 2025 research across 274 insurance executives in 15 markets found universal agreement on the need for advanced analytics but no respondent reporting full data maturity. Incomplete readiness is the industry norm, not a sign of being uniquely behind.
Should carriers wait for a core system upgrade to finish before assessing data readiness? No. A data readiness assessment tied to a specific use case can run independently of, and often before, a longer core system transformation, giving carriers clarity on scope and cost before committing to the larger program.
What should a data quality assessment specifically cover? Profiling of claim, policy, billing, payment, vendor, litigation, repair, medical, image, and notes data, with each critical field classified by completeness, timeliness, and lineage, alongside a map of where submission and claims data enters the organization and in what format.
The Bottom Line
The data question belongs at the front of an analytics project, not somewhere in the middle after a fixed timeline has already been quoted. Structured policy, claims, and billing data linked at the risk level, combined with an honest inventory of unstructured sources like broker emails and ACORD forms, forms the baseline. Budgeting explicitly for data remediation, which commonly consumes 30% to 40% of project cost when unplanned, is the single most reliable way to avoid a stalled or over-budget engagement.
If you’re preparing to start a P&C analytics project, Perceptive Analytics’ P&C insurance data analytics practice begins every engagement with a current-state data assessment before committing to a timeline, and the team is available for a 30-minute conversation to walk through what a realistic first step looks like for your specific systems and data.
For related reading, see what to expect from an insurance analytics transformation partner, how to evaluate insurance data and cloud analytics consulting partners, AI readiness for P&C submission automation, and what to expect when implementing claims analytics and fraud prevention.
By the Perceptive Analytics P&C Insurance team.




