Executive Summary

Healthcare organizations are entering a data growth cycle that is materially different from the one that shaped many of today’s analytics environments. Electronic health records, claims, clinical trials, medical imaging, pathology, genomics, remote monitoring, connected devices and real-world evidence are all adding data, but they are not adding it in the same shape or on the same schedule. An industry estimate cited by IntuitionLabs projects healthcare data volume to grow by roughly 36% per year over a five-year period. The underlying L.E.K. analysis projected global healthcare data to rise from about 2,300 exabytes in 2020 to 10,800 exabytes in 2025. The 36% figure is therefore best understood as a cited industry projection, not as a measured 2026 growth rate. [1][2]

The more revealing signal is what is happening inside genomics. The NCBI Sequence Read Archive grew from 47.04 GB in May 2007 to 27.93 PB by February 2024, an increase of approximately 620,000 times. The point is not that every healthcare dataset will grow at anything like that rate. The point is that modern healthcare data contains workloads whose scale, structure and computational requirements can change faster than the architecture around them. [3]

Most legacy analytics environments were designed around individual applications or functional domains. Clinical systems served clinical operations. Claims platforms served reimbursement workflows. Research repositories served scientists. Commercial warehouses served reporting teams. Imaging systems managed large files separately from structured records. These systems can each work well within their original boundaries, yet the boundaries become expensive when a business question crosses them. Repeated integration, duplicate extracts, semantic reconciliation and separate governance controls create an analytics friction tax that rises as data grows.

AI has made the problem harder to hide. A model can process information quickly, but it cannot make a poorly described dataset trustworthy. Production AI needs clear definitions, reliable source data, permissions, provenance and a way to distinguish current authoritative information from duplicates or obsolete records. The organizations rearchitecting now are therefore not simply buying more storage. They are building a foundation in which storage, compute, metadata, identity, governance and reusable analytical products can scale together.

36%

Cited annual growth projection for healthcare data volumes over a five-year period.

L.E.K. analysis, later cited by IntuitionLabs

10,800 EB

Projected global healthcare data volume for 2025 in the L.E.K. analysis.

L.E.K. Consulting

620,000x

Approximate growth in NCBI SRA data from May 2007 to Feb. 2024.

Koreeda et al., Genes 2025

 

“The challenge is no longer simply storing more healthcare data. It is making the next analytical workload easier to build than the last one.”

The 36% Growth Problem Has Become an Architecture Problem

Healthcare data growth is often described as a storage problem because storage is the easiest part to see. More records, more images, more genomic files and more device events appear to imply that the organization simply needs a larger data lake or a larger warehouse. That framing is incomplete. The real problem is that data growth increases the number of relationships the architecture must maintain. A thousand additional files can be cheap to store while still being expensive to classify, catalog, secure, move and reconcile. A new clinical data source can be technically easy to ingest while creating weeks of work for analysts who now have to understand a new identifier system and determine how it differs from the patient or provider entities already used elsewhere.

The underlying data is also becoming more heterogeneous. Structured transactions can sit beside semi-structured JSON, high-resolution imaging, laboratory measurements, genomic sequences, physician notes and streaming device telemetry. The same organization may need interactive SQL for management reporting, distributed processing for genomics, specialized environments for machine learning and controlled access for clinical or regulatory workloads. Treating all of those data types as if they were conventional relational tables is likely to create unnecessary copying or transformations.

This is why the architecture problem is fundamentally about data movement and reuse. Every time a dataset is moved to a different platform for a different consumer, the organization has another opportunity to change its definition, lose metadata, introduce latency or create a new security boundary. Growth multiplies the number of these movements. A scalable architecture must therefore reduce the amount of movement required to answer recurring questions while allowing the computational layer to change as workloads change.


Figure 1: The data growth problem in two numbers

Sources: L.E.K. Consulting [1]; Koreeda et al., Genes 2025 [3].

The Hidden Cost of Scaling the Wrong Architecture

The financial impact of rapid data growth rarely appears as one dramatic infrastructure bill. It shows up as repeated work. Data engineering teams build more ingestion pipelines. Analysts create extracts because the governed environment is difficult to query at the right level of detail. Scientists keep local copies because permissions or compute limits make the official source hard to use. Governance teams have to document multiple copies of the same information. Data scientists clean and transform the same source data again because the first project produced an analysis-specific output rather than a reusable asset. Each task can be rational in isolation. Together, they create an operating model in which a growing share of the data budget is spent on making old work available again.

IQVIA has highlighted this problem in its health data transformation work. The company has cited an estimate that 97% of healthcare data goes unused after initial creation and has described a typical healthcare organization as having patient data distributed across as many as 15 disparate systems. These are vendor-reported figures and should not be treated as universal benchmarks, but they illustrate the structural problem: volume does not translate into value when the organization cannot reliably find, standardize and connect the information it already owns or has access to. [4]

A clinical-development team makes the problem easy to visualize. Suppose the team wants to combine trial data, laboratory results, medical imaging and a real-world dataset. In a fragmented environment, the work begins with finding owners, requesting access, extracting or replicating data, harmonizing identifiers, reconciling study definitions, documenting transformations and securing approvals. The analysis itself comes after that preparation. If another study needs a similar combination six months later, the team may repeat most of the exercise because the first project created a one-off analytical dataset. The organization has effectively paid for the same data preparation twice.

The architecture becomes more expensive still when every downstream system creates its own copy. Copies can be useful for legitimate reasons, including performance, isolation or regulatory requirements. The issue is uncontrolled duplication. A modern analytics foundation should make the default path a governed shared asset, with copies created intentionally when there is a clear reason and an explicit lifecycle for those copies.

Figure 2: Why traditional analytics architectures strain under data growth

Conceptual model based on recurring integration, duplication and governance patterns described in [2], [4] and [5].

Why More Data Requires a Different Analytics Foundation

The answer to higher data volume is not to move every dataset into one giant warehouse. Nor is it to create an unrestricted data lake and assume that users will somehow find their way through it. The more durable model separates the concerns that legacy architectures often bundle together. Storage should scale independently from compute. Multiple analytical engines should be able to work against governed data where appropriate. Metadata should describe the data in a way that both people and machines can use. Identity and access controls should operate consistently. And the outputs most likely to be reused should be exposed as analytical products rather than left as project-specific transformations.

This is where lakehouse approaches and open table formats have gained attention in life sciences. IntuitionLabs describes a lakehouse as a way to combine scalable object storage and flexible data formats with warehouse-style capabilities such as transactional reliability, schema management and governance. Apache Iceberg is particularly relevant because it can provide a table abstraction over open storage and allow multiple compatible engines to work against the same underlying data. The practical benefit is less about the label of the architecture and more about reducing the number of physical copies needed when different workloads need access to the same information. [5]

The architecture also needs to acknowledge that not all workloads have the same performance profile. A clinical dashboard may need predictable interactive SQL. Genomic processing may need large bursts of distributed compute. A machine-learning workflow may require specialized libraries or accelerators. A partner-facing use case may require tightly scoped sharing. A modern foundation should therefore provide common governance and metadata without forcing all workloads into one computational shape.

AI Has Turned the Data Foundation Into a Strategic Capability

AI has accelerated the need for better data architecture because AI systems amplify both the value and the weaknesses of the underlying data. A generative system answering a scientific question needs to distinguish authoritative information from duplicated or obsolete content. A predictive model needs stable definitions and historical context so that changes in data capture do not get mistaken for changes in the phenomenon being modeled. An automated workflow needs identity, permissions and an audit trail so the organization can understand what information was used and under which controls.

EPAM’s life sciences R&D analysis makes a similar point by emphasizing findability, accessibility, interoperability and reusability as prerequisites for AI-enabled research. It also highlights centralized metadata, semantic layers and modular data services as part of the data foundation required for scalable AI. The implication is straightforward: the data platform has become part of the AI product. Model performance cannot be separated from the quality, context and governance of the data that feeds the model. [6]

Merck offers a practical illustration from clinical development. Its modern clinical data platform brings clinical, operational, regulatory and safety information into a unified, metadata-driven environment designed to support analytics and AI. In related generative AI work, Merck reports that the time required to create a fully human-reviewed first draft of a clinical study report fell from an average of 180 hours to 80 hours, while errors in specified categories were reduced by 50%. Those are company-reported results from defined use cases, not universal productivity benchmarks, but they show how the data layer and the AI layer become more valuable when they are designed together. [7][8]

Why “Move Everything to the Cloud” Is Not Enough

Cloud adoption can provide elasticity, faster provisioning and a more flexible cost model, but it does not automatically remove fragmentation. A company can migrate a collection of on-premises databases into the cloud and still have the same domains, copies, pipelines and incompatible definitions. The infrastructure has moved; the operating model has not. McKinsey has argued that cloud value in life sciences is not limited to infrastructure efficiency. The larger opportunity comes from enabling analytics, standardization, automation and business innovation, with migration and rearchitecture decisions made at the workload or domain level. [9]

This distinction matters in healthcare because some systems have legitimate reasons to remain specialized. Validated processes, regulatory requirements, contractual constraints, acquisition history and technical dependencies all affect the appropriate target state. A modern architecture therefore does not require an immediate retirement of every legacy platform. It requires the surrounding data ecosystem to become more deliberate. Legacy systems can remain in place while their role becomes narrower, their interfaces become cleaner and their data becomes easier to discover and govern.

The stronger measure of cloud modernization is therefore not the percentage of workloads moved. It is whether data can be reused more often, whether governance can be applied more consistently, whether new analytical workloads can be onboarded without creating another integration project and whether the platform team can make the next use case easier than the previous one.

PERCEPTIVE ANALYTICS PERSPECTIVE

At Perceptive Analytics, we recommend approaching healthcare data growth as a data architecture and analytics operating-model problem rather than a storage procurement exercise. The starting point is a domain and workload assessment: map where critical data lives, how much of it is reused, how identifiers and definitions differ, which pipelines are duplicated, what governance controls exist, which workloads need elastic or specialized compute and where scientists or analysts lose time before analysis begins. From that fact base, the organization can decide which domains should become shared analytical foundations, which datasets should be treated as reusable products, which workloads should move, and which systems should remain specialized. The technology decision follows the workload. Snowflake, Databricks, open table formats such as Iceberg and existing platforms can all play a role, but the objective is the same: reduce data movement, make trust reusable and ensure that each new analytical or AI initiative builds on capabilities the organization already paid to create.

What a Scalable Analytics Framework Changes

A conventional modernization program often reports infrastructure milestones: terabytes migrated, servers retired, pipelines rebuilt or cloud accounts provisioned. Those metrics matter, but they do not prove that the organization can handle growth better. A scalable analytics framework measures whether the amount of work required to turn new data into a trusted decision is falling.

That means tracking the number of domains with clear ownership, the number of duplicate pipelines retired, the proportion of high-value data with usable metadata and lineage, the time required to onboard a new source, the number of analytical products that can reuse the same foundation and the amount of manual reconciliation still required between systems. These metrics connect architecture to business friction. They also create a more durable investment case because each shared capability can be measured by how many downstream workloads it supports.

What To Do Instead: A Scalable Healthcare Analytics Framework

A practical modernization program should begin with the role each layer of the healthcare data ecosystem needs to play. The framework below focuses on five capabilities that compound over time: understanding the estate, building a reusable foundation, making data understandable, embedding governance and exposing trusted data for analytics and AI. The goal is not to create a single monolithic platform. It is to create an architecture in which growth adds reusable capability instead of another isolated stack.



Figure 3: A scalable healthcare analytics architecture

Conceptual framework developed for this article using principles described in [5], [6] and [9].

01 Map the Data Estate and Prioritize High-Value Domains

The first step is to establish a working map of where critical data lives, who owns it, how it is transformed, which consumers use it and where duplicate copies exist. This should go beyond an application inventory. A patient, provider, study, sample or product can have different identifiers and definitions across systems. A domain-level view exposes those differences before they are reproduced in a new platform. It also lets leadership prioritize by impact. A dataset supporting dozens of analytical workloads and multiple AI initiatives deserves a different investment profile from a specialized source used by one stable process.

02 Build a Scalable and Reusable Foundation

The second step is to establish storage and processing capabilities that can accommodate structured, semi-structured and unstructured data without forcing every workload into one physical model. Modern lakehouse approaches are useful because storage and compute can evolve independently and multiple engines can work with governed shared data. The objective is not to build one enormous repository. It is to make high-value information available to the workloads that need it while minimizing unnecessary duplication. When the next clinical or analytics use case can reuse the same foundation, the economics of modernization improve because the organization is extending an existing capability rather than funding another isolated stack.

03 Make Metadata, Lineage and Interoperability First-Class

Storage does not make data reusable. Users need to know what a dataset represents, where it came from, which transformations were applied, what its limitations are, who owns it and whether it can be used for a particular purpose. A modern metadata layer should cover cataloging, lineage, business and scientific definitions, identifiers, ownership and usage context. Interoperable formats and standardized interfaces also matter because they reduce the number of custom pipelines required when another analytical engine or application needs the same information. The practical result is a platform where data becomes easier to discover and harder to misunderstand.

04 Design Governance Into the Data Flow

Healthcare data governance cannot be a final approval step. Patient information, clinical records, genomic information, intellectual property and regulated reporting data may all require different access, masking, retention and audit controls. The platform should therefore integrate identity, role-based access, data classification, masking, auditability and quality rules into the normal data flow. This changes governance from a manual queue into reusable infrastructure. An approved control can be inherited by another product instead of being reinvented for every project.

05 Create Reusable Data Products for Analytics and AI

The final step is to expose trusted, well-documented data as reusable products instead of leaving each analytics team to start from raw sources. A clinical-trial product could combine standardized study data, laboratory information, metadata and quality indicators in a form that supports trial analytics and AI-enabled workflows. A patient or RWE product could unify de-identified longitudinal information with clear provenance and usage constraints. A manufacturing product could bring process, quality and equipment information together for real-time analytics. The point is to make the next use case consume an existing capability rather than rebuild the preparation layer.

Case Study: Medidata’s Move to a Modern Lakehouse Architecture

Medidata provides a useful example of what happens when a life sciences organization treats architecture as a shared capability rather than another data store. In its AWS case study, Medidata describes a next-generation platform designed to serve thousands of clinical trials worldwide. Its earlier environment had pipelines and data copies distributed across technologies and storage systems, which meant that access controls and operational complexity had to be managed in multiple places. Growing data volumes also increased the effort required for observability and root-cause analysis. [10]

The rearchitecture centered on Apache Iceberg as a common data layer, with Apache Flink for distributed stream processing, Kafka for the streaming backbone and AWS Glue Data Catalog for centralized metadata. The design allowed multiple consumers to work from a consistent view of data rather than creating a separate downstream copy for each consumer. This is important because it attacks the repeated integration problem at its source. The platform becomes the reusable layer, while the consumer becomes an additional use case rather than an entirely new pipeline project.

The reported outcomes illustrate the operating impact. AWS and Medidata say pipeline latency was reduced from numerous hours and sometimes days to minutes, helping customers realize a reported 99% performance gain from data ingestion to analytics. They also report that Iceberg interoperability enabled a single copy of data to satisfy multiple use cases, allowing the team to focus maintenance and observability on a subset of operations that was five times smaller than before. These are vendor and customer reported results from a specific architecture, not universal benchmarks. The broader lesson is that reducing consumer-specific copies can lower both the amount of data movement and the operational surface area that engineers need to maintain. [10]

The governance dimension is equally important. Medidata describes reducing the need to build custom access-control layers at each integration point by using centralized cloud identity mechanisms, while keeping the data within its virtual private cloud. For healthcare and life sciences, that combination matters. A scalable analytics architecture only creates value when it can increase data availability without making the organization less confident about who can access the data, what they are using or how the data was produced.



Figure 4: Medidata case study, from fragmented pipelines to a unified lakehouse architecture

Source: AWS / Medidata case study [10]. Reported figures are specific to the described implementation.

Proof Points: What the Industry Is Already Doing

AstraZeneca illustrates the scale side of the equation. Its Centre for Genomics Research reports more than 1.7 million human genomes with matched clinical insights and more than 680,000 genomes from understudied global communities. The company also reports that human genetics research has supported more than 80 pipeline decisions since 2017 and that more than 310 AstraZeneca clinical trials are contributing to its genomic datasets. [11] The significance for architecture is straightforward. When population-scale genomic data is combined with clinical context and multi-omics information, the data environment becomes part of the scientific capability itself. Researchers need to move from raw observations to hypotheses without repeatedly rebuilding the underlying data preparation.

IQVIA illustrates the data-utilization side. Its Health Data Transformation Platform is designed to standardize and integrate heterogeneous healthcare information, including structured and unstructured data, while incorporating de-identification and privacy-preserving methods. IQVIA has argued that healthcare organizations spend a large share of their effort cleaning and wrangling data before analysis. Whether an individual organization matches those percentages is less important than the recurring pattern: if preparation consumes most of the available analytical capacity, growth in raw data can produce growth in workload without a corresponding increase in insight. [4]

Merck illustrates the clinical-development side. Its modern clinical data platform is described by AWS as a unified, metadata-driven environment that brings clinical, operational, regulatory and safety data together and supports standardization, automated third-party ingestion, security controls, centralized metadata and self-service analytics. The platform is intended to support trial planning, patient enrollment, site performance and risk-based monitoring while establishing a foundation for predictive modeling and generative AI. [8]

Taken together, these examples point in the same direction. AstraZeneca demonstrates the scale and diversity of modern biological data. IQVIA demonstrates the cost of fragmented data preparation. Merck demonstrates the value of a governed clinical data foundation. Medidata demonstrates how a shared data layer can reduce integration and maintenance overhead. None of these cases implies that every organization needs an identical technology stack. They show why the architecture needs to make data easier to reuse as volume, variety and analytical demand increase.

From Data Growth to Data Capability

The move from data growth management to data capability also changes the role of the data-platform team. Instead of operating primarily as infrastructure administrators, the platform organization increasingly becomes a provider of reusable services: governed data products, metadata and cataloging, standardized ingestion, identity, quality controls, interoperable storage formats and analytical access patterns.

This is an important shift because the value of a platform compounds only when new work becomes easier. If every new source requires a new security model, every dashboard requires a custom transformation, and every AI application needs another copy of the same clinical dataset, the organization is scaling infrastructure without scaling capability. A mature platform should make the next source easier to onboard, the next use case easier to launch and the next governance review easier to complete.

That is also why architecture labels matter less than architecture behavior. Whether the organization describes its environment as a lakehouse, warehouse, data lake, data mesh or hybrid model is secondary. What matters is whether the environment can absorb new data while reducing duplication, preserving provenance, supporting multiple analytical workloads and making trusted information reusable.

Conclusion

Healthcare data volumes are growing rapidly, but the bigger challenge is that the data is becoming more diverse, more interconnected and more important to analytical and AI workloads. The cited 36% annual growth figure captures the scale of the trend, while the roughly 620,000-fold increase in NCBI’s Sequence Read Archive shows how quickly some biological datasets can expand. The architectural challenge is therefore not just keeping up with storage. It is keeping up with the relationships, controls and analytical demands that come with that growth.

Traditional architectures absorb growth through more pipelines, more copies and more reconciliation. That can work for isolated workloads, but it becomes increasingly expensive when the same information needs to support clinical development, research, commercial analytics, real-world evidence and AI. Moving those silos to the cloud may improve infrastructure, but it does not remove the underlying friction.

The practical response is to build capability rather than simply migrate infrastructure. Map the data estate. Prioritize high-value domains. Build a scalable shared foundation. Make metadata, lineage and interoperability first-class. Embed governance into the data flow. Expose trusted information as reusable data products. Protect specialized systems where they still serve a genuine business, scientific or regulatory purpose, but stop allowing them to define the architecture of everything around them.

For leadership, the question is moving from ‘How do we store more healthcare data?’ to ‘How do we make every new unit of data easier to discover, govern and use?’ The organizations that answer that question well will not necessarily be the ones that migrate the most systems. They will be the ones that create a data foundation capable of supporting more workloads without proportionally increasing the amount of integration and reconciliation required to use it.

Perceptive Analytics helps healthcare and life sciences organizations modernize data ecosystems through scalable data engineering, cloud platforms, advanced analytics and AI. By integrating fragmented data, establishing governed foundations and building reusable analytical capabilities, organizations can turn growing data volumes into a durable analytics capability.

“The goal of a scalable healthcare data platform is not to store every new byte more cheaply. It is to make the next question easier to answer because the foundation for it already exists.”

References & Sources

  • L.E.K. Consulting, “Tapping Into New Potential: Realising the Value of Data in the Healthcare Sector,” 2023. Reports the projection that global healthcare data would increase from about 2,300 exabytes in 2020 to 10,800 exabytes in 2025, equivalent to approximately 36% annual growth. https://www.lek.com/sites/default/files/PDFs/tapping-new-potential.pdf
  • IntuitionLabs, “Pharma R&D Data Lakehouses: Databricks, Snowflake & Iceberg,” 2026. Summarizes the healthcare data-growth estimate and explains lakehouse architecture using Databricks, Snowflake and Apache Iceberg in life sciences. https://intuitionlabs.ai/articles/pharma-data-lakehouse-databricks-snowflake-iceberg
  • Koreeda, T., Honda, H., & Onami, J.-i., “Snowflake Data Warehouse for Large-Scale and Diverse Biological Data Management and Analysis,” Genes, 2025. Documents NCBI SRA growth from 47.04 GB in May 2007 to 27.93 PB in February 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11765040/
  • IQVIA, “Automate Data Transformation to Accelerate Healthcare Insights,” 2022. Describes healthcare data fragmentation, the company’s reported 97% unused-data estimate and its Health Data Transformation Platform. https://www.iqvia.com/blogs/2022/11/automate-data-transformation-to-accelerate-healthcare-insights
  • IntuitionLabs, “Pharma R&D Data Lakehouses: Databricks, Snowflake & Iceberg,” 2026. Covers lakehouse architecture, open formats, ACID transactions, governance and interoperability in life sciences. https://intuitionlabs.ai/articles/pharma-data-lakehouse-databricks-snowflake-iceberg
  • EPAM, “R&D Revolution in Life Sciences: Designing Data Platforms to Enable AI,” August 2025. Covers data silos, FAIR data, metadata, semantic layers, modular data services and AI-ready R&D platforms. https://www.epam.com/insights/blogs/r-and-d-revolution-in-life-sciences-designing-data-platforms-to-enable-ai
  • Merck, “Merck Expands Innovative Internal Generative AI Solutions Helping to Deliver Medicines to Patients Faster,” June 2025. Reports the reduction from 180 hours to 80 hours for a fully human-reviewed clinical study report first draft and a 50% reduction in specified error categories. https://www.merck.com/news/merck-expands-innovative-internal-generative-ai-solutions-helping-to-deliver-medicines-to-patients-faster/
  • AWS, “Highlights from the 2025 AWS Life Sciences Symposium’s Clinical Trials track,” June 2025. Describes Merck’s modern clinical data platform, metadata-driven architecture, governance, analytics and AI foundations. https://aws.amazon.com/blogs/industries/highlights-from-the-2025-aws-life-sciences-symposiums-clinical-trials-track/
  • McKinsey, “The Case for Cloud in Life Sciences,” October 2021. Discusses domain-based cloud adoption, analytics, standardization and migration versus remediation or rearchitecture. https://www.mckinsey.com/industries/life-sciences/our-insights/the-case-for-cloud-in-life-sciences
  • AWS, “Medidata’s journey to a modern lakehouse architecture on AWS,” November 2025. Describes Medidata’s Apache Iceberg architecture, streaming pipelines, unified clinical data view and reported performance improvements. https://aws.amazon.com/blogs/big-data/medidatas-journey-to-a-modern-lakehouse-architecture-on-aws/
  • AstraZeneca, Centre for Genomics Research / Genomics. Reports more than 1.7 million human genomes with matched clinical insights, 680,000+ genomes from understudied global communities, 80+ pipeline decisions supported and 310+ clinical trials fueling genomic datasets. https://www.astrazeneca.com/r-d/science-and-technologies/genomics.html

Submit a Comment

Your email address will not be published. Required fields are marked *