Executive Summary

 

Life sciences organizations are reaching a point where data-platform architecture is becoming a direct constraint on scientific and operational performance. Genomic sequencing, clinical trials, laboratory systems, imaging, real-world evidence, electronic health records, connected devices and commercial operations are all generating more information, but the harder problem is that these datasets are increasingly diverse and need to be used together. One illustration of the scale shift is the NCBI Sequence Read Archive, a genomics repository, which grew from 47.04 GB in May 2007 to 27.93 PB by February 2024, roughly a 620,000× increase, while AstraZeneca now reports more than 1.7 million human genomes with matched clinical insights supporting R&D.

The architectural problem is therefore no longer simply storage capacity. Traditional environments were often organized around individual applications and functions: clinical data lived in clinical systems, laboratory information in LIMS, research data in specialist repositories, and commercial information in separate platforms. That model can work when teams operate largely within their own boundaries. It becomes much harder when a researcher needs to connect genomics with phenotype, a clinical team needs to combine trial data with laboratory and real-world evidence, or an AI application needs governed access to several sources at once. Every additional connection introduces integration work, duplicated transformations, inconsistent definitions and another place where access and lineage have to be managed.

AI has made this weakness more visible. A capable model cannot compensate for data that is difficult to find, poorly described, inconsistently defined or impossible to trace back to an authoritative source. Industry analysis from EPAM similarly frames findability, accessibility, interoperability and reusability as central barriers to AI-driven R&D. The implication is important: AI readiness is increasingly a data-architecture problem. The organizations moving now are not merely lifting existing databases into the cloud. They are creating reusable foundations built around scalable storage, interoperable formats, metadata, lineage, governance, identity and data products.

The strategic goal should not be to eliminate every legacy system. It should be to stop legacy architecture from determining what the organization can do with its data. A modern platform gives different workloads access to a governed foundation while allowing compute and applications to evolve independently. In practice, this means mapping the data estate, prioritizing high-value domains, reducing duplicate pipelines, moving compute closer to data, making metadata and lineage first-class capabilities, and turning trusted datasets into reusable products. The outcome is not simply a newer technology stack; it is a lower-friction operating model in which every new scientific or AI use case can build on capabilities the organization has already created.

620,000×

Approximate growth in NCBI Sequence Read Archive data from 2007 to Feb. 2024.

Koreeda et al., Genes 2025

27.93 PB

Published SRA data volume reported for February 2024.

Koreeda et al., Genes 2025

1.7M+

Human genomes with matched clinical insights reported by AstraZeneca.

AstraZeneca, Centre for Genomics Research

“The value of a modern data platform is not that it stores more information. It is that the next scientific question can reuse the foundation built for the last one.”

The Data Growth Problem Has Become an Architecture Problem

For much of the industry’s digital evolution, data environments were designed around applications rather than around the reuse of data itself. A clinical platform was optimized for clinical operations, a laboratory system for laboratory workflows, a research repository for a particular scientific community and a commercial environment for commercial reporting. Those systems were often individually successful, but their boundaries became expensive when the questions being asked stopped respecting the same boundaries. Modern drug discovery increasingly combines genomic, transcriptomic, proteomic, phenotypic and clinical information. Clinical development increasingly brings together trial data, laboratory results, imaging, biomarkers, external datasets and real-world evidence. The architecture has to support these combinations without forcing every team to build a new integration layer from scratch.

The underlying data also behaves differently. Genomic sequences can be extremely large and computationally intensive. Clinical data may be structured around controlled standards and study lifecycles. Imaging and pathology introduce large unstructured or semi-structured objects. Real-world data may arrive from external partners with different identifiers, refresh schedules and quality characteristics. Even when two systems contain a field called “patient,” “study,” “sample” or “outcome,” the meaning, identifier, ownership and permissible use may differ. Simply adding more storage does not solve this semantic problem. It can make it more difficult by allowing more copies of slightly different information to accumulate.

AstraZeneca’s population-genomics work illustrates the scale of the new operating environment: its Centre for Genomics Research reports more than 1.7 million human genomes with matched clinical insights, alongside multi-omics work spanning transcriptomics, proteomics and metabolomics.

This is why the current rearchitecture cycle is different from earlier infrastructure upgrades. The question is not whether an organization can keep adding capacity to its existing systems. The question is whether the underlying architecture can make growing volumes of increasingly heterogeneous information usable across multiple consumers. A platform that solves that problem can turn data growth into a reusable capability. A platform that does not will continue to absorb growth as another collection of pipelines, extracts, copies and reconciliation tasks.

Figure 1: Moving from fragmented data silos to a governed, reusable data platform

The Hidden Cost of Keeping Data in Silos

The financial impact of fragmented data architecture rarely appears as one large infrastructure invoice. It shows up as repeated work. Data engineers build separate pipelines to move similar information between systems. Scientists create local copies because the authoritative dataset is difficult to find or access. Analysts reconcile definitions before they can start analysis. Data scientists clean and transform the same source data again for another use case. Governance teams have to understand multiple copies, permissions and control mechanisms rather than one clearly managed asset. Each individual task may look reasonable, but the cumulative effect is a data-friction tax that can become a material operating cost.

Consider a clinical-development team preparing an analysis that needs trial data, laboratory results and an external real-world dataset. In a fragmented environment, the team may first identify owners for each source, request access separately, extract or replicate the data, standardize identifiers, reconcile study and patient definitions, document transformations and obtain approvals. If the same combination is needed for another study six months later, much of that work may be repeated because the first project produced an analysis-specific dataset rather than a reusable data product. The organization therefore pays repeatedly for the same data preparation rather than investing once in a governed capability.

The problem becomes even more pronounced when AI enters the picture. EPAM’s analysis of R&D data platforms emphasizes that pharma’s AI challenge is not simply model availability; it is the ability to make data findable, accessible, interoperable and reusable. An AI system can process huge volumes of information, but it still needs to know what a dataset represents, whether the data is current, what transformations were applied, which sources are authoritative and what the permitted use is. Poor metadata and duplicated datasets do not disappear when a model is introduced. They become inputs to a larger system and can multiply the cost of establishing trust in the output.

The distinction between a repository and a platform is therefore important. A repository answers the question, “Where can we store this?” A platform answers a broader set of questions: “Can people discover it? Can they understand it? Can they access it appropriately? Can different engines use it without another copy? Can we trace how it changed? Can the next use case reuse the work?” That difference is often where the business case for rearchitecture becomes strongest. The objective is not to move data for the sake of movement; it is to reduce the amount of work required every time the organization wants to create value from data.

AI Has Made the Data Foundation Strategic

AI has accelerated the urgency because life sciences organizations are moving beyond isolated demonstrations toward AI-assisted workflows across discovery, clinical development and operations. Models can support target identification, biomarker discovery, patient selection, trial optimization, scientific search, document generation and other activities. But the model itself is only one component. A production AI application also needs reliable source data, clear context, traceability, access controls and mechanisms for handling uncertainty. Without those foundations, organizations can demonstrate impressive outputs without being able to deploy them consistently or responsibly at enterprise scale.

A generative AI system answering a scientific question, for example, needs more than access to a large document repository. It needs to distinguish authoritative information from obsolete or duplicated content, understand relationships between datasets and documents, and preserve enough provenance that a scientist can investigate where an answer came from. A predictive model has a similar requirement: training data must have consistent definitions and appropriate historical context, otherwise model performance can be distorted by changes in data capture rather than by the scientific signal the organization is trying to learn.

This is why metadata, catalogs, lineage, semantic definitions, identity and governance are increasingly being treated as core architectural capabilities rather than optional services. EPAM’s R&D framework emphasizes centralized metadata, semantic layers, modular data services and governed experimentation environments precisely because AI requires machine-interpretable context as well as raw data. The architecture needs to make the right data easier for both people and machines to discover and harder to misunderstand.

Figure 2: The data foundation stack, from raw sources to AI-ready capability

Merck provides a practical illustration of the architectural shift: its modern clinical data platform brings fragmented clinical, operational, regulatory and safety data into a unified, metadata-driven environment designed to support analytics and AI.

Why ‘Move Everything to the Cloud’ Is Not Enough

Cloud adoption is an important part of the modernization story, but treating cloud migration as the destination can leave the underlying problem intact. A fragmented collection of databases and applications can remain fragmented after the servers are moved. Teams can still maintain separate pipelines, duplicate datasets and incompatible definitions; only the location of the infrastructure has changed. McKinsey’s life sciences analysis makes a similar point by emphasizing that there is no single migration path: organizations need to weigh domain-specific use cases and decide where migration, remediation or rearchitecture makes the most sense.

A modern platform should instead separate the concerns that have historically been bundled together. Storage should be able to scale independently from compute. Different analytical engines should be able to work against governed data without creating a downstream copy for every consumer. Metadata and lineage should provide a common understanding of what the data means. Identity and access controls should be applied consistently. Applications should be able to evolve without forcing the organization to redesign the underlying data environment each time.

This flexibility matters because life sciences workloads are unusually heterogeneous. Genomic processing may require bursts of distributed compute. Business intelligence may require predictable interactive performance. Machine-learning teams may need specialized environments and accelerators. Clinical systems may have strict availability, validation and lifecycle requirements. A shared data foundation does not mean every team must use the same engine. It means that teams can use the right compute for the job while operating against data that is governed and reusable.

There will also be legitimate reasons for some systems to remain on-premises or in specialized environments. Regulatory requirements, validated processes, technical constraints, acquisition history and organizational readiness all influence the appropriate target state. The objective is therefore not to eliminate legacy technology at any cost. It is to prevent legacy architecture from becoming the default answer to every new data requirement. A system can remain in place while its role becomes narrower, its interfaces become cleaner and the surrounding data estate becomes easier to operate.

The strongest modernization programmes consequently treat cloud as an architectural enabler rather than a finish line. The measure of success is not the percentage of workloads that moved to a hyperscaler. It is whether scientists can find and use trusted data faster, whether duplicated pipelines decline, whether governance becomes more scalable and whether new analytical or AI applications can be delivered without rebuilding the same integration machinery again.

PERCEPTIVE ANALYTICS PERSPECTIVE

At Perceptive Analytics, we recommend starting data-platform modernization with a business and data architecture assessment rather than with a technology decision. The first question should be which scientific or business decisions are currently being slowed down by the way data is stored, accessed, governed or connected. From there, the organization can map critical domains, identify duplicated pipelines and datasets, document movement between systems, assess manual reconciliation and determine which workloads need elastic or specialized compute.

The result is a fact base for deciding what should be rearchitected first, what should be integrated, and what should simply remain as it is. The objective is a reusable foundation in which every high-value initiative leaves behind common ingestion, metadata, identity, quality, governance and data-product capabilities that the next initiative can reuse.

What To Do Instead: A Modern Life Sciences Data Platform Framework

A practical modernization programme should not begin with a wholesale replacement of existing systems. It should begin by defining the role each layer of the data ecosystem needs to play. The framework below focuses on five capabilities that compound over time: understanding the estate, creating a scalable foundation, making data understandable, embedding governance and exposing trusted data as reusable products. This approach allows leadership to modernize around measurable friction while protecting workloads that still depend on specialized or validated environments.

Figure 3: A modern life sciences data platform framework

01 Map the Data Estate and Prioritize High-Value Domains

The first step is to establish a clear view of where critical data lives, who owns it, how it is transformed, what quality issues exist and which scientific or business processes depend on it. This is more difficult than producing an inventory of applications because the same concept can be represented differently across systems. A patient, study, compound, sample or outcome may have multiple identifiers, definitions and ownership models. A domain-level map makes those differences visible before they are reproduced in a new platform. It also allows leadership to prioritize modernization based on business impact: a dataset that supports dozens of studies and multiple AI use cases deserves a different investment profile from a specialized source used by one low-change process.

02 Build a Scalable and Reusable Data Foundation

The second step is to establish storage and processing capabilities that can accommodate structured, semi-structured and unstructured information without forcing every workload into one physical model. Modern lakehouse approaches are useful because storage and compute can evolve independently and multiple analytical engines can access common data. The objective is not to build one giant repository. It is to reduce unnecessary duplication and make the underlying data available to the workloads that need it. When a new clinical or research use case can reuse the same governed data layer, the economics change: the organization is extending an existing capability rather than funding another isolated stack.

03 Make Metadata, Lineage and Interoperability First-Class

Storage alone does not make data reusable. Scientists and AI systems need to know what a dataset represents, where it came from, how it was transformed, what its limitations are, who owns it and whether it can be used for a particular purpose. A modern metadata layer should therefore cover cataloging, lineage, scientific and business definitions, identifiers, ownership and usage context. Interoperable formats and standardized interfaces matter just as much because they reduce the number of custom pipelines required when another analytical engine or application needs the same information. The practical result is a platform where data becomes easier to discover and harder to misunderstand.

04 Design Governance Into the Data Flow

In life sciences, governance cannot be treated as a final review step. Patient information, clinical data, genomic information, intellectual property and regulated records can all have different access and retention requirements. The platform therefore needs identity, role-based access, masking, auditability, quality controls and lineage to operate as part of the normal data flow. This is a meaningful operating-model change. Instead of recreating compliance controls for every new project, teams can inherit approved controls from the platform. Governance becomes scalable infrastructure rather than a manual queue that grows every time another dataset is introduced.

05 Create Reusable Data Products for Analytics and AI

The final step is to expose trusted, well-documented datasets as reusable data products rather than leaving each analytics team to start from raw sources. A clinical-trial data product might combine standardized study information, laboratory data, metadata and quality indicators in a form that supports trial analytics, safety analysis and AI-enabled workflows. A genomics data product could connect genomic observations with phenotype or clinical context while preserving provenance and permissions. This changes the economics of scale: scientists spend less time rebuilding data preparation and more time applying domain expertise to the scientific question. Each new domain can also strengthen shared metadata, identity, governance and integration capabilities.

Case Study: Medidata’s Move to a Modern Lakehouse Architecture

Medidata offers a useful example of why rearchitecture is different from simply adding another data store. In its AWS case study, the company describes a next-generation platform intended to serve thousands of clinical trials worldwide. Its earlier environment had data pipelines and copies distributed across technologies and storage systems, which meant that access controls and operational complexity had to be managed in multiple places. As data volumes increased, the organization also faced growing observability and root-cause-analysis challenges.

The rearchitecture centered on Apache Iceberg as a common data layer, with Apache Flink handling distributed stream processing, Kafka providing the streaming backbone and the AWS Glue Data Catalog providing centralized metadata. Instead of creating consumer-specific downstream copies, the new design allows multiple consumers to work from a consistent view of data. This is a significant architectural shift because it attacks the repeated integration work at its source: the platform becomes the reusable layer rather than every downstream consumer becoming a new pipeline project.

The reported results illustrate why this matters operationally. Medidata says pipeline latency was reduced from numerous hours and sometimes days to minutes, helping customers realize a reported 99% performance gain from ingestion to analytics. It also reports that Iceberg’s interoperability enabled a single copy of data to satisfy multiple use cases, allowing the team to focus maintenance and observability on a subset of operations that was five times smaller than before. These figures are vendor-reported outcomes from Medidata’s case study rather than universal benchmarks, but the underlying lesson is broadly relevant: reducing copies and consumer-specific integrations can improve both performance and the amount of engineering effort required to keep the platform running.

The security dimension is equally important. Medidata describes how the architecture reduced the need to build custom access-control layers at each integration point, instead centralizing authorization through familiar cloud identity mechanisms. The data remained within the company’s virtual private cloud. For life sciences organizations, that combination of interoperability and governance is critical. A modern platform only creates value if it can make data easier to use without making the organization less confident about who can use it, what they are using or how it was produced.

Figure 4: Medidata case study, from fragmented pipelines to a unified lakehouse architecture

Proof Points: What the Industry Is Already Doing

AstraZeneca illustrates the scale side of the equation. Its Centre for Genomics Research now reports more than 1.7 million human genomes with matched clinical insights, with the dataset spanning a broad international collaboration network. The company also describes population multi-omics work that layers transcriptomics, proteomics and metabolomics onto genomic information. At this scale, data access and computational efficiency are not peripheral IT concerns; they directly influence how quickly scientists can turn population-scale observations into hypotheses, targets and biomarkers.

Merck illustrates the clinical-development side. Its modern clinical data platform brings clinical, operational, regulatory and safety information into a unified, metadata-driven environment and adds standardization, automated third-party ingestion, centralized metadata, role-based access and self-service analytics. The platform is intended to support trial planning, patient enrollment, site performance and risk-based monitoring, while also creating a foundation for predictive modeling, protocol optimization, synthetic control arms and generative AI. In related generative-AI work described by the company, an internal platform reduced the time required to create a fully human-reviewed first draft of a clinical study report from an average of 180 hours to 80 hours, while reducing errors in the draft by 50% across specified categories. These figures are company-reported outcomes rather than universal productivity benchmarks, but they illustrate how governed data, domain expertise, model capabilities and redesigned processes can work together as a repeatable capability.

Moderna provides an earlier but still useful example of the business value of cloud-enabled data and analytics. McKinsey reported that Moderna was able to deliver its first clinical batch of a COVID-19 vaccine candidate to the U.S. National Institutes of Health for a Phase I trial 42 days after the initial sequencing of the virus, with cloud technology helping the company move rapidly without rebuilding infrastructure for each new requirement. The example is not a claim that cloud alone creates speed. It demonstrates what becomes possible when scalable infrastructure, data access and operating processes are aligned around a time-critical scientific objective.

Together, these examples point to the same architectural direction from different angles. AstraZeneca demonstrates the scale and diversity of modern biological data. Merck demonstrates the need for governed, reusable clinical data. Medidata demonstrates how a shared data layer can reduce integration and maintenance overhead. Moderna demonstrates the value of scalable infrastructure when the organization needs to move quickly. None of these cases argues for one universal technology stack. They show why the architecture must become flexible enough to support different workloads while preserving a common foundation of trust.

From Data Migration to Data Capability

The shift from migration to capability also changes what the data-platform team is expected to provide. Instead of functioning primarily as infrastructure operators, the team becomes a provider of reusable capabilities: governed data products, metadata services, standardized interfaces, quality controls, identity services and analytical access patterns.

This operating model makes the platform team an internal enabler rather than a queue for infrastructure requests. Scientific and business teams can focus on the questions they need to answer while inheriting technical and governance capabilities that have already been established. The practical test is whether the platform team can make the next data source, analytical workload or AI use case easier to onboard without rebuilding the same foundations.

Leadership should therefore resist making the technology label the centre of the strategy. Whether the implementation is described as a lakehouse, warehouse, data lake, mesh or another architecture is less important than whether the platform team is building reusable capabilities that reduce friction between data and decisions. Modernization becomes successful when each investment strengthens a common foundation and makes the next investment easier.

Conclusion

Life sciences organizations are not rearchitecting their data platforms simply because on-premises infrastructure has become unfashionable. They are doing it because the nature of the data problem has changed. Data volumes are expanding rapidly, scientific datasets are becoming more diverse, research increasingly crosses functional boundaries and AI is raising the value of proprietary information. The rapid growth of genomics data is one illustration of that scale shift, but the architectural challenge extends across the broader life sciences data estate.

But storage capacity is only the visible part of the challenge. The harder question is whether organizations can find, connect, govern and use their data quickly enough to support the next generation of scientific and business decisions. Fragmented platforms create a recurring tax through duplicate pipelines, repeated data preparation, inconsistent definitions, manual reconciliation and project-specific controls. Moving those same silos into the cloud may improve infrastructure, but it does not remove the underlying friction.

The practical response is therefore a shift from migration to capability building. Map the data estate. Prioritize the domains where friction is constraining high-value decisions. Build a scalable foundation that can serve multiple workloads. Make metadata, lineage and interoperability first-class capabilities. Embed governance into the data flow. Expose trusted information as reusable data products. Protect the specialized systems that still serve a genuine regulatory or operational purpose, but stop allowing those systems to dictate the architecture of everything around them.

For leadership, the decision is moving from “How do we store more data?” to “How do we build an architecture that allows our data to create more value?” The organizations that answer that question well will not necessarily be the ones that migrate the most systems. They will be the ones that create a data foundation capable of supporting multiple data types, analytical engines and AI applications while preserving governance, provenance and scientific trust. In that model, rearchitecture is not an IT refresh. It is an investment in the organization’s ability to turn data into scientific progress.

Perceptive Analytics partners with life sciences organizations to modernize data ecosystems through scalable data engineering, advanced analytics and AI-enabled solutions. By helping organizations integrate fragmented data, establish governed data foundations and build reusable analytical capabilities, we enable teams to turn growing data volumes into a strategic asset for research, development and operations.

“The value of a modern data platform is not that it stores more information. It is that the next scientific question can reuse the foundation built for the last one.”

References & Sources


Submit a Comment

Your email address will not be published. Required fields are marked *