Executive Summary

Statistical Analysis System, commonly referred to as Statistical Analysis Software and abbreviated as ‘SAS’ remains deeply embedded in life sciences because it solves problems that are unusually difficult to solve in regulated environments. Clinical statistical programming, statistical analysis, table/listing/figure generation, regulatory submissions and other controlled workflows depend on mature statistical capabilities, validation practices, auditability and decades of institutional knowledge. For many organizations, the question is therefore not why SAS has survived; it is why a platform that remains operationally valuable can also become increasingly expensive to maintain as the rest of the enterprise moves elsewhere.

The answer begins with architecture. Life sciences organizations increasingly use cloud platforms such as Snowflake and Databricks for enterprise data, large-scale processing and AI/ML, while Python and SQL have become central to modern analytics engineering and data science. SAS does not necessarily disappear from this environment. Instead, it increasingly operates alongside these platforms. SAS’s expanded support for Databricks Spark in 2025 and its broader integration direction around cloud data platforms reflect this shift. The result is a hybrid estate in which SAS may continue to run the regulated analytics that depend on it while other workloads increasingly live on modern cloud infrastructure.

That coexistence changes the economics. The visible cost is the SAS license, but the larger operating cost can sit in the layers around it: duplicated data, SAS-specific transformations, interfaces, validation of legacy applications, specialist administration, hard-to-replace talent, manual handoffs and the effort required to keep multiple analytical environments consistent. None of these costs necessarily appears as a line item called ‘SAS cost’. They emerge as slower change, additional engineering work and growing dependency on people who understand both the legacy environment and the regulatory context.

The strategic implication is not that life sciences companies should simply ‘get off SAS’. A wholesale migration can itself create regulatory, operational and financial risk. The more useful question is: where does SAS provide unique value, and where is the organization maintaining SAS simply because it has always been there? The right modernization strategy protects the regulated core, identifies redundant layers, reduces unnecessary data movement, modernizes selectively and builds a workforce capable of operating across SAS, Python, SQL and cloud platforms. In that model, modernization becomes an exercise in reducing analytics debt and improving the economics of the estate rather than replacing one platform with another.

$1,500 → Millions

Range of SAS licensing costs, from per-user pricing to full enterprise implementations.

SAS, Sales & Licensing FAQ

66% vs. 9%

Share of 2025 US data-scientist job postings mentioning Python versus SAS.

O*NET Online / Lightcast

SAS + Snowflake + Databricks

Customers increasingly operate across external data platforms alongside SAS Viya.

Techzine Global

Why Organizations Continue to Stay on SAS

Before examining the cost, it is important to understand the reason organizations continue to stay. In regulated life sciences, technology choices are not made on technical capability alone. A statistical environment must support reproducible analyses, controlled changes, validation, audit trails, established submission processes and teams that understand how to demonstrate that the output is reliable. A platform that has been used across many submission cycles accumulates not only code but also procedures, templates, review practices and organizational knowledge. Replacing it therefore means changing an operating model, not simply converting programs from one language to another.

This is particularly important in clinical development. A SAS program may sit inside a chain that begins with data standards and data collection, continues through derivations and statistical analyses, and ends with tables, listings and figures used in a regulatory submission. Even where a newer technology can perform the same calculation, the organization must still establish confidence in the new process, validate it appropriately, train users, document the change and ensure that historical knowledge is not lost. For this reason, SAS can remain the sensible choice for specific regulated workloads even when the surrounding enterprise architecture has modernized.

The problem begins when historical dependence becomes a default architecture. A company may have moved its enterprise data warehouse to Snowflake, built a lakehouse on Databricks, adopted Python for machine learning and established cloud-based data engineering practices, while continuing to treat SAS as the natural home for any workload that touches statistical data. Over time, the organization starts building connectors and duplicate pipelines around that decision. SAS is no longer just an application used for a specific purpose; it becomes a permanent boundary between teams, data platforms and skill sets. The cost of staying is therefore determined less by the existence of SAS itself than by how much of the wider operating model has been built around keeping SAS at the center.

The Costs of Staying on SAS

  • License and Infrastructure Cost

The first and most visible cost is the SAS contract. SAS uses different pricing models depending on the offering, including user-, capacity-, transaction- and revenue-based approaches. Its own licensing FAQ indicates that specific offerings can begin at around $1,500 per user, while fully implemented enterprise offerings can reach millions of dollars. Those figures provide useful context, but they do not answer the question a large pharmaceutical company actually needs to ask: what does the SAS-enabled operating model cost in total?

A large SAS estate can require infrastructure, administration, environment management, testing and ongoing maintenance in addition to the software itself. More importantly, the license can remain constant even when the workload mix changes. If SAS is still licensed broadly while some users increasingly work in Python, SQL, Databricks or Snowflake, the organization may be paying for capacity or access that no longer corresponds to the work being performed. A workload-level view is therefore more informative than an enterprise-wide license number. The relevant measure is the cost of SAS per active workload and the value that workload creates, rather than simply the amount paid to the vendor each year.

  • Integration and Data Movement Cost

The second cost appears when SAS is not the system where the organization’s primary data lives. Consider a modern pharmaceutical data landscape: clinical data may be managed in an enterprise cloud platform, genomic and real-world data may reside in cloud storage or a lakehouse, data engineering may be performed with Spark, and AI teams may build models in Python. If a regulated analysis still has to run inside SAS, the organization needs a reliable path for getting the right data into the SAS environment and getting the results back out.

That path creates engineering work. Pipelines have to transform data into SAS-compatible structures, interfaces must be monitored, permissions must be managed and changes in upstream schemas must be tested. When the same business logic is implemented once in the enterprise data platform and again in SAS, the organization also creates two places where a definition can change. The result is not necessarily a visible failure; it is often a slower change process and more reconciliation work. As the number of data sources grows, the number of dependencies grows with it. This is one reason the economic impact of a legacy SAS architecture can increase even when the underlying SAS program itself has barely changed.

  • Talent and Knowledge Cost

SAS expertise remains valuable in life sciences, but it is increasingly specialized. The 2025 job-posting data cited in this briefing illustrates the split: SAS appeared in 66% of US biostatistician postings, while Python appeared in 17%; in the broader data-scientist market, Python appeared in 66% of postings compared with 9% for SAS. The numbers do not mean SAS is becoming irrelevant. They show that the SAS talent pool is more concentrated around specialist statistical and regulated roles, while the broader analytics labor market is increasingly oriented around Python, SQL and cloud technologies.

For a pharmaceutical company with a large legacy estate, this creates a specific talent problem. The ideal employee may need to understand SAS programming, clinical statistics, regulatory expectations, validation, legacy macros and the organization’s historical business rules. That combination is considerably narrower than the market for a data scientist who works primarily with Python, SQL, cloud infrastructure and machine learning. Recruitment can therefore take longer, salary pressure can increase, and the organization can become dependent on a relatively small group of specialists.

The more serious risk is knowledge concentration. A senior SAS programmer may understand not only the code but also why an apparently redundant step exists, which historical submission requirement it supports and which downstream process will break if it is removed. That knowledge is difficult to document completely. When the individual leaves, the company may lose architectural understanding along with coding capacity. Modernization then becomes more expensive because the organization must first rediscover what its existing estate actually does.

  • Analytics Debt and the Cost of Duplication

Analytics debt is the cumulative cost created when an organization keeps an analytical architecture working without simplifying the dependencies underneath it. SAS is particularly susceptible to this pattern because stable programs can continue producing correct results for years. That stability makes the debt difficult to see. A clinical report still runs, a submission is still produced and a statistical program still returns the expected result, so there is little immediate incentive to change it. Meanwhile, the enterprise around that program continues to evolve.

A data engineering team may build a modern pipeline into Snowflake or Databricks, while another team creates a SAS-specific extract from the same source. A Python team may reproduce a transformation for an AI use case. Governance may be implemented differently in each environment. Monitoring, access controls, development processes and production deployments may also diverge. None of these decisions is necessarily unreasonable in isolation. The problem is the cumulative effect: the organization ends up maintaining several versions of the same data and business logic, along with the interfaces that reconcile them.

That debt compounds because every new data source or analytical use case adds another dependency. If a dataset must be transformed separately for SAS and a cloud platform, a change to the source can require multiple validations. If a definition changes, teams must determine which environment is authoritative. If a new analyst needs access, several permissions and workflows may need to be coordinated. Over time, the organization spends more effort maintaining connections between analytics systems than creating new analytical capability. This is why analytics debt should be measured quarterly rather than discovered only when a major migration or transformation exposes it.

Why SAS Supporting Databricks Matters

SAS’s expanded support for Databricks Spark in 2025 is significant because it reflects the architecture life sciences organizations are already moving toward. The problem it addresses is straightforward: companies increasingly want their large-scale data processing and storage to remain in a cloud data platform, but they may still have SAS models and statistical workflows that are important to the business. Moving every SAS workload to a new platform can be expensive and risky; forcing all enterprise data back into SAS creates duplication and movement.

The emerging model is to bring computation closer to the data. SAS can work with Databricks so that SAS models can be published to and executed in the Databricks environment rather than requiring every analytical step to happen inside a traditional SAS-centred architecture. Similar integration efforts around Snowflake and Microsoft Fabric point to the same broader direction. The practical value is not simply that ‘SAS now connects to Databricks’. It is that the boundary between the regulated SAS layer and the modern enterprise data platform can become thinner.

For a life sciences organization, that can reduce unnecessary data movement and make it easier to preserve existing SAS capabilities while modernizing the surrounding architecture. It does not automatically eliminate the underlying cost problem. If a company simply adds Databricks without retiring redundant SAS pipelines, it can increase complexity. The benefit appears when integration is used as part of deliberate workload segmentation: keep the workflows that genuinely depend on SAS, move suitable processing closer to enterprise data, and retire duplicate components instead of allowing both environments to expand indefinitely.

From Cost Diagnosis to Strategic Reframe

At this point, the obvious conclusion might be that SAS has become too expensive and should be removed. That would be an overly simplistic response. The costs described above are not evidence that SAS has no value; they are evidence that the way SAS is positioned in the architecture matters. A validated clinical programming workflow can be highly valuable even when the surrounding enterprise data platform has changed. The issue arises when the organization keeps SAS responsible for workloads that no longer require its unique capabilities simply because those workloads were historically built there.

This is the gap between application modernization and operating-model modernization. Replacing SAS programs one by one can create a large migration project without addressing the architecture that created the cost in the first place. The better approach is to separate the regulated core from the redundant layer around it. Clinical statistical programming, TLF generation and regulatory submission workflows may continue to require a controlled SAS environment. Enterprise data engineering, large-scale data processing, AI/ML development, real-world data preparation and many commercial analytics workloads may be better suited to cloud-native platforms.

The goal is therefore not to run two complete stacks forever. It is to reduce the SAS footprint to the workloads where it provides differentiated value. Done properly, coexistence can lower rather than increase cost: fewer unnecessary SAS seats, less custom integration, fewer duplicated transformations and a smaller amount of legacy code that must be understood and maintained. The strategic question becomes ‘where should SAS remain?’ rather than ‘how do we eliminate SAS?’.

PERCEPTIVE ANALYTICS PERSPECTIVE

Modernization should begin with a workload-level assessment rather than a wholesale migration. For each SAS workload, quantify license cost, specialist dependency, data movement, duplicated logic, maintenance effort and migration feasibility. The result is a map of where SAS creates unique regulatory or statistical value and where the organization is paying for complexity that can be removed.

What a Modernization Framework Changes

A modernization framework is necessary because the cost problem is distributed across technology, people and process. Without a framework, organizations tend to choose one of two extremes: preserve the existing SAS estate indefinitely or launch a large migration programme based on the assumption that every workload should move. Both approaches can destroy value. The first allows analytics debt to compound; the second can move validated workloads unnecessarily and create new operational and regulatory risk.

A workload-segmentation framework creates a common decision process. Every SAS workload can be evaluated against business value, regulatory criticality, maintenance burden, architectural duplication, talent dependency and migration feasibility. This changes the conversation from a technology preference into an evidence-based portfolio decision. It also makes modernization incremental: teams can remove low-value duplication while protecting the workflows that genuinely need the stability and controls of SAS.

What To Do Instead: A SAS Modernization Framework

01 . Protect the regulated core

Start by identifying clinical, statistical and submission workloads where SAS provides validated capabilities or where the cost and risk of changing the workflow are disproportionate to the potential benefit. The objective is not to preserve everything labelled ‘SAS’; it is to protect what is genuinely regulatory-critical. This gives leadership a safe boundary for modernization and prevents cost reduction from becoming a compliance exercise.

02 . Identify the redundant layer

Map where SAS duplicates capabilities already available in Snowflake, Databricks, Python, SQL or enterprise data engineering platforms. Look for duplicate datasets, repeated transformations, parallel models, manual exports and SAS programs that exist primarily because the organization never retired an older process. Removing this layer can produce benefits without touching the regulated core: fewer pipelines to maintain, fewer reconciliations and fewer points of failure.

03 . Move compute closer to data

Where practical, reduce unnecessary movement between the enterprise data platform and SAS. This is where integrations such as SAS support for Databricks become strategically useful. Keeping large datasets and heavy processing in the environment designed to handle them can reduce copying and improve scalability, while SAS remains available for the workloads that require it. The change is architectural rather than cosmetic: the data platform becomes the system of record and SAS becomes a specialized analytical component.

04 . Modernize selectively

Prioritize workloads using a combination of business value, maintenance burden and migration feasibility. High-maintenance, low-differentiation workloads are often the best starting point because the organization can remove recurring cost without disrupting a critical regulatory process. Conversely, a highly validated submission workflow may have a low modernization priority even if it runs on old technology. This sequencing helps the company realize benefits early and use the savings and learning to fund more complex changes.

05 . Build a dual-skilled workforce

The workforce strategy should not be to replace every SAS programmer with a Python developer. That approach risks losing regulatory knowledge that is difficult to recreate. Instead, organizations should build teams that combine SAS, Python, SQL, cloud and life sciences domain expertise. Cross-training reduces key-person dependency, gives teams a path to modernize workloads gradually and makes the organization less exposed to a single specialist talent pool.

06 . Measure analytics debt quarterly

Analytics debt should be treated as an operating metric rather than a one-time transformation topic. Leadership can track SAS license cost per active workload, number and usage of SAS programs, SAS-specific data pipelines, duplicate datasets, manual handoffs, specialist FTE dependency, average change effort and migration readiness. Reviewing these measures quarterly makes architectural drift visible. It also gives management a way to distinguish genuine modernization progress from simply adding another platform alongside the old one.

Figure 1: The SAS workload segmentation framework

Proof Points: Modernizing Without Abandoning Governance

The direction of travel is already visible in regulated organizations. Regeneron is upgrading from SAS 9 on SAS Grid to SAS Viya in the cloud, with work focused on the specific requirements of a regulated pharmaceutical environment. SAS describes how Regeneron and SAS co-developed ways to batch-run hundreds of programs and logs in Viya while meeting strict submission standards, and how the organization built a scalable Statistical Computing Environment around Viya. The example is useful because it demonstrates that modernization does not have to mean abandoning the controls that made the legacy environment valuable. Instead, the controlled capability can be redesigned on a cloud platform with the operating requirements of the pharmaceutical environment kept intact.

A second example is Chiesi Group. Chiesi moved from a PC-based architecture to SAS Viya deployed with SAS Managed Cloud Services on Microsoft Azure. According to SAS’s customer story, the organization needed a shared environment, standardized processes and strong documentation of the validation of the system installation and the lifecycle of the data and programs generating analysis. The move was therefore not simply an infrastructure refresh. It changed how teams worked, improved data exchange and standardization, and supported the efficient delivery of trial results to regulatory authorities. The lesson for organizations carrying legacy SAS environments is that modernization can be framed around stronger integration, standardization and cloud operating practices rather than a binary choice between keeping SAS unchanged and eliminating it.

These examples also reinforce an important distinction. Modernization can happen within the SAS ecosystem when the objective is to improve the operating model, but the same workload-segmentation principle still applies when the enterprise is deciding what should remain on SAS and what should move to other platforms. The destination is not the point; the architectural role of each workload is.

Regeneron, for example, is upgrading from SAS 9 on SAS Grid to SAS Viya in the cloud. The modernization is being designed around the requirements of a regulated pharmaceutical environment, including batch execution of large numbers of programs and a scalable Statistical Computing Environment.

Figure 2: Regeneron SAS Viya modernization

Conclusion

SAS is not going away from life sciences, and the organizations that treat it as if it must disappear immediately may create a new set of risks while trying to solve an old cost problem. SAS continues to provide value in statistical programming, clinical analytics and regulated workflows where validation, auditability and accumulated domain knowledge matter. The question is not whether that value exists. The question is whether the organization is paying for far more SAS capability, supporting infrastructure and specialist dependency than the regulated core actually requires.

The hidden costs of staying on SAS are therefore architectural as much as contractual. They include duplicated data and logic, integration work, specialist talent, manual reconciliation, legacy maintenance and the opportunity cost of making every new analytical capability work around an increasingly complex boundary. These costs accumulate gradually, which is why they are easy to overlook in an annual licensing discussion. Analytics debt is not a single invoice; it is the recurring effort required to keep yesterday’s architecture connected to tomorrow’s data and analytics environment.

The practical response is selective modernization. Protect the regulated core. Identify and retire redundant SAS functionality. Move data and compute closer together. Build teams that understand both SAS and modern analytics technologies. Measure the debt quarterly and use those measures to guide investment. In other words, keep SAS where its regulatory and statistical value is distinctive, while removing the dependencies that make the rest of the enterprise harder to hire for, integrate, scale and evolve.


Perceptive Analytics works together with life sciences companies to transform their analytics estate using scalable data engineering, analytics, AI-driven automation and cloud data platforms in order to enable organizations to decrease their reliance on platforms while ensuring regulatory processes are maintained.

“The cost of staying on SAS is no longer just the license. It is the cost of maintaining an architecture, talent model and operating model around a technology that increasingly has to coexist with the modern data stack.”

References & Sources


Submit a Comment

Your email address will not be published. Required fields are marked *