By Adhya Khare, Marketing Coordinator

This article is for HPC, research-computing, and data-infrastructure leaders evaluating long-term or archival storage. It covers the DKRZ deployment’s scale, migration approach, storage economics, and buying lessons.

The German Climate Computing Center (DKRZ, Deutsches Klimarechenzentrum) has brought a new hierarchical storage management system online for one of the world’s largest climate simulation archives. The new system is built on Versity ScoutAM software, delivered in partnership with NEC, and the scale of the project is what makes it notable: one exabyte of climate simulation data across a seven year contract term. 

The quantity of data is significant. But the larger story is not simply that DKRZ built a large tape archive. The price per GB of high-performance scratch storage has risen fivefold or more under AI demand, and it is crushing HPC budgets. That pressure is driving a wave of demand for storage offload: fast, efficient archival storage that can move inactive data out of the most expensive tier.

At a time when AI infrastructure demand is raising the price and constraining the availability of enterprise flash and disk, the ability to offload to more efficient storage has become strategically important. A high-performance archive is no longer just where an organization preserves old data. It is what allows the performance tier to stay focused on active computation.

Offloading Lustre Scratch to Tape: Shrinking the Most Expensive Storage Tier

DKRZ’s Levante supercomputer utilizes a 130 PB global Lustre file system for project data, model output, and temporary processing. This is the high-performance working tier where researchers generate and analyze large datasets.

It is high-performance storage that DKRZ cannot afford to use as indefinite storage for completed simulation output. Every terabyte that remains on Lustre after active processing is finished consumes space that could otherwise support running models, new projects, and live scientific workloads. The archive provides an efficient exit path: finished output can be quickly moved out of the performance tier while remaining available when researchers need it again.

ScoutAM transfers data between DKRZ’s HPC systems and the archive at aggregate speeds of up to 30 GB/s (over 2 PB per day). A ScoutAM managed 8.2 PB combined SSD and HDD cache absorbs peak ingest loads and keeps a portion of archived data immediately accessible without mounting tape.

That changes the role of the archive in the storage hierarchy.

Instead of treating tape as a distant destination that users avoid because retrieval may interrupt their work, DKRZ can use the archive as an active capacity tier. Data moves off Lustre quickly enough to free premium storage, while the cache allows frequently recalled datasets to return without waiting for physical media. The objective is not simply to place more data on tape. It is to keep Levante’s performance storage focused on its intended job: active computation and processing.

This matters well beyond DKRZ, and the timing is not incidental. AI infrastructure demand has driven an unprecedented storage supply crunch: enterprise SSD contract prices have risen sharply through 2026, and both major hard drive makers have sold their entire 2026 production to hyperscale cloud customers, with long-term agreements locking in supply through 2028. In that market, expanding the performance tier every time data grows is an increasingly expensive strategy. A high-throughput archive offers another option: increase the data you can retain without expanding your most expensive.

Zero Data Migration for a 254 PB Archive

The most striking part of the DKRZ deployment is what DKRZ did not have to do: migrate the data. DKRZ brought its entire 254 PB archive onto the new system with no data migration. No data was rewritten. The legacy archive stayed in place and stayed readable while new data began landing in an open format on current-generation media, and the cutover was measured in hours, not the years a full migration cycle would have demanded.

Inside DKRZ: Germany’s National Climate Archive

Founded in 1987 with founding scientific director Nobel laureate Klaus Hasselmann and headquartered in Hamburg, DKRZ is Germany’s national supercomputing center for climate research, serving more than 1,000 climate and earth-system researchers. Every major climate model, simulation, and dataset produced by the country’s leading research institutions passes through its systems.

That responsibility has produced one of the largest scientific data archives in the world. The environment spans ten tape libraries, including three Spectra Logic TFinity libraries, with 130 LTO drives and capacity for as many as 100,000 tape cartridges. Inside it sits 254 PB of simulation data generated across decades of research and multiple generations of computing infrastructure.

That data is not static. Scientists revisit previous model output to reproduce findings, compare historical simulations, train new analytical systems, and support increasingly high-resolution research. The archive, therefore, has to do more than preserve data. It has to move data quickly enough to function as an active extension of the HPC environment.

The Archive That Has to Outlive Everything

Most storage procurements are written to answer one question: how do we get in? How fast can we deploy, how much will it hold, how quickly can we start writing data. The team at the German Climate Computing Center (DKRZ) started from the opposite end. Before they chose a platform for an exabyte-scale archive, they asked how they would one day get out of it — and they wrote the answer into the contract.

It sounds pessimistic. It is actually the most optimistic thing a buyer of long-term storage can do. An archive meant to outlive four decades of climate science will outlive vendors, media generations, file formats, and the careers of the people who build it. Planning the exit on day one is how you make sure the data survives all of them.

Vendor lock-in is an abstract phrase until you’re the one holding a quarter of an exabyte you can’t easily move. DKRZ had spent more than a decade first on IBM HPSS and then on another vendor, and the systems team had learned exactly where the traps are. They wanted a decisive break from proprietary tape formats and appliance-style deployments. Not because those systems had failed to store the data, but because they made the data dependent on the tooling that wrote it.

Here is the distinction that separates a durable archive from a fragile one. In a proprietary system, the bytes on tape are only meaningful to the software that put them there even if those bytes are written in an open format like LTFS. Lose access to that software — through an end-of-life notice, a license dispute, an acquisition, a price increase you can’t absorb — and you don’t just lose a vendor relationship. You lose the ability to read your own data. For a national climate record that is meant to serve the decades into the future, that is an unacceptable single point of failure.

Your Data On Your Terms: DKRZ’s Tender Requirements

DKRZ ran a formal public tender with a list of technical requirements shaped by hard-won operational experience. Versity ScoutAM, delivered in partnership with NEC, cleared every mandatory requirement:

  • 100% local admin autonomy. No proprietary abstraction between the team and the system. The people who run the archive answer to no external gatekeeper to operate it.
  • Direct tape access in an open format. The bytes on tape are readable without the vendor’s software in the loop. This is the technical foundation the exit strategy rests on.
  • A flexible policy engine capable of modeling the full range of workflows that Germany’s climate community depends on 
  • Site licensing with no per-server or per-capacity surprises. Cost that scales predictably as the archive grows toward an exabyte, with no metering that turns growth into a renegotiation.

Notice what these have in common. None of them is a feature in the usual sense. They are structural guarantees that DKRZ will still control its own archive in fifteen years, regardless of what happens to any vendor, including Versity. The most important thing a storage vendor can offer an archive of this permanence is a credible promise that the customer doesn’t need the vendor to read its own data. Versity’s answer is architectural: the metadata required to reconstruct the archive lives on the tapes themselves, in open format, and can be dumped and read without specialized tools. On top of that, Versity provides a free, read-only license in perpetuity, so DKRZ can always access its own archive even if it never buys another Versity license again. DKRZ can recover its data or read its own tapes with entirely different software, with no dependency on proprietary tooling.

Why the Migration Took Hours, Not Years

The clearest test of an open architecture is what happens when you try to move. Closed platforms are built to make that hard, often locking data behind proprietary formats that force a full migration before you can leave. Openness inverts that.

Rewriting 254 PB across five media generations would have been a multi-year, high-risk undertaking. Instead of moving the data, ScoutAM ingested the existing metadata and treats the legacy archive as a read-only tier. New writes land in the open Versity format on current-generation media, while every legacy file remains fully accessible in place, transparent to users. Nobody has to wait years for a tape rewrite, and nothing goes dark during the transition.

An archive built on open formats is cheap to leave, which is exactly why moving into one is cheap too. The same property that lets you walk away later is what lets you arrive without a painful, years-long migration now. Lock-in cuts both ways. A platform that is hard to leave is usually hard to join. Openness is the property that makes both directions painless.

Self-Service Access: Making the Archive Usable

An exit strategy protects the data. But DKRZ also needed the archive to become something researchers work with rather than around, which is why they chose Versity.

Historically, locating archived data at DKRZ meant involving an administrator or writing custom metadata queries against the catalog. That model doesn’t scale into the AI era, and it isn’t how researchers want to spend their time. ScoutAM includes metadata capabilities so scientists can search, locate, and retrieve their own data directly, with automatic metadata extraction and custom tagging built in.

Crucially, this self-service layer doesn’t reintroduce the dependency the team worked so hard to design out. It is native to ScoutAM, and the metadata it works from lives in open format on the archive itself. The convenience layer and the durability guarantee are the same architecture, not competing ones. You don’t trade away portability to get usability.

What Buyers Can Learn From the DKRZ Deployment

DKRZ’s archive now rests on a foundation built to serve the next four decades of climate science the way the last one served the previous four. You do not need an exabyte of climate data to use what the deployment demonstrates. It comes down to three questions worth putting to any long-term storage decision.

Does it earn its keep? In a market where high-performance storage is both expensive and scarce, an archive fast enough to pull finished data off premium disk is the lever that keeps your costliest tier small. Prioritize throughput and cache, not just capacity.

Can you leave? Decide how you will exit before you decide how you will arrive. Insist on open, self-describing formats and direct access to your own media, and treat any structural dependency on a single vendor’s software as the risk it is.

Can people use it? An archive that needs an administrator and a custom query for every retrieval is a vault, not a resource. Self-service search built on open metadata is what turns decades of stored data back into active science.

Performance, openness, and usability are not three separate products. They are one architecture doing three jobs at once: keeping data affordable to hold, readable for decades, and reachable by the people who need it. To evaluate your own environment against these criteria, talk to the Versity team about a storage assessment.

Frequently Asked Questions

How did DKRZ modernize 254 PB without data migration?

Instead of rewriting the data onto new media, ScoutAM imported the existing archive’s metadata and treats the legacy tapes as a read-only tier. New data is written in the open Versity format, while legacy files stay accessible in place. Test imports completed in a few hours rather than the multiple years a full rewrite would take.

How does the archive reduce HPC storage costs?

The archive offloads finished data from DKRZ’s 130 PB Lustre performance tier at up to 30 GB/s, backed by an 8.2 PB cache. Moving completed output off expensive flash and disk lets DKRZ grow total capacity without expanding its most costly storage tier, which matters as enterprise flash and disk prices climb under AI-driven demand.

Can DKRZ read its data without Versity software?

Yes. The metadata needed to reconstruct the archive lives on the tapes in an open format that can be read without specialized tools, and Versity provides a free read-only license in perpetuity. DKRZ can access its own archive even if it never buys another Versity license.

What was DKRZ using before ScoutAM?

DKRZ previously ran proprietary hierarchical storage management platforms, including Stronglink and IBM HPSS, for more than a decade. The team moved to an open-format system specifically to avoid the risk of data being locked behind vendor software.

Read more here

Rise to the challenge

Connect with Versity today to find out how we can tailor a solution to keep your organization’s data safe and accessible as you advance your mission.