HPC Climate Modeling
DKRZ, Germany’s Climate Archive, Rebuilt for the Exabyte Era

Exabyte-Scale Climate Archive
Unified archive designed to grow from 254 PB toward a 500-million-file, 1 EB
254 PB Zero-Data Migration
Legacy tape archive brought online in hours — no rewrite, no downtime
Self-Service Archive Access
Scientists can search directly, with automatic metadata extraction and custom tagging
Overview
The German Climate Computing Center (DKRZ, Deutsches Klimarechenzentrum) is Germany’s national supercomputing center for climate research. Founded in 1987 by Nobel laureate Klaus Hasselmann and headquartered in Hamburg, DKRZ provides high-performance computing and data services to more than 5,000 climate and earth-system researchers.
Every major climate model, simulation, and dataset produced by Germany’s leading research institutions passes through DKRZ’s systems. Its tape archive holds 254 PB of climate research data, grows by roughly 30 PB every year, and spans two sites, two library vendors, and five generations of LTO media.
Climate science is increasingly a data-reuse discipline. Researchers now revisit decades of simulation output to reproduce results, feed AI-driven weather and climate workflows, and run higher-resolution models that produce data at greater volume, velocity, and file count than any previous system was designed to handle. To support this shift, DKRZ needed an archive that could operate not as a passive preservation tier, but as an active scientific resource. One that is searchable, policy-driven, open, and ready for the next generation of climate research.

Challenges
Legacy archive system unable to keep up with concurrency and small-file workloads in a large scale data environment
Fragmented, multi-vendor archive across two sites, two library vendors, and five LTO generations
Proprietary tape formats and appliance-style deployments with no documented exit strategy
Growing volumes of small files from modern climate formats like Zarr — notoriously inefficient on tape
User-driven data movement forcing researchers to reason about storage tiers instead of science
Unpredictable restore times as peak days pushed 200 TB in and 100 TB out of the archive
Solution
DKRZ ran a formal public tender with a demanding list of technical requirements shaped by hard-won operational experience. Versity ScoutAM cleared every mandatory requirement and gave DKRZ what previous systems couldn’t: 100% local administrative control, direct tape access in an open format, and a flexible policy engine engineered from day one for exabyte-scale.
Rewriting 254 PB across five media generations would have been a multi-year, high-risk undertaking. Instead, ScoutAM ingested the existing metadata and treats the legacy archive as a read-only tier. New writes land in the open Versity format on current-generation media, while legacy data remains fully accessible in place, transparent to users.
ScoutAM’s orchestration engine also solves one of the fastest-growing headaches in climate research: small files. Modern formats like Zarr favor huge numbers of small files, a pattern notoriously inefficient on tape. Containerization was a hard requirement in DKRZ’s procurement, and ScoutAM addresses it directly by automatically packing small files into larger objects on tape.
Furthermore, DKRZ’s users now have a direct, self-service layer for accessing the archive. On legacy systems, retrieval of big data sets often meant asking an administrator to run a query and stage files first. Here, scientists can search, locate, and retrieve their own data with automatic metadata extraction and custom tagging built in. The archive shifts from a passive long-term store to an active, queryable knowledge base that plugs directly into modern HPC and AI workflows.

“DKRZ’s mission is to be a partner to climate science, and our vision is to reliably unlock the potential of accelerating technological progress for climate research. Delivering on that vision requires a data infrastructure we can trust for decades — which is why we chose Versity.”
Prof. Dr. Thomas Ludwig
Managing Director, DKRZ
Results
✓ Exabyte-scale archive designed for a 500-million-file namespace
✓ Zero-data migration brought a 254 PB tape archive online in hours
✓ Self-service search, automatic metadata extraction, and custom tagging natively within ScoutAM
✓ Non-proprietary, open-format architecture with 100% local admin autonomy
✓ Automated dual-site replication between Hamburg and Garching for defined data sets
More Case Studies
TACC Builds 1 EB Archive for AI Research with Versity ScoutAM
Discover how TACC built an exabyte-scale, AI-ready archive powered by Versity ScoutAM to keep pace with Horizon’s massive data demands.
Accelerating Discovery: TGen Research Supercharged with Versity Archival Solution
Discover how Versity supports TGen – Part of City of Hope in managing vast amounts of critical research data efficiently…