Berkelium Cluster Specifications

128+
Compute Nodes
96x
Enterprise GPUs (A100/H100/L40S)
4.5 PB
Ceph & NVMe Storage
100 Gbps
ESnet / Interconnect Fabric

Compute & Acceleration Node Pools

GPU Accelerated

NVIDIA Hopper & Ampere Pool

Dedicated and MIG-sliced GPU nodes optimized for PyTorch, TensorFlow, Jax, and Cryo-EM structural simulations.

  • AcceleratorsNVIDIA H100 (80GB SXM5), A100 (80GB), L40S
  • CPUs / Node64-core AMD EPYC 9554
  • Node Memory768 GB DDR5 ECC
  • Local Scratch2x 3.84TB NVMe (RAID0)
High-Memory CPU

Scale-Up Compute Pool

Large memory footprint nodes for genome assembly, graph analytics, in-memory databases, and large batch pipelines.

  • CPUs / Node128-core Dual AMD EPYC Bergamo
  • Node Memory1.5 TB to 2.0 TB DDR5
  • Workload TypeMemory-intensive / MPI
  • PreemptionGuaranteed / QoS Burstable
Interactive & Web

Standard Service Pool

General application hosting, JupyterHub instances, web APIs, scientific portals, and pipeline orchestrators.

  • CPUs / Node64-core Intel Xeon Platinum
  • Node Memory512 GB DDR4/DDR5
  • IngressAutomatic TLS via Let's Encrypt
  • Availability99.9% Multi-Zone HA

Storage Architecture

Berkelium provides low-latency, scalable storage classes configured specifically for massive scientific data formats (HDF5, NetCDF, Parquet, TIFF):

cephfs-replicated (ReadWriteMany)

Triple-replicated Ceph filesystem. Allows simultaneous read/write access across dozens of pods across all nodes.

nvme-scratch-local (ReadWriteOnce)

Ultra-fast direct PCIe NVMe storage for temporary intermediate model training data and high-IOPS shuffles.

LBNL Science Data Bridge

Read-only NFS/POSIX mounts connecting directly to central laboratory instrument repositories and NERSC archives.

Security & Multi-Tenancy

Engineered from the ground up for zero-trust scientific multi-tenancy compliant with DOE and UC Berkeley standards:

LBNL OneID / OIDC Authentication

Direct integration with Berkeley Lab identity systems. No manual key exchange; authorization based on division and group membership.

Strict Namespace Isolation & Quotas

Granular ResourceQuotas and LimitRanges prevent noisy-neighbor issues and cap GPU consumption according to grant allocations.

Automated Secret Management

HashiCorp Vault Agent sidecars automatically inject sensitive credentials into container memory without storing secrets on disk.

Need compute resources for your research group?

Provision a dedicated Kubernetes namespace with custom CPU, memory, and GPU quotas through our self-service portal.