Getting Started with Berkelium
Welcome to Berkelium, the multi-tenant scientific Kubernetes cluster managed by the Science IT department at Lawrence Berkeley National Laboratory (LBNL).
Berkelium provides LBNL researchers, experimental beamlines (ALS), biosciences groups, climate modelers, and particle physicists with modern cloud-native computing, GPU acceleration (NVIDIA H100, A100, L40S), massive parallel storage, and self-service orchestration.
Documentation Quick Index
Section titled βDocumentation Quick IndexβExplore our comprehensive guides modelled for scientific computing:
- Getting Access & SSO β LBNL OneID authentication and requesting your first namespace.
- Using Berkelium β CLI vs. Web Portal, tool installations (
kubectl,kubelogin). - Resource Hierarchy β How divisions, grants, and namespace quotas are structured.
- Cluster Policies β Fair-share scheduling, idle container rules, and security policies.
- Deployed Services β Prometheus, Grafana, Harbor Registry, Vault, and KubeRay.
- FAQ & Support β Common questions, office hours, and Slack channels.
- Kubernetes for HPC β Bridge the gap between Slurm batch scripts and container pods.
- Docker & Apptainer β Building reproducible research containers.
- Basic Kubernetes β Launching, inspecting, and deleting pods and services.
- Batch Jobs β Running parameter sweeps and parallel array jobs.
- Persistent Volumes β Mounting persistent storage into your containers.
- Debugging Pods β Reading logs, diagnosing
CrashLoopBackOff, and interactive shells.
- GPU Accelerated Pods β Requesting NVIDIA H100, A100, and MIG slices.
- CPU-Only & High-Memory β 128-core AMD EPYC Bergamo nodes with 1.5TB RAM.
- Coder Development Environments β Self-hosted cloud IDEs with VS Code, JetBrains, and AI agents.
- Interactive Pods β Guidelines for Jupyter notebooks, VS Code, and GUI desktops.
- Distributed Ray Clusters β Multi-node Ray orchestration with the KubeRay operator.
- Distributed PyTorch β Distributed Data Parallel (DDP) and Kubeflow Training.
- JupyterHub Service β Launching shared JupyterLab instances for research labs.
- CephFS Shared Filesystem β Multi-node shared
ReadWriteManyvolumes. - Local NVMe Scratch β Ultra-low-latency scratch storage for high-IOPS model training.
- Data Movement & Transfer β Globus endpoints, fast rsync, and beamline POSIX bridges.
- Purge Policies β Inactive data retention and snapshot guidelines.
- Exposing Web Services β Ingress controllers, routing, and automatic SSL/TLS certificates.
- LoadBalancer VIPs & ESnet β Direct 100Gbps connectivity through ESnet data fabrics.
Need Help or Office Hours?
Section titled βNeed Help or Office Hours?β- Slack: Join
#berkelium-userson the LBNL Slack workspace. - Office Hours: Every Tuesday & Thursday (11am β 12pm PT) via Zoom.
- Help Desk Ticket: Submit an inquiry at Science IT Consultation.
