Skip to content
42 changes: 21 additions & 21 deletions docs/clusters/alpine/alpine-hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,15 +16,15 @@ All Alpine nodes are available to all users. For full details about node access,

| Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric |
| --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- |
| {{ alpine_ucb_total_64_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_128_core_256GB_cpu_nodes }} Milan CPU | amilan | x86_64 AMD Milan | 2 | 128 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_64_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_128_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 128 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_48_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 2 | 48 | 1 | 21.5 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_64_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 1 | 64 | 1 | 16 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_mi100_gpu_nodes }} Milan AMD GPU | ami100 | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | AMD MI100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_a100_gpu_nodes }} Milan NVIDIA GPU | aa100 | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_gh200_gpu_nodes }} Grace CPU NVIDIA Hopper GPU | gh200<br><br>Note: these nodes are only available upon request, please submit a [support request form](https://colorado.service-now.com/req_portal?id=ucb_sc_rc_form). | ARM Neoverse V2 | 1 | 72 | 1 | 6.6 | NVIDIA Hopper GPU | 1 | 1.8 T SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_acompile_nodes }} Milan CPU compile nodes | acompile | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_64_core_256GB_cpu_nodes_atesting }} Milan CPU test nodes; pulls from CU amilan pool | atesting | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_acompile_nodes }} AMD CPU compile nodes | acompile | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_64_core_256GB_cpu_nodes_atesting }} AMD CPU test nodes; pulls from CU's `acpu` pool | atesting | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_ucb_total_a100_test_gpu_nodes }} Milan NVIDIA GPU testing node | aa100 (requested using the gpu-testing QoS) | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 (each split by MIG) | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_ucb_total_mi100_test_gpu_nodes }} Milan AMD GPU testing nodes; pulls from ami100 pool | ami100 (requested using the gpu-testing QoS) | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | AMD MI100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE |

Expand All @@ -39,7 +39,7 @@ All Alpine nodes are available to all users. For full details about node access,

| Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric |
| --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- |
| {{ alpine_amc_total_64_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_amc_total_64_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_amc_total_64_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 1 | 64 | 1 | 16 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_amc_total_128_core_2TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 2 | 128 | 1 | 16 | N/A | 0 | 70G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_amc_total_a100_gpu_nodes }} Milan NVIDIA GPU | aa100 | x86_64 AMD Milan | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE |
Expand All @@ -48,7 +48,7 @@ All Alpine nodes are available to all users. For full details about node access,
:::

```{note}
**CU Anschutz job submission limit:** CU Anschutz users on Alpine are subject to a campus-wide hard limit of **200 concurrent jobs** across all QoS types (including `normal`, which system-wide allows up to 1000 jobs/user). This limit applies to all job types and partitions. It was implemented to reduce queue wait times for all Anschutz users given limited core-hour availability. If you need to run large numbers of jobs, consider using [GNU Parallel](https://github.com/kf-cuanschutz/CU-Anschutz-HPC-documentation/blob/main/Office-hours-presentation-files/GNU_parallel_presentation.pdf) as a workaround.
**CU Anschutz job submission limit:** CU Anschutz users on Alpine are subject to a campus-wide hard limit of **200 concurrent jobs** across all QoS types (including `cpu-normal`, which system-wide allows up to 1000 jobs/user). This limit applies to all job types and partitions. It was implemented to reduce queue wait times for all Anschutz users given limited core-hour availability. If you need to run large numbers of jobs, consider using [GNU Parallel](https://github.com/kf-cuanschutz/CU-Anschutz-HPC-documentation/blob/main/Office-hours-presentation-files/GNU_parallel_presentation.pdf) as a workaround.
```

### Colorado State University contribution
Expand All @@ -60,8 +60,8 @@ All Alpine nodes are available to all users. For full details about node access,

| Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric |
| --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- |
| {{ alpine_csu_total_48_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 2 | 48 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_csu_total_32_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 2 | 32 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
| {{ alpine_csu_total_48_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 48 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) |
| {{ alpine_csu_total_32_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 32 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE |
:::


Expand All @@ -86,13 +86,13 @@ Resources are requested within jobs by passing in SLURM directives, or resource

| Partition | Description | # of nodes | cores/node | RAM/core (GB) | Billing_weight/core |
| --------- | ---------------------------- | ---------- | ---------- | ------------- | ------------------- |
| amilan | AMD Milan (default) | {{ alpine_total_amilan_nodes }} | 32 or 48 or 64 or 128 | {{ alpine_standard_ram_per_core }} | 1 |
| acpu | AMD CPU nodes (default) | {{ alpine_total_acpu_nodes }} | 32 or 48 or 64 or 128 | {{ alpine_standard_ram_per_core }} | 1 |
| ami100 | GPU-enabled (3x AMD MI100) | {{ alpine_total_ami100_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.1<sup>3</sup> |
| aa100 | GPU-enabled (3x NVIDIA A100)<sup>4</sup>. For select nodes, MIG has been enabled providing 6x 20 GB NVIDIA A100 MIG instances. | {{ alpine_total_aa100_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.1<sup>3</sup> |
| al40 | GPU-enabled (3x NVIDIA L40)<sup>4</sup> | {{ alpine_total_al40_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.1<sup>3</sup> |
| amem<sup>1</sup> | High-memory | {{ alpine_total_amem_nodes }} | 48 or 64 or 128 | 16<sup>2</sup> | 4.0 |
| acompile | AMD Milan compile nodes | {{ alpine_total_acompile_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | N/A |
| atesting | AMD Milan test nodes | {{ alpine_total_atesting_cpu_nodes }}; Pulls from CU amilan pool | 64 | {{ alpine_standard_ram_per_core }} | 0.025 |
| acompile | AMD CPU compile nodes | {{ alpine_total_acompile_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | N/A |
| atesting | AMD CPU test nodes | {{ alpine_total_atesting_cpu_nodes }}; Pulls from CU's `acpu` pool | 64 | {{ alpine_standard_ram_per_core }} | 0.025 |
| gh200 | NVIDIA Grace-Hopper (GH200) nodes<br><br>Note: this partition is only available upon request, please submit a [support request form](https://colorado.service-now.com/req_portal?id=ucb_sc_rc_form). | {{ alpine_ucb_total_gh200_gpu_nodes }} | 72 | 6.65 | Billed at roughly twice the rate of our A100s |

```{important}
Expand All @@ -110,7 +110,7 @@ Resources are requested within jobs by passing in SLURM directives, or resource

All users, regardless of institution, should specify partitions as follows:
```bash
--partition=amilan
--partition=acpu
--partition=aa100
--partition=ami100
--partition=al40
Expand All @@ -119,14 +119,14 @@ All users, regardless of institution, should specify partitions as follows:

### Quality of Service (qos)

**Quality of Service or QoS is used to constrain or modify the characteristics that a job can have.** For example, by selecting the `long` QoS, a user can place the job in a **lower priority queue** with a max wall time increased from 24 hours to 7 days.
**Quality of Service or QoS is used to constrain or modify the characteristics that a job can have.** For example, by selecting the `cpu-long` QoS, a user can place the job in a **lower priority queue** with a max wall time increased from 24 hours to 7 days.

#### Available QoS for Alpine:

| QOS name | Description | Max walltime | Max jobs/user | Max hardware/user | Valid Partitions |
| ----------- | -------------------------- | --------------- | ------------- | ------------------ | ---------------- |
| normal | Standard QoS for non-testing partitions | 1 day | 1000 | 128 nodes | amilan |
| long | Longer wall times | 7 days | 200 | 20 nodes | amilan |
| cpu-normal | Standard QoS for non-testing partitions | 1 day | 1000 | 128 nodes | acpu |
| cpu-long | Longer wall times | 7 days | 200 | 20 nodes | acpu |
| mem-normal | Standard QoS for High-memory jobs | 24 hours | 1000 | 256 CPU cores | amem |
| mem-long | QoS for longer running High-memory jobs | 7 days | 200 | 185 CPU cores | amem |
| gpu-normal | Standard QoS for GPU jobs | 24 hours | 1000 | see [Available GRES on Alpine](#available-gres-on-alpine) | aa100,ami100,al40 |
Expand All @@ -142,20 +142,20 @@ All users, regardless of institution, should specify partitions as follows:
`````{tab-set}
:sync-group: tabset-ex-qos-req

````{tab-item} Requesting the normal partition
:sync: ex-qos-req-normal-partition
````{tab-item} Requesting the cpu-normal QoS
:sync: ex-qos-req-cpu-normal-partition

```bash
--qos=normal
--qos=cpu-normal
```

````

```` {tab-item} Requesting the long partition
:sync: ex-qos-req-long-partition
```` {tab-item} Requesting the cpu-long QoS
:sync: ex-qos-req-cpu-long-partition

```bash
--qos=long
--qos=cpu-long
```

````
Expand Down
2 changes: 1 addition & 1 deletion docs/clusters/alpine/quick-start.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ section.
## Cluster Summary
### Nodes
The Alpine cluster is made up of different types of nodes. A general overview of these nodes is as follows:
- **CPU nodes**: {{ alpine_total_256GB_cpu_nodes }} AMD Milan compute nodes with 256 GB RAM
- **CPU nodes**: {{ alpine_total_256GB_cpu_nodes }} AMD compute nodes with 256 GB RAM
- **GPU nodes**: a mixture of {{ alpine_total_gpu_nodes }} NVIDIA and AMD GPUs
- **High-memory nodes**: {{ alpine_total_hi_mem_cpu_nodes }} high-memory nodes with 1 TB of memory or more

Expand Down
14 changes: 7 additions & 7 deletions docs/clusters/alpine/slurm_directive_ex.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,8 @@ Below are some examples of SLURM directives that can be used in your batch scrip
To run a 32-core job for 24 hours on a single Alpine CPU node:

```bash
#SBATCH --partition=amilan
#SBATCH --qos=normal
#SBATCH --partition=acpu
#SBATCH --qos=cpu-normal
#SBATCH --nodes=1
#SBATCH --ntasks=32
#SBATCH --time=24:00:00
Expand All @@ -28,11 +28,11 @@ To run a 32-core job for 24 hours on a single Alpine CPU node:
To run a 56-core job (28 cores/node) across two Alpine CPU nodes in the low-priority qos for seven days:

```bash
#SBATCH --partition=amilan
#SBATCH --partition=acpu
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=28
#SBATCH --time=7-00:00:00
#SBATCH --qos=long
#SBATCH --qos=cpu-long
```

````
Expand Down Expand Up @@ -72,16 +72,16 @@ To run a 42-core job for 2 hours on a single Alpine NVIDIA GPU node, using 2 40

## Full Example Job Script

Run a 1-hour job on 4 cores on an Alpine CPU node with the normal qos that runs a python script using a custom conda environment.
Run a 1-hour job on 4 cores on an Alpine CPU node with the `cpu-normal` QoS that runs a python script using a custom conda environment.

```
#!/bin/bash

#SBATCH --partition=amilan
#SBATCH --partition=acpu
#SBATCH --job-name=example-job
#SBATCH --output=example-job.%j.out
#SBATCH --time=01:00:00
#SBATCH --qos=normal
#SBATCH --qos=cpu-normal
#SBATCH --nodes=1
#SBATCH --ntasks=4
#SBATCH --mail-type=ALL
Expand Down
4 changes: 2 additions & 2 deletions docs/compute/modules.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,8 +102,8 @@ loads Anaconda into the environment is provided below:
#SBATCH --nodes=1
#SBATCH --time=00:01:00
#SBATCH --ntasks=1
#SBATCH --partition=amilan
#SBATCH --qos=normal
#SBATCH --partition=acpu
#SBATCH --qos=cpu-normal
#SBATCH --job-name=test-job
#SBATCH --output=test-job.%j.out

Expand Down
14 changes: 7 additions & 7 deletions docs/compute/monitoring-resources.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,15 +102,15 @@ This will display the output:
job stats for user ralphie over past 35 days
jobid jobname partition qos account cpus state start-date-time elapsed wait
-------------------------------------------------------------------------------------------------------------------
8483382 sys/dash amilan normal ucb-gener+ 1 TIMEOUT 2021-09-14T09:32:09 01:00:16 0 hrs
8487254 test.sh amilan normal ucb-gener+ 1 COMPLETE 2021-09-14T13:21:12 00:00:02 0 hrs
8483382 sys/dash acpu cpu-normal ucb-gener+ 1 TIMEOUT 2021-09-14T09:32:09 01:00:16 0 hrs
8487254 test.sh acpu cpu-normal ucb-gener+ 1 COMPLETE 2021-09-14T13:21:12 00:00:02 0 hrs
8487256 interact ahub interacti+ ucb-gener+ 1 TIMEOUT 2021-09-14T13:22:11 12:00:22 0 hrs
8508557 acompile acompile compile ucb-gener+ 2 COMPLETE 2021-09-16T10:41:45 00:00:00 0 hrs
8508561 test.sh amilan normal ucb-gener+ 24 CANCELLE 2021-09-22T10:07:03 00:00:00 143 hrs
8508569 test amilan normal ucb-gener+ 4096 FAILED 2021-09-16T10:42:46 00:00:00 0 hrs
8508575 test amilan normal ucb-gener+ 8192 FAILED 2021-09-16T10:43:17 00:00:00 0 hrs
8508593 test amilan normal ucb-gener+ 4096 CANCELLE 2021-09-16T10:44:47 00:00:00 0 hrs
8508604 test amilan normal ucb-gener+ 2048 CANCELLE 2021-09-16T10:45:40 00:00:00 0 hrs
8508561 test.sh acpu cpu-normal ucb-gener+ 24 CANCELLE 2021-09-22T10:07:03 00:00:00 143 hrs
8508569 test acpu cpu-normal ucb-gener+ 4096 FAILED 2021-09-16T10:42:46 00:00:00 0 hrs
8508575 test acpu cpu-normal ucb-gener+ 8192 FAILED 2021-09-16T10:43:17 00:00:00 0 hrs
8508593 test acpu cpu-normal ucb-gener+ 4096 CANCELLE 2021-09-16T10:44:47 00:00:00 0 hrs
8508604 test acpu cpu-normal ucb-gener+ 2048 CANCELLE 2021-09-16T10:45:40 00:00:00 0 hrs
8512083 spawner- ahub interacti+ ucb-gener+ 1 TIMEOUT 2021-09-16T16:55:37 04:00:23 0 hrs
8579077 acompile acompile compile ucb-gener+ 1 COMPLETE 2021-09-24T15:26:32 00:00:47 0 hrs
8627076 acompile acompile compile ucb-gener+ 24 CANCELLE 2021-10-04T12:17:30 00:10:03 0 hrs
Expand Down
Loading