diff --git a/docs/clusters/alpine/alpine-hardware.md b/docs/clusters/alpine/alpine-hardware.md index 72b79210b..ee5fb51df 100644 --- a/docs/clusters/alpine/alpine-hardware.md +++ b/docs/clusters/alpine/alpine-hardware.md @@ -16,15 +16,15 @@ All Alpine nodes are available to all users. For full details about node access, | Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric | | --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- | -| {{ alpine_ucb_total_64_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | -| {{ alpine_ucb_total_128_core_256GB_cpu_nodes }} Milan CPU | amilan | x86_64 AMD Milan | 2 | 128 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | +| {{ alpine_ucb_total_64_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | +| {{ alpine_ucb_total_128_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 128 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | | {{ alpine_ucb_total_48_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 2 | 48 | 1 | 21.5 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_ucb_total_64_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 1 | 64 | 1 | 16 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_ucb_total_mi100_gpu_nodes }} Milan AMD GPU | ami100 | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | AMD MI100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_ucb_total_a100_gpu_nodes }} Milan NVIDIA GPU | aa100 | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_ucb_total_gh200_gpu_nodes }} Grace CPU NVIDIA Hopper GPU | gh200

Note: these nodes are only available upon request, please submit a [support request form](https://colorado.service-now.com/req_portal?id=ucb_sc_rc_form). | ARM Neoverse V2 | 1 | 72 | 1 | 6.6 | NVIDIA Hopper GPU | 1 | 1.8 T SSD | 2x25 Gb Ethernet +RoCE | -| {{ alpine_ucb_total_acompile_nodes }} Milan CPU compile nodes | acompile | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | -| {{ alpine_ucb_total_64_core_256GB_cpu_nodes_atesting }} Milan CPU test nodes; pulls from CU amilan pool | atesting | x86_64 AMD Milan | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | +| {{ alpine_ucb_total_acompile_nodes }} AMD CPU compile nodes | acompile | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | +| {{ alpine_ucb_total_64_core_256GB_cpu_nodes_atesting }} AMD CPU test nodes; pulls from CU's `acpu` pool | atesting | x86_64 AMD | 1 or 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | | {{ alpine_ucb_total_a100_test_gpu_nodes }} Milan NVIDIA GPU testing node | aa100 (requested using the gpu-testing QoS) | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 (each split by MIG) | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_ucb_total_mi100_test_gpu_nodes }} Milan AMD GPU testing nodes; pulls from ami100 pool | ami100 (requested using the gpu-testing QoS) | x86_64 AMD Milan | 2 | 64 | 1 | {{ alpine_standard_ram_per_core }} | AMD MI100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE | @@ -39,7 +39,7 @@ All Alpine nodes are available to all users. For full details about node access, | Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric | | --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- | -| {{ alpine_amc_total_64_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | +| {{ alpine_amc_total_64_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_amc_total_64_core_1TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 1 | 64 | 1 | 16 | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | | {{ alpine_amc_total_128_core_2TB_cpu_nodes }} Milan High-Memory | amem | x86_64 AMD Milan | 2 | 128 | 1 | 16 | N/A | 0 | 70G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | | {{ alpine_amc_total_a100_gpu_nodes }} Milan NVIDIA GPU | aa100 | x86_64 AMD Milan | 1 | 64 | 1 | {{ alpine_standard_ram_per_core }} | NVIDIA A100 | 3 | 416G SSD | 2x25 Gb Ethernet +RoCE | @@ -48,7 +48,7 @@ All Alpine nodes are available to all users. For full details about node access, ::: ```{note} -**CU Anschutz job submission limit:** CU Anschutz users on Alpine are subject to a campus-wide hard limit of **200 concurrent jobs** across all QoS types (including `normal`, which system-wide allows up to 1000 jobs/user). This limit applies to all job types and partitions. It was implemented to reduce queue wait times for all Anschutz users given limited core-hour availability. If you need to run large numbers of jobs, consider using [GNU Parallel](https://github.com/kf-cuanschutz/CU-Anschutz-HPC-documentation/blob/main/Office-hours-presentation-files/GNU_parallel_presentation.pdf) as a workaround. +**CU Anschutz job submission limit:** CU Anschutz users on Alpine are subject to a campus-wide hard limit of **200 concurrent jobs** across all QoS types (including `cpu-normal`, which system-wide allows up to 1000 jobs/user). This limit applies to all job types and partitions. It was implemented to reduce queue wait times for all Anschutz users given limited core-hour availability. If you need to run large numbers of jobs, consider using [GNU Parallel](https://github.com/kf-cuanschutz/CU-Anschutz-HPC-documentation/blob/main/Office-hours-presentation-files/GNU_parallel_presentation.pdf) as a workaround. ``` ### Colorado State University contribution @@ -60,8 +60,8 @@ All Alpine nodes are available to all users. For full details about node access, | Count & Type | Partition | Processor | Sockets | Cores (total) | Threads per Core | RAM per Core (GB) | GPU type | GPU count | Local Disk Capacity & Type | Fabric | | --------------------- | ------------------- | ---------------- | :-------: | :-------------: | :------------: | :-------------: | ----------- | :---------: | -------------------------- | -------------------------------------------- | -| {{ alpine_csu_total_48_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 2 | 48 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | -| {{ alpine_csu_total_32_core_256GB_cpu_nodes }} Milan General CPU | amilan | x86_64 AMD Milan | 2 | 32 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | +| {{ alpine_csu_total_48_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 48 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | HDR-100 InfiniBand (200Gb inter-node fabric) | +| {{ alpine_csu_total_32_core_256GB_cpu_nodes }} AMD CPU | acpu | x86_64 AMD | 2 | 32 | 1 | {{ alpine_standard_ram_per_core }} | N/A | 0 | 416G SSD | 2x25 Gb Ethernet +RoCE | ::: @@ -86,13 +86,13 @@ Resources are requested within jobs by passing in SLURM directives, or resource | Partition | Description | # of nodes | cores/node | RAM/core (GB) | Billing_weight/core | | --------- | ---------------------------- | ---------- | ---------- | ------------- | ------------------- | -| amilan | AMD Milan (default) | {{ alpine_total_amilan_nodes }} | 32 or 48 or 64 or 128 | {{ alpine_standard_ram_per_core }} | 1 | +| acpu | AMD CPU nodes (default) | {{ alpine_total_acpu_nodes }} | 32 or 48 or 64 or 128 | {{ alpine_standard_ram_per_core }} | 1 | | ami100 | GPU-enabled (3x AMD MI100) | {{ alpine_total_ami100_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.13 | | aa100 | GPU-enabled (3x NVIDIA A100)4. For select nodes, MIG has been enabled providing 6x 20 GB NVIDIA A100 MIG instances. | {{ alpine_total_aa100_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.13 | | al40 | GPU-enabled (3x NVIDIA L40)4 | {{ alpine_total_al40_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | 6.13 | | amem1 | High-memory | {{ alpine_total_amem_nodes }} | 48 or 64 or 128 | 162 | 4.0 | -| acompile | AMD Milan compile nodes | {{ alpine_total_acompile_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | N/A | -| atesting | AMD Milan test nodes | {{ alpine_total_atesting_cpu_nodes }}; Pulls from CU amilan pool | 64 | {{ alpine_standard_ram_per_core }} | 0.025 | +| acompile | AMD CPU compile nodes | {{ alpine_total_acompile_nodes }} | 64 | {{ alpine_standard_ram_per_core }} | N/A | +| atesting | AMD CPU test nodes | {{ alpine_total_atesting_cpu_nodes }}; Pulls from CU's `acpu` pool | 64 | {{ alpine_standard_ram_per_core }} | 0.025 | | gh200 | NVIDIA Grace-Hopper (GH200) nodes

Note: this partition is only available upon request, please submit a [support request form](https://colorado.service-now.com/req_portal?id=ucb_sc_rc_form). | {{ alpine_ucb_total_gh200_gpu_nodes }} | 72 | 6.65 | Billed at roughly twice the rate of our A100s | ```{important} @@ -110,7 +110,7 @@ Resources are requested within jobs by passing in SLURM directives, or resource All users, regardless of institution, should specify partitions as follows: ```bash ---partition=amilan +--partition=acpu --partition=aa100 --partition=ami100 --partition=al40 @@ -119,14 +119,14 @@ All users, regardless of institution, should specify partitions as follows: ### Quality of Service (qos) -**Quality of Service or QoS is used to constrain or modify the characteristics that a job can have.** For example, by selecting the `long` QoS, a user can place the job in a **lower priority queue** with a max wall time increased from 24 hours to 7 days. +**Quality of Service or QoS is used to constrain or modify the characteristics that a job can have.** For example, by selecting the `cpu-long` QoS, a user can place the job in a **lower priority queue** with a max wall time increased from 24 hours to 7 days. #### Available QoS for Alpine: | QOS name | Description | Max walltime | Max jobs/user | Max hardware/user | Valid Partitions | | ----------- | -------------------------- | --------------- | ------------- | ------------------ | ---------------- | -| normal | Standard QoS for non-testing partitions | 1 day | 1000 | 128 nodes | amilan | -| long | Longer wall times | 7 days | 200 | 20 nodes | amilan | +| cpu-normal | Standard QoS for non-testing partitions | 1 day | 1000 | 128 nodes | acpu | +| cpu-long | Longer wall times | 7 days | 200 | 20 nodes | acpu | | mem-normal | Standard QoS for High-memory jobs | 24 hours | 1000 | 256 CPU cores | amem | | mem-long | QoS for longer running High-memory jobs | 7 days | 200 | 185 CPU cores | amem | | gpu-normal | Standard QoS for GPU jobs | 24 hours | 1000 | see [Available GRES on Alpine](#available-gres-on-alpine) | aa100,ami100,al40 | @@ -142,20 +142,20 @@ All users, regardless of institution, should specify partitions as follows: `````{tab-set} :sync-group: tabset-ex-qos-req -````{tab-item} Requesting the normal partition -:sync: ex-qos-req-normal-partition +````{tab-item} Requesting the cpu-normal QoS +:sync: ex-qos-req-cpu-normal-partition ```bash ---qos=normal +--qos=cpu-normal ``` ```` -```` {tab-item} Requesting the long partition -:sync: ex-qos-req-long-partition +```` {tab-item} Requesting the cpu-long QoS +:sync: ex-qos-req-cpu-long-partition ```bash ---qos=long +--qos=cpu-long ``` ```` diff --git a/docs/clusters/alpine/quick-start.md b/docs/clusters/alpine/quick-start.md index 1b91eb2ef..78d651239 100644 --- a/docs/clusters/alpine/quick-start.md +++ b/docs/clusters/alpine/quick-start.md @@ -30,7 +30,7 @@ section. ## Cluster Summary ### Nodes The Alpine cluster is made up of different types of nodes. A general overview of these nodes is as follows: -- **CPU nodes**: {{ alpine_total_256GB_cpu_nodes }} AMD Milan compute nodes with 256 GB RAM +- **CPU nodes**: {{ alpine_total_256GB_cpu_nodes }} AMD compute nodes with 256 GB RAM - **GPU nodes**: a mixture of {{ alpine_total_gpu_nodes }} NVIDIA and AMD GPUs - **High-memory nodes**: {{ alpine_total_hi_mem_cpu_nodes }} high-memory nodes with 1 TB of memory or more diff --git a/docs/clusters/alpine/slurm_directive_ex.md b/docs/clusters/alpine/slurm_directive_ex.md index 06fa50c3f..0a3cff7ad 100644 --- a/docs/clusters/alpine/slurm_directive_ex.md +++ b/docs/clusters/alpine/slurm_directive_ex.md @@ -13,8 +13,8 @@ Below are some examples of SLURM directives that can be used in your batch scrip To run a 32-core job for 24 hours on a single Alpine CPU node: ```bash -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks=32 #SBATCH --time=24:00:00 @@ -28,11 +28,11 @@ To run a 32-core job for 24 hours on a single Alpine CPU node: To run a 56-core job (28 cores/node) across two Alpine CPU nodes in the low-priority qos for seven days: ```bash -#SBATCH --partition=amilan +#SBATCH --partition=acpu #SBATCH --nodes=2 #SBATCH --ntasks-per-node=28 #SBATCH --time=7-00:00:00 -#SBATCH --qos=long +#SBATCH --qos=cpu-long ``` ```` @@ -72,16 +72,16 @@ To run a 42-core job for 2 hours on a single Alpine NVIDIA GPU node, using 2 40 ## Full Example Job Script -Run a 1-hour job on 4 cores on an Alpine CPU node with the normal qos that runs a python script using a custom conda environment. +Run a 1-hour job on 4 cores on an Alpine CPU node with the `cpu-normal` QoS that runs a python script using a custom conda environment. ``` #!/bin/bash -#SBATCH --partition=amilan +#SBATCH --partition=acpu #SBATCH --job-name=example-job #SBATCH --output=example-job.%j.out #SBATCH --time=01:00:00 -#SBATCH --qos=normal +#SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks=4 #SBATCH --mail-type=ALL diff --git a/docs/compute/modules.md b/docs/compute/modules.md index 12e3ff8cf..3d4cbcf5e 100644 --- a/docs/compute/modules.md +++ b/docs/compute/modules.md @@ -102,8 +102,8 @@ loads Anaconda into the environment is provided below: #SBATCH --nodes=1 #SBATCH --time=00:01:00 #SBATCH --ntasks=1 -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --job-name=test-job #SBATCH --output=test-job.%j.out diff --git a/docs/compute/monitoring-resources.md b/docs/compute/monitoring-resources.md index 5482d2388..68ddc1d9f 100644 --- a/docs/compute/monitoring-resources.md +++ b/docs/compute/monitoring-resources.md @@ -102,15 +102,15 @@ This will display the output: job stats for user ralphie over past 35 days jobid jobname partition qos account cpus state start-date-time elapsed wait ------------------------------------------------------------------------------------------------------------------- -8483382 sys/dash amilan normal ucb-gener+ 1 TIMEOUT 2021-09-14T09:32:09 01:00:16 0 hrs -8487254 test.sh amilan normal ucb-gener+ 1 COMPLETE 2021-09-14T13:21:12 00:00:02 0 hrs +8483382 sys/dash acpu cpu-normal ucb-gener+ 1 TIMEOUT 2021-09-14T09:32:09 01:00:16 0 hrs +8487254 test.sh acpu cpu-normal ucb-gener+ 1 COMPLETE 2021-09-14T13:21:12 00:00:02 0 hrs 8487256 interact ahub interacti+ ucb-gener+ 1 TIMEOUT 2021-09-14T13:22:11 12:00:22 0 hrs 8508557 acompile acompile compile ucb-gener+ 2 COMPLETE 2021-09-16T10:41:45 00:00:00 0 hrs -8508561 test.sh amilan normal ucb-gener+ 24 CANCELLE 2021-09-22T10:07:03 00:00:00 143 hrs -8508569 test amilan normal ucb-gener+ 4096 FAILED 2021-09-16T10:42:46 00:00:00 0 hrs -8508575 test amilan normal ucb-gener+ 8192 FAILED 2021-09-16T10:43:17 00:00:00 0 hrs -8508593 test amilan normal ucb-gener+ 4096 CANCELLE 2021-09-16T10:44:47 00:00:00 0 hrs -8508604 test amilan normal ucb-gener+ 2048 CANCELLE 2021-09-16T10:45:40 00:00:00 0 hrs +8508561 test.sh acpu cpu-normal ucb-gener+ 24 CANCELLE 2021-09-22T10:07:03 00:00:00 143 hrs +8508569 test acpu cpu-normal ucb-gener+ 4096 FAILED 2021-09-16T10:42:46 00:00:00 0 hrs +8508575 test acpu cpu-normal ucb-gener+ 8192 FAILED 2021-09-16T10:43:17 00:00:00 0 hrs +8508593 test acpu cpu-normal ucb-gener+ 4096 CANCELLE 2021-09-16T10:44:47 00:00:00 0 hrs +8508604 test acpu cpu-normal ucb-gener+ 2048 CANCELLE 2021-09-16T10:45:40 00:00:00 0 hrs 8512083 spawner- ahub interacti+ ucb-gener+ 1 TIMEOUT 2021-09-16T16:55:37 04:00:23 0 hrs 8579077 acompile acompile compile ucb-gener+ 1 COMPLETE 2021-09-24T15:26:32 00:00:47 0 hrs 8627076 acompile acompile compile ucb-gener+ 24 CANCELLE 2021-10-04T12:17:30 00:10:03 0 hrs diff --git a/docs/conf.py b/docs/conf.py index 054a45da4..3d4a87e40 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -56,7 +56,7 @@ 'alpine_csu_total_32_core_256GB_cpu_nodes': '49', # Alpine hardware page, partition section substitutions - 'alpine_total_amilan_nodes': '403', + 'alpine_total_acpu_nodes': '403', 'alpine_total_ami100_nodes': '8', 'alpine_total_aa100_nodes': '12', 'alpine_total_al40_nodes': '3', diff --git a/docs/open_ondemand/configuring_apps.md b/docs/open_ondemand/configuring_apps.md index 5a2177275..f53f2ab7d 100644 --- a/docs/open_ondemand/configuring_apps.md +++ b/docs/open_ondemand/configuring_apps.md @@ -29,7 +29,7 @@ Unfortunately, specifying these options can be overwhelming! To help users make | Cluster | The HPC cluster you would like to run on. Possible options are [Alpine](../clusters/alpine/index.md) and [Blanca](../clusters/blanca/blanca.md). | | Account | The account you would like to use. If you do not have a project allocation, then CU Boulder users specify `ucb-general`; CSU users specify `csu-general`; RMACC users specify `rmacc-general`; and AMC users provide `amc-general`. If you have a project allocation you can use this allocation e.g. `ucbXXX_asc1`. Blanca users should use their `blanca-` partition name. | | Partition | Specifies a particular node type to use. For example, you can provide `ahub` for quicker access or utilize another [partition on Alpine](../clusters/alpine/alpine-hardware.md#partitions). Blanca users should use their `blanca-` partition. | -| QoS Name | Quality of Service (QoS) constrains or modifies certain job characteristics. On most Alpine partitions you can specify `normal` for jobs of up to 24 hours and `long` for jobs of up to 7 days in duration. For more information see [Alpine QoS](../clusters/alpine/alpine-hardware.md#quality-of-service-qos). Blanca users should specify their `blanca-` partition name for QoS. | +| QoS Name | Quality of Service (QoS) constrains or modifies certain job characteristics. On most Alpine partitions you can specify `cpu-normal` for jobs of up to 24 hours and `cpu-long` for jobs of up to 7 days in duration. For more information see [Alpine QoS](../clusters/alpine/alpine-hardware.md#quality-of-service-qos). Blanca users should specify their `blanca-` partition name for QoS. | | Time| The duration of the job, in hours. This is dependent on both the partition and the QoS on Alpine (see above). On Blanca, users may specify jobs of up to 7 days (168 hours) in duration. | | Number of cores | The number of physical CPU cores for the job. Interactive job applications may use up to 16 cores, if using the `ahub` partition. All jobs are limited to a single compute node. | | Reservation | A reservation reserves resources for jobs being executed by select users and/or accounts. Reservations are rare on our system, but can sometimes be granted for courses utilizing HPC resources or the testing of specialty hardware. | diff --git a/docs/programming/MPIBestpractices.md b/docs/programming/MPIBestpractices.md index c6b48c814..458bcaee4 100644 --- a/docs/programming/MPIBestpractices.md +++ b/docs/programming/MPIBestpractices.md @@ -115,8 +115,8 @@ Simply select the Compiler and MPI wrapper you wish to use and place it in a job #!/bin/bash #SBATCH --nodes=2 #SBATCH --time=04:00:00 -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --constraint=ib #SBATCH --ntasks=128 #SBATCH --job-name=mpi-job diff --git a/docs/running-jobs/batch-jobs.md b/docs/running-jobs/batch-jobs.md index 8ac7bcf73..ba729dc37 100644 --- a/docs/running-jobs/batch-jobs.md +++ b/docs/running-jobs/batch-jobs.md @@ -17,10 +17,10 @@ sbatch Because job scripts specify the desired resources for your job, you won't need to specify any resources on the command line. You can, however, overwrite or add any job parameter by providing the specific resource as a flag within `sbatch` command: ```bash -sbatch --partition=amilan +sbatch --partition=acpu ``` -Running this command would force your job to run on the amilan partition *no matter what your job script specified*. +Running this command would force your job to run on the `acpu` partition *no matter what your job script specified*. ## Making a Job Script diff --git a/docs/running-jobs/interactive-jobs.md b/docs/running-jobs/interactive-jobs.md index 1bfc549c4..13e3fb340 100644 --- a/docs/running-jobs/interactive-jobs.md +++ b/docs/running-jobs/interactive-jobs.md @@ -9,10 +9,10 @@ To run an interactive job on Research Computing resources, request an interactiv The primary flags we recommend users specify are the `partition` flag and the `time` flag. These flags will specify partition and amount of time for your job respectively. The `sinteractive` command is run as follows: ```bash -sinteractive --partition=amilan --time=00:10:00 --ntasks=1 --nodes=1 --qos=normal +sinteractive --partition=acpu --time=00:10:00 --ntasks=1 --nodes=1 --qos=cpu-normal ``` -This will run an interactive job to the Slurm queue that will start a terminal session that will run on one core of one node on the amilan partition for ten minutes. Once the session has started you can run any application or script you may need from the command line. For example, if you load the Python module using `module load python` and then type `python`, you will open an interactive python shell on a compute node (rather than the login nodes, which is forbidden). When you are finished with your interactive job, you can end the session by typing `exit`. If you do not end your session, the interactive job will run for the full time requested, which will use up part of your allocation. +This will run an interactive job to the Slurm queue that will start a terminal session that will run on one core of one node on the `acpu` partition for ten minutes. Once the session has started you can run any application or script you may need from the command line. For example, if you load the Python module using `module load python` and then type `python`, you will open an interactive python shell on a compute node (rather than the login nodes, which is forbidden). When you are finished with your interactive job, you can end the session by typing `exit`. If you do not end your session, the interactive job will run for the full time requested, which will use up part of your allocation. ```{seealso} Check out this [page](job-resources.md) for a list of Slurm directives that can be used with interactive jobs. diff --git a/docs/running-jobs/job-arrays.md b/docs/running-jobs/job-arrays.md index e4021e0db..20b8b0353 100644 --- a/docs/running-jobs/job-arrays.md +++ b/docs/running-jobs/job-arrays.md @@ -254,8 +254,8 @@ Here's how it works: ```bash #!/bin/bash #SBATCH --time=00:00:10 -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --job-name=Array_Example_Multiple_Files @@ -286,8 +286,8 @@ You can use this method with any program that takes command-line arguments, not ```bash #!/bin/bash #SBATCH --time=00:00:10 - #SBATCH --partition=amilan - #SBATCH --qos=normal + #SBATCH --partition=acpu + #SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks=1 #SBATCH --job-name=cars diff --git a/docs/running-jobs/job-resources.md b/docs/running-jobs/job-resources.md index a4ca1cd09..589263377 100644 --- a/docs/running-jobs/job-resources.md +++ b/docs/running-jobs/job-resources.md @@ -9,14 +9,14 @@ Job scripts, the `sbatch` command, and the `sinteractive` command support many d | Type | Description | Flag | Example | | :---------------------- | :--------------------------------------------- | :------------------------- | :---------------------------- | | [Allocation](../clusters/alpine/allocations.md) | Specify an allocation account | `--account=`
| `--account=ucb###_asc1`
| -| Partition | Specify a partition ([see table below](#partitions)) | `--partition=`
| `--partition=amilan`
| +| Partition | Specify a partition ([see table below](#partitions)) | `--partition=`
| `--partition=acpu`
| | Sending email | Receive an email at the beginning or the end of a job | `--mail-type=`
| `--mail-type=BEGIN,END`
| | Email address | Email address to receive the email | `--mail-user=`
| `--mail-user=ralphie@colorado.edu`
| | Number of nodes | The number of nodes needed to run the job | `--nodes=<#>`
| `--nodes=1`
| | Number of tasks | The ***total*** number of processes needed to run the job | `--ntasks=<#>`
| `--ntasks=4`
| | Tasks per node | The number of processes you wish to assign to each node (only needed for multi-node jobs) | `--ntasks-per-node=<#>`
| `--ntasks-per-node=4`
| | Total memory | The total memory (per node requested) required for the job.
Using `--mem` does not alter the number of cores allocated to the job, but you will be charged for the number of cores corresponding to the proportion of total memory requested.
Units of `--mem` can be specified with the suffixes: K,M,G,T (default M)| `--mem=<#>`
|`--mem=25G`
| -| Quality of service | Specify a QoS ([see table below](#quality-of-service)) | `--qos=`
| `--qos=normal`
| +| Quality of service | Specify a QoS ([see table below](#quality-of-service)) | `--qos=`
| `--qos=cpu-normal`
| | Wall time | The max amount of time your job will run for | `--time=`
| `--time=03:00:00`
| | Job Name | Name your job so you can identify it in the queue | `--job-name=`
| `--job-name=Census-Data-Analysis`
| | Job Array | Specify the range of values to use for the Job Array indexes. Learn more about [job arrays here](./job-arrays.md). | `--array=-` | `--array=0-5`
. | @@ -27,6 +27,6 @@ Nodes with the same hardware configuration are grouped into partitions. You will ## Quality of Service -Quality of Service (QoS) is used to constrain or modify the characteristics that a job can have. This could come in the form of specifying a QoS to request for a longer run time or a high priority queue for condo owned nodes. For example, by selecting the `long` QoS, a user can place the job in a lower priority queue with a max wall time increased from 24 hours to 7 days. A list of QoS codes available on Alpine can be found on our [Alpine Hardware](../clusters/alpine/alpine-hardware.md#quality-of-service-qos) page. +Quality of Service (QoS) is used to constrain or modify the characteristics that a job can have. This could come in the form of specifying a QoS to request for a longer run time or a high priority queue for condo owned nodes. For example, by selecting the `cpu-long` QoS, a user can place the job in a lower priority queue with a max wall time increased from 24 hours to 7 days. A list of QoS codes available on Alpine can be found on our [Alpine Hardware](../clusters/alpine/alpine-hardware.md#quality-of-service-qos) page. diff --git a/docs/running-jobs/running-apps-with-jobs.md b/docs/running-jobs/running-apps-with-jobs.md index bfec4a1d6..3d44144c8 100644 --- a/docs/running-jobs/running-apps-with-jobs.md +++ b/docs/running-jobs/running-apps-with-jobs.md @@ -49,10 +49,10 @@ Another method of running applications on Research Computing resources is throug You can request an interactive job by using the `sinteractive` command. Similar to `sbatch`, resources must be requested. With `sinteractive`, this is done via the command line through the use of flags. You will need to, at a minimum, include the `--partition`, `--qos`, and `--time` flags. We encourage using the `--ntasks` and `--nodes` as well, otherwise the job will default to 1 task and 1 node. ```bash -sinteractive --partition=amilan --qos=normal --time=00:10:00 --ntasks=4 --nodes=1 +sinteractive --partition=acpu --qos=cpu-normal --time=00:10:00 --ntasks=4 --nodes=1 ``` -The example above will submit an interactive job requesting the `amilan` partition with 4 cores on one node with the `normal` quality of service (QoS) for ten minutes. Once the interactive session has started, you will be provided a terminal session on a compute node. Within this session, you can run any interactive terminal application you may need from the command line. +The example above will submit an interactive job requesting the `acpu` partition with 4 cores on one node with the `cpu-normal` quality of service (QoS) for ten minutes. Once the interactive session has started, you will be provided a terminal session on a compute node. Within this session, you can run any interactive terminal application you may need from the command line. ```{important} Be careful when setting `--ntasks` and ensure you also set `--nodes`. If `--nodes` is not set, Slurm may spread your job across multiple nodes. Also, be aware that GPU-based interactive jobs must set `--nodes=1` and cannot currently run across multiple nodes. diff --git a/docs/software/gaussian.md b/docs/software/gaussian.md index 726d4410d..d21c887e8 100644 --- a/docs/software/gaussian.md +++ b/docs/software/gaussian.md @@ -31,8 +31,8 @@ __Example SMP job script:__ #!/bin/bash #SBATCH --job-name=g16-test -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks-per-node=1 #SBATCH --cpus-per-task=64 @@ -60,7 +60,7 @@ g16 -m=20gb -p=${SLURM_CPUS_PER_TASK} my_input.com ### Multi-node parallel jobs (Linda) -In order to run on more than 64 cores in the `amilan` partition on Alpine, your job will need to span multiple nodes using the Linda network parallel communication model. We advise using one Linda worker per node, with multiple (up to 64) SMP cores per node. The nodes on which `g16` will run will be determined once the job starts, before invoking `g16` The batch script example below demonstrates how to run `g16` across multiple nodes. A sample input file is at the bottom of this documentation page. +In order to run on more than 64 cores in the `acpu` partition on Alpine, your job will need to span multiple nodes using the Linda network parallel communication model. We advise using one Linda worker per node, with multiple (up to 64) SMP cores per node. The nodes on which `g16` will run will be determined once the job starts, before invoking `g16` The batch script example below demonstrates how to run `g16` across multiple nodes. A sample input file is at the bottom of this documentation page. __Example Linda Parallel job script__ @@ -68,8 +68,8 @@ __Example Linda Parallel job script__ #!/bin/bash #SBATCH --job-name=g16-test -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --nodes=2 #SBATCH --ntasks-per-node=1 #SBATCH --cpus-per-task=64 @@ -109,7 +109,7 @@ Not all G16 computations scale efficiently beyond a single node! According to th ### G16 on Alpine NVIDIA GPUs -Please see the [Gaussian GPU documentation](https://gaussian.com/running/?tabid=5)] for information on how configure Gaussian input files to run on GPUs. CURC presently does not have example job scripts for running Gaussian on GPUs. The Gaussian GPU documentation will also enable you to determine whether the A100 GPUs in Alpine's `aa100` partition will be effective for your calculations. In many cases, SMP parallelization across all of the cores in an amilan node will provide better speedup than offloading computational work to a GPU. +Please see the [Gaussian GPU documentation](https://gaussian.com/running/?tabid=5)] for information on how configure Gaussian input files to run on GPUs. CURC presently does not have example job scripts for running Gaussian on GPUs. The Gaussian GPU documentation will also enable you to determine whether the A100 GPUs in Alpine's `aa100` partition will be effective for your calculations. In many cases, SMP parallelization across all of the cores in an `acpu` node will provide better speedup than offloading computational work to a GPU. ```{warning} G16 can not use the AMD MI100 GPUs in Alpine's `ami100` partition. diff --git a/docs/software/matlab.md b/docs/software/matlab.md index fcd267864..845361340 100644 --- a/docs/software/matlab.md +++ b/docs/software/matlab.md @@ -193,8 +193,7 @@ end ``` Now all we have left to do is modify our batch script to specify that -we want to run 4 tasks on the node (we can use up to 64 cores on each -‘amilan' node on Alpine). We can also change the name of the job and the +we want to run 4 tasks on the node (one can request a larger amount of tasks by utilizing an `acpu` node on Alpine). We can also change the name of the job and the output file if we choose. ```bash diff --git a/docs/software/vasp.md b/docs/software/vasp.md index 832458eee..d1d8a533d 100644 --- a/docs/software/vasp.md +++ b/docs/software/vasp.md @@ -35,8 +35,8 @@ make ```bash #!/bin/bash -#SBATCH --partition=amilan -#SBATCH --qos=normal +#SBATCH --partition=acpu +#SBATCH --qos=cpu-normal #SBATCH --nodes=1 #SBATCH --ntasks=2 #SBATCH --time=1:00:00 diff --git a/node_counter_scripts/run.py b/node_counter_scripts/run.py index d79db07d3..4d71185f4 100644 --- a/node_counter_scripts/run.py +++ b/node_counter_scripts/run.py @@ -106,7 +106,7 @@ partition_names = get_all_partition_names() for partition_name in partition_names: - if partition_name == 'amilan*': - partition_name = 'amilan' + if partition_name == 'acpu*': + partition_name = 'acpu' num_nodes = get_num_nodes_partition(partition_name) print(f"Partition {partition_name} has {num_nodes} nodes") \ No newline at end of file