diff --git a/doc/source/admin/cluster.md b/doc/source/admin/cluster.md index d3e87d96153..9b3a0829bb8 100644 --- a/doc/source/admin/cluster.md +++ b/doc/source/admin/cluster.md @@ -16,6 +16,7 @@ Galaxy is known to work with: * [HTCondor](http://research.cs.wisc.edu/htcondor/) * [Slurm](https://slurm.schedmd.com/) * [Galaxy Pulsar](#pulsar) (formerly LWR) +* [AWS Batch](https://aws.amazon.com/batch) It should also work with [any other DRM](http://www.drmaa.org/implementations.php) which implements a [DRMAA](http://www.drmaa.org) interface. If you successfully run Galaxy with a DRM not listed here, please let us know via an email to the [galaxy-dev mailing list](https://galaxyproject.org/mailing-lists/). @@ -307,6 +308,63 @@ Torque attributes can be defined in either their short (e.g. [qsub(1B)](http://c Most options available to `qsub(1b)` and `pbs_submit(3b)` are supported. Exceptions include `-o/Output_Path`, `-e/Error_Path`, and `-N/Job_Name` since these PBS job attributes are set by Galaxy. +## AWS Batch + +Runs jobs via the [AWS Batch](https://aws.amazon.com/batch/). Built on top of AWS Elastic Container Service (ECS), AWS Batch enables users to run hundreds of thousands of jobs with simple configuration. + +#### Dependencies + +AWS Batch job runner requirs AWS Elastic File System (EFS) being mounted as a shared file system that enables Galaxy and job containers to read and write files. In the best pratice, Galaxy is installed an AWS EC2 instance and an EFS is mounted to the EC2 as a local drive. Job-related paths, such as objects, jobs_directory, tool_directory and so on, need to be placed on the EFS drive. +In addition, Galaxy admin needs to configure Batch compute environment, Batch job queue and proper AWS IAM roles, and provision them as destination parameters. +AWS Batch job runner requires [boto3](https://pypi.org/project/boto3/) installed in Galaxy environment. + +#### Parameters and Configuration + +AWS Batch job runner sends jobs to Batch compute environment that is composed of either Fargate or EC2. While Fargate provides a series of lightweigt compute resources (up to 4 vcpu and 30 GB memeory), the EC2 offers more abroad choices. With `auto_platform` enabled, this runner supports mapping to the best fit type of resources based on the provisioned `vcpu` and `memory`, i.e., Fargate is preferred over EC2 when `vcpu` and `memory` don't go beyond the limits (4 and 30 gb, respectively). If the power of `GPU` is needed for a destination, a job queue built on top of GPU-enabled compute environment must be provisoned. + +```xml + + + + + + + + + true + + arn_for_Fargate_job_queue, arn_for_EC2_job_queue + + arn:aws:iam::xxxxxxxxxxxxxxxxxx + 1 + 2048 + + fs-xxxxxxxxxxxxxx + /mnt/efs/fs1 + + 1.4.0 + + true + + + true + + arn_for_gpu_job_queue + + arn:aws:iam::xxxxxxxxxxxxxxxxxx + 4 + 20000 + + 1 + fs-xxxxxxxxxxxxxx + /mnt/efs/fs1 + + + +``` + ## Submitting Jobs as the Real User Galaxy runs as a process on your server as whatever user starts the server - usually an account created for the purpose of running Galaxy. Jobs will be submitted to your cluster(s) as this user. In environments where users in Galaxy are guaranteed to be users on the underlying system (i.e. Galaxy is configured to use external authentication), it may be desirable to submit jobs to the cluster as the user logged in to Galaxy rather than Galaxy's system user.