The Ops Manager is responsible for facilitating workloads such as backing up data, monitoring database performance and more. To make your multi-cluster Ops Manager and the Application Database deployment resilient to entire data center or zone failures, deploy the Ops Manager Application and the Application Database on multiple Kubernetes clusters.
Prerequisites
Before you begin the following procedure, perform the following actions:
Install
kubectl.Complete the GKE Clusters procedure or the equivalent.
Complete the TLS Certificates procedure or the equivalent.
Complete the Istio Service mesh procedure or the equivalent.
Complete the Deploy the MongoDB Operator procedure.
Set the required environment variables as follows:
# This script builds on top of the environment configured in the setup guides. # It depends (uses) the following env variables defined there to work correctly. # If you don't use the setup guide to bootstrap the environment, then define them here. # ${K8S_CLUSTER_0_CONTEXT_NAME} # ${K8S_CLUSTER_1_CONTEXT_NAME} # ${K8S_CLUSTER_2_CONTEXT_NAME} # ${OM_NAMESPACE} # Defaults for the test RustFS S3 storage. # If you use your own S3 storage - override any of these. export S3_OPLOG_BUCKET_NAME="${S3_OPLOG_BUCKET_NAME:-s3-oplog-store}" export S3_SNAPSHOT_BUCKET_NAME="${S3_SNAPSHOT_BUCKET_NAME:-s3-snapshot-store}" export S3_ENDPOINT="${S3_ENDPOINT:-rustfs.rustfs.svc.cluster.local}" export S3_ACCESS_KEY="${S3_ACCESS_KEY:-rustfsadmin}" export S3_SECRET_KEY="${S3_SECRET_KEY:-rustfsadmin123}" export OPS_MANAGER_VERSION="8.0.5" export APPDB_VERSION="8.0.5-ent"
Source Code
You can find all included source code in the MongoDB Kubernetes Operator repository.
Procedure
Generate TLS certificates.
kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" apply -f - <<EOF apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: om-cert spec: dnsNames: - om-svc.${OM_NAMESPACE}.svc.cluster.local duration: 240h0m0s issuerRef: name: my-ca-issuer kind: ClusterIssuer renewBefore: 120h0m0s secretName: cert-prefix-om-cert usages: - server auth - client auth --- apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: om-db-cert spec: dnsNames: - "*.${OM_NAMESPACE}.svc.cluster.local" duration: 240h0m0s issuerRef: name: my-ca-issuer kind: ClusterIssuer renewBefore: 120h0m0s secretName: cert-prefix-om-db-cert usages: - server auth - client auth EOF
Install Ops Manager.
At this point, you have prepared the environment and the Kubernetes Operator to deploy the Ops Manager resource.
Create the necessary credentials for the Ops Manager admin user that the Kubernetes Operator will create after deploying the Ops Manager Application instance:
1 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" --namespace "${OM_NAMESPACE}" create secret generic om-admin-user-credentials \ 2 --from-literal=Username="admin" \ 3 --from-literal=Password="Passw0rd@" \ 4 --from-literal=FirstName="Jane" \ 5 --from-literal=LastName="Doe" Deploy the simplest
MongoDBOpsManagercustom resource possible (with TLS enabled) on a single member cluster, which is also known as the operator cluster.This deployment is almost the same as the deployment for the single-cluster mode, but with
spec.topologyandspec.applicationDatabase.topologyset toMultiCluster.Deploying this way shows that a single Kubernetes cluster deployment is a special case of a multi-Kubernetes cluster deployment on a single Kubernetes member cluster. You can start deploying the Ops Manager Application and the Application Database on as many Kubernetes clusters as necessary from the beginning, and don't have to start with the deployment with only a single member Kubernetes cluster.
At this point, you have prepared the Ops Manager deployment to span more than one Kubernetes cluster, which you will do later in this procedure.
1 kubectl apply --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" -f - <<EOF 2 apiVersion: mongodb.com/v1 3 kind: MongoDBOpsManager 4 metadata: 5 name: om 6 spec: 7 topology: MultiCluster 8 version: "${OPS_MANAGER_VERSION}" 9 adminCredentials: om-admin-user-credentials 10 externalConnectivity: 11 type: LoadBalancer 12 security: 13 certsSecretPrefix: cert-prefix 14 tls: 15 ca: ca-issuer 16 clusterSpecList: 17 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 18 members: 1 19 applicationDatabase: 20 version: "${APPDB_VERSION}" 21 topology: MultiCluster 22 security: 23 certsSecretPrefix: cert-prefix 24 tls: 25 ca: ca-issuer 26 clusterSpecList: 27 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 28 members: 3 29 backup: 30 enabled: false 31 EOF Wait for the Kubernetes Operator to pick up the work and reach the
status.applicationDatabase.phase=Pendingstate. Wait for both the Application Database and Ops Manager deployments to complete.1 echo "Waiting for Application Database to reach Pending phase..." 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.applicationDatabase.phase}'=Pending opsmanager/om --timeout=30s Deploy Ops Manager. The Kubernetes Operator deploys Ops Manager by performing the following steps. It:
Deploys the Application Database's replica set nodes and waits for the MongoDB processes in the replica set to start running.
Deploys the Ops Manager Application instance with the Application Database's connection string and waits for it to become ready.
Adds the Monitoring MongoDB Agent containers to each Application Database's Pod.
Waits for both the Ops Manager Application and the Application Database Pods to start running.
1 echo "Waiting for Application Database to reach Running phase..." 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.applicationDatabase.phase}'=Running opsmanager/om --timeout=1200s 3 echo; echo "Waiting for Ops Manager to reach Running phase..." 4 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.opsManager.phase}'=Running opsmanager/om --timeout=1200s 5 echo; echo "MongoDBOpsManager resource" 6 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get opsmanager/om 7 echo; echo "Pods running in cluster ${K8S_CLUSTER_0_CONTEXT_NAME}" 8 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods 9 echo; echo "Pods running in cluster ${K8S_CLUSTER_1_CONTEXT_NAME}" 10 kubectl --context "${K8S_CLUSTER_1_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods Now that you have deployed a single-member cluster in a multi-cluster mode, you can reconfigure this deployment to span more than one Kubernetes cluster.
On the second member cluster, deploy two additional Application Database replica set members and one additional instance of the Ops Manager Application:
1 kubectl apply --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" -f - <<EOF 2 apiVersion: mongodb.com/v1 3 kind: MongoDBOpsManager 4 metadata: 5 name: om 6 spec: 7 topology: MultiCluster 8 version: "${OPS_MANAGER_VERSION}" 9 adminCredentials: om-admin-user-credentials 10 externalConnectivity: 11 type: LoadBalancer 12 security: 13 certsSecretPrefix: cert-prefix 14 tls: 15 ca: ca-issuer 16 clusterSpecList: 17 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 18 members: 1 19 - clusterName: "${K8S_CLUSTER_1_CONTEXT_NAME}" 20 members: 1 21 applicationDatabase: 22 version: "${APPDB_VERSION}" 23 topology: MultiCluster 24 security: 25 certsSecretPrefix: cert-prefix 26 tls: 27 ca: ca-issuer 28 clusterSpecList: 29 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 30 members: 3 31 - clusterName: "${K8S_CLUSTER_1_CONTEXT_NAME}" 32 members: 2 33 backup: 34 enabled: false 35 EOF Wait for the Kubernetes Operator to pick up the work (pending phase):
1 echo "Waiting for Application Database to reach Pending phase..." 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.applicationDatabase.phase}'=Pending opsmanager/om --timeout=30s 3 4 echo "Waiting for Ops Manager to reach Pending phase..." 5 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.opsManager.phase}'=Pending opsmanager/om --timeout=600s Wait for the Kubernetes Operator to finish deploying all components:
1 echo "Waiting for Application Database to reach Running phase..." 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.applicationDatabase.phase}'=Running opsmanager/om --timeout=1200s 3 echo; echo "Waiting for Ops Manager to reach Running phase..." 4 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.opsManager.phase}'=Running opsmanager/om --timeout=1200s 5 echo; echo "MongoDBOpsManager resource" 6 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get opsmanager/om 7 echo; echo "Pods running in cluster ${K8S_CLUSTER_0_CONTEXT_NAME}" 8 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods 9 echo; echo "Pods running in cluster ${K8S_CLUSTER_1_CONTEXT_NAME}" 10 kubectl --context "${K8S_CLUSTER_1_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods
Enable backup.
In a multi-Kubernetes cluster deployment of the Ops Manager Application, you can configure only S3-based backup storage. This procedure refers to S3_* defined in env_variables.sh.
Optional. Deploy S3-compatible storage for testing.
This step provides a script that deploys a simple RustFS instance for testing purposes. You can skip this step if you have AWS S3 or other S3-compatible buckets available. Adjust the
S3_*variables accordingly in env_variables.sh in this case.Note
RustFS is only for testing and isn't suitable for production. For production, use your own S3 storage.
1 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" create namespace "${RUSTFS_NAMESPACE}" --dry-run=client -o yaml | \ 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" apply -f - 3 4 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${RUSTFS_NAMESPACE}" delete job rustfs-create-buckets --ignore-not-found=true || true 5 6 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${RUSTFS_NAMESPACE}" apply -f - <<EOF 7 apiVersion: cert-manager.io/v1 8 kind: Certificate 9 metadata: 10 name: rustfs-cert 11 spec: 12 dnsNames: 13 - rustfs.${RUSTFS_NAMESPACE}.svc.cluster.local 14 duration: 240h0m0s 15 issuerRef: 16 name: my-ca-issuer 17 kind: ClusterIssuer 18 renewBefore: 120h0m0s 19 secretName: rustfs-tls 20 usages: 21 - server auth 22 --- 23 apiVersion: v1 24 kind: Service 25 metadata: 26 name: rustfs 27 labels: 28 app: rustfs 29 spec: 30 selector: 31 app: rustfs 32 ports: 33 - name: s3-https 34 port: 443 35 targetPort: 9000 36 - name: s3 37 port: 9000 38 targetPort: 9000 39 --- 40 apiVersion: apps/v1 41 kind: Deployment 42 metadata: 43 name: rustfs 44 labels: 45 app: rustfs 46 spec: 47 replicas: 1 48 selector: 49 matchLabels: 50 app: rustfs 51 template: 52 metadata: 53 labels: 54 app: rustfs 55 annotations: 56 # RustFS needs no mesh sidecar; keep behavior identical with and 57 # without Istio. 58 sidecar.istio.io/inject: "false" 59 spec: 60 securityContext: 61 runAsNonRoot: true 62 runAsUser: 10001 63 runAsGroup: 10001 64 fsGroup: 10001 65 seccompProfile: 66 type: RuntimeDefault 67 containers: 68 - name: rustfs 69 image: quay.io/rustfs/rustfs:1.0.0 70 securityContext: 71 allowPrivilegeEscalation: false 72 capabilities: 73 drop: ["ALL"] 74 runAsNonRoot: true 75 env: 76 - name: RUSTFS_ACCESS_KEY 77 value: "${S3_ACCESS_KEY}" 78 - name: RUSTFS_SECRET_KEY 79 value: "${S3_SECRET_KEY}" 80 - name: RUSTFS_VOLUMES 81 value: /data 82 - name: RUSTFS_ADDRESS 83 value: 0.0.0.0:9000 84 - name: RUSTFS_CONSOLE_ENABLE 85 value: "false" 86 - name: RUSTFS_TLS_PATH 87 value: /opt/tls 88 ports: 89 - name: s3 90 containerPort: 9000 91 readinessProbe: 92 httpGet: 93 scheme: HTTPS 94 path: /health 95 port: 9000 96 initialDelaySeconds: 5 97 periodSeconds: 3 98 volumeMounts: 99 - name: data 100 mountPath: /data 101 - name: tls 102 mountPath: /opt/tls 103 readOnly: true 104 volumes: 105 - name: data 106 emptyDir: {} 107 - name: tls 108 secret: 109 secretName: rustfs-tls 110 items: 111 - key: tls.crt 112 path: rustfs_cert.pem 113 - key: tls.key 114 path: rustfs_key.pem 115 --- 116 apiVersion: batch/v1 117 kind: Job 118 metadata: 119 name: rustfs-create-buckets 120 spec: 121 backoffLimit: 1 122 template: 123 metadata: 124 annotations: 125 # istio-proxy keeps running after the aws-cli container exits, so an 126 # injected Job never reaches Complete. 127 sidecar.istio.io/inject: "false" 128 spec: 129 restartPolicy: Never 130 containers: 131 - name: aws-cli 132 image: public.ecr.aws/aws-cli/aws-cli:latest 133 env: 134 - name: AWS_ACCESS_KEY_ID 135 value: "${S3_ACCESS_KEY}" 136 - name: AWS_SECRET_ACCESS_KEY 137 value: "${S3_SECRET_KEY}" 138 - name: AWS_DEFAULT_REGION 139 value: us-east-1 140 - name: S3_ENDPOINT 141 value: "https://${S3_ENDPOINT}" 142 - name: S3_OPLOG_BUCKET_NAME 143 value: "${S3_OPLOG_BUCKET_NAME}" 144 - name: S3_SNAPSHOT_BUCKET_NAME 145 value: "${S3_SNAPSHOT_BUCKET_NAME}" 146 command: ["sh", "-ec"] 147 args: 148 - | 149 aws_opts="--endpoint-url \$S3_ENDPOINT --no-verify-ssl --cli-connect-timeout 5 --cli-read-timeout 10" 150 attempt=0 151 until aws \$aws_opts s3api list-buckets; do 152 attempt=\$((attempt + 1)) 153 if [ "\$attempt" -ge 24 ]; then 154 echo "RustFS endpoint \$S3_ENDPOINT not reachable after \$attempt attempts" >&2 155 exit 1 156 fi 157 sleep 5 158 done 159 aws \$aws_opts s3api create-bucket --bucket "\$S3_OPLOG_BUCKET_NAME" 160 aws \$aws_opts s3api create-bucket --bucket "\$S3_SNAPSHOT_BUCKET_NAME" 161 EOF 162 163 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${RUSTFS_NAMESPACE}" wait --for=condition=available deployment/rustfs --timeout=300s 164 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${RUSTFS_NAMESPACE}" wait --for=condition=complete job/rustfs-create-buckets --timeout=240s Before you configure and enable backup, create secrets:
s3-access-secret- contains S3 credentials.s3-ca-cert- contains a CA certificate that issued the bucket's server certificate. The sample RustFS deployment in this procedure serves a cert-manager-issued certificate, so the secret contains the CA certificate from therustfs-tlssecret. Because this CA certificate is not publicly trusted, you must provide it so that Ops Manager can trust the connection.
If you use publicly trusted certificates, you may skip this step and remove the values from the
spec.backup.s3Stores.customCertificateSecretRefsandspec.backup.s3OpLogStores.customCertificateSecretRefssettings.1 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" create secret generic s3-access-secret \ 2 --from-literal=accessKey="${S3_ACCESS_KEY}" \ 3 --from-literal=secretKey="${S3_SECRET_KEY}" 4 5 # RustFS serves a cert-manager certificate; OM must trust the CA that signed it. 6 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" create secret generic s3-ca-cert \ 7 --from-literal=ca.crt="$(kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${RUSTFS_NAMESPACE}" get secret rustfs-tls -o jsonpath="{.data['ca\.crt']}" | base64 --decode)"
Re-deploy Ops Manager with backup enabled.
The Kubernetes Operator can configure and deploy all components, the Ops Manager Application, the Backup Daemon instances, and the Application Database's replica set nodes in any combination on any member clusters for which you configure the Kubernetes Operator.
To illustrate the flexibility of the multi-Kubernetes cluster deployment configuration, deploy only one Backup Daemon instance on the third member cluster and specify zero Backup Daemon members for the first and second clusters.
1 kubectl apply --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" -f - <<EOF 2 apiVersion: mongodb.com/v1 3 kind: MongoDBOpsManager 4 metadata: 5 name: om 6 spec: 7 topology: MultiCluster 8 version: "${OPS_MANAGER_VERSION}" 9 adminCredentials: om-admin-user-credentials 10 externalConnectivity: 11 type: LoadBalancer 12 security: 13 certsSecretPrefix: cert-prefix 14 tls: 15 ca: ca-issuer 16 clusterSpecList: 17 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 18 members: 1 19 backup: 20 members: 0 21 - clusterName: "${K8S_CLUSTER_1_CONTEXT_NAME}" 22 members: 1 23 backup: 24 members: 0 25 - clusterName: "${K8S_CLUSTER_2_CONTEXT_NAME}" 26 members: 0 27 backup: 28 members: 1 29 applicationDatabase: 30 version: "${APPDB_VERSION}" 31 topology: MultiCluster 32 security: 33 certsSecretPrefix: cert-prefix 34 tls: 35 ca: ca-issuer 36 clusterSpecList: 37 - clusterName: "${K8S_CLUSTER_0_CONTEXT_NAME}" 38 members: 3 39 - clusterName: "${K8S_CLUSTER_1_CONTEXT_NAME}" 40 members: 2 41 backup: 42 enabled: true 43 s3Stores: 44 - name: my-s3-block-store 45 s3SecretRef: 46 name: "s3-access-secret" 47 pathStyleAccessEnabled: true 48 s3BucketEndpoint: "${S3_ENDPOINT}" 49 s3BucketName: "${S3_SNAPSHOT_BUCKET_NAME}" 50 customCertificateSecretRefs: 51 - name: s3-ca-cert 52 key: ca.crt 53 s3OpLogStores: 54 - name: my-s3-oplog-store 55 s3SecretRef: 56 name: "s3-access-secret" 57 s3BucketEndpoint: "${S3_ENDPOINT}" 58 s3BucketName: "${S3_OPLOG_BUCKET_NAME}" 59 pathStyleAccessEnabled: true 60 customCertificateSecretRefs: 61 - name: s3-ca-cert 62 key: ca.crt 63 EOF Wait until the Kubernetes Operator finishes its configuration:
1 echo; echo "Waiting for Backup to reach Running phase..." 2 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.backup.phase}'=Running opsmanager/om --timeout=1200s 3 echo "Waiting for Application Database to reach Running phase..." 4 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.applicationDatabase.phase}'=Running opsmanager/om --timeout=1200s 5 echo; echo "Waiting for Ops Manager to reach Running phase..." 6 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" wait --for=jsonpath='{.status.opsManager.phase}'=Running opsmanager/om --timeout=1200s 7 echo; echo "MongoDBOpsManager resource" 8 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get opsmanager/om 9 echo; echo "Pods running in cluster ${K8S_CLUSTER_0_CONTEXT_NAME}" 10 kubectl --context "${K8S_CLUSTER_0_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods 11 echo; echo "Pods running in cluster ${K8S_CLUSTER_1_CONTEXT_NAME}" 12 kubectl --context "${K8S_CLUSTER_1_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods 13 echo; echo "Pods running in cluster ${K8S_CLUSTER_2_CONTEXT_NAME}" 14 kubectl --context "${K8S_CLUSTER_2_CONTEXT_NAME}" -n "${OM_NAMESPACE}" get pods
Create credentials for the Kubernetes Operator.
To configure credentials, you must create an Ops Manager organization, generate programmatic API keys in the Ops Manager UI, and create a secret with your Load Balancer IP. See Create Credentials for the Kubernetes Operator to learn more.