Just use ECS please

Introduction

Kubernetes is great for a larger company, but sometimes you want to keep the upkeep and overhead down. This is where ECS (Elastic Container Service) shines. I currently work for a company whose entire engineering team is about 12 people, and I’m the sole person on the more “infrastructure” side of things. Part of my goal was to balance cost, going-public requirements, and maintainability.

ECS is AWS’s platform for container management. Tasks in ECS are equivalent to pods in Kubernetes.

Deployments

The examples and steps below are for GitHub Actions, but you can probably apply them to whatever system you’re using.

  1. Pull down the task-definition.json:

    aws ecs describe-task-definition \
      --task-definition "${{ inputs.environment }}-${{ inputs.app_name }}" \
      --region ${{ env.REGION }} --query taskDefinition > task-definition.json
  2. Combine the task definition with the new info:

    - name: Fill in the new image ID in the Amazon ECS task definition
      id: task-def
      uses: aws-actions/amazon-ecs-render-task-definition@v1
      env:
        VERSION: ${{ env.VERSION }}
      with:
        task-definition: task-definition.json
        container-name: ${{ inputs.app_name }}
        image: "${{ secrets.AWS_ACCOUNT_ID }}.dkr.ecr.${{ env.REGION }}.amazonaws.com/${{ inputs.app_name }}:${{ inputs.version || env.VERSION }}"
        environment-variables: |
          DD_VERSION=${{ inputs.version || env.VERSION }}
  3. Deploy the new task definition:

    - name: Deploy Amazon ECS task definition
      uses: aws-actions/amazon-ecs-deploy-task-definition@v1
      with:
        task-definition: ${{ steps.task-def.outputs.task-definition }}
        service: "${{ inputs.environment }}-${{ inputs.app_name }}"
        cluster: "${{ inputs.environment }}-ecs-cluster"
        wait-for-service-stability: true

Fargate

Fargate is pretty nice, especially for Java apps, because you can use a larger percentage of your system resources than on a standard instance. The general rule of thumb is that no more than 50% of your resources go to the JVM, so a 4GB heap needs an instance with at least 8GB of RAM (and you usually want to go even lower if your app isn’t well optimized). With Fargate you can align your resources much more closely with the app’s needs.

Cost

Most of the cost is on the CPU side, so keep that in mind.

Estimates

Resource Calculation Roughly
CPU (0.04048 × 730) × number of vCPUs ~$30/month per vCPU
Memory (0.004445 × 730) × GB of RAM ~$3/month per GB

You can get further savings with an AWS Savings Plan. One note: while on a 1-to-1 basis this may look more expensive than an EC2 instance, keep in mind that you’re saving the cost of managing the security of those EC2 instances’ packages and access.

Task and service CPU/memory

With Fargate you match CPU with a memory amount. Here’s a table of the current options.1

CPU Memory (GB) Other notes
1 2–8 1GB increments
2 4–16 1GB increments
4 8–30 1GB increments
8 16–60 4GB increments
16 32–120 8GB increments

Memory needs to be given as a binary value: GB × 2^10. For example, 6GB is 6 × 2^10 = 6144.

Observability

In ECS, a task is a set of one or more containers. To get logging and metrics from your main app container, you’ll need two “sidecars” (containers that live alongside your app container in the task):

  • log-router
  • datadog-agent

By doing this you lose the ability (as of early 2024) to send the logs to CloudWatch.

Datadog

While this is specifically about Datadog, I imagine it applies to other observability tooling too, give or take some small details.

Logging

Logging can be handled entirely by the Datadog agent sidecar. With the example below, it will automatically pull in all logs from the container. You need to use the awsfirelens log driver and amazon/aws-for-fluent-bit.

APM

This depends on your application’s language, so I’d recommend looking at Datadog’s documentation for it. Because the datadog-agent is set up as a sidecar, everything should tie together without much effort.

Example

Here’s an example of the above as a Terraform config, using the terraform-aws-modules/ecs/aws//modules/service module at v5.9.1.

Container definitions

container_definitions = {
    (local.app) = {
      cpu       = var.app_cpu
      memory    = var.app_memory
      essential = true
      enable_autoscaling       = true
      autoscaling_min_capacity = 1
      autoscaling_max_capacity = 2
      autoscaling_policies     = {}
      image                    = "${ACCOUNT_ID}.dkr.ecr.${data.aws_region.current.name}.amazonaws.com/${data.aws_ecr_image.app_image.repository_name}:${data.aws_ecr_image.app_image.image_tag}"
      readonly_root_filesystem = false
      log_configuration = {
        logDriver = "awsfirelens"
        options = {
          Name           = "datadog"
          Host           = "http-intake.logs.datadoghq.com"
          apikey         = jsondecode(data.aws_secretsmanager_secret_version.selected.secret_string)["DD_API_KEY"]
          dd_service     = local.app
          dd_source      = local.dd_source
          dd_message_key = "log"
          dd_tags        = "host:${local.env}-ecs,env:${local.env}"
          TLS            = "on"
          provider       = "ecs"
          retry_limit    = "2"
        }
      }
      health_check = {
        retries = 10
        command = ["CMD-SHELL", "curl -f http://localhost:${local.app_port}${local.health_check} || exit 1"]
        timeout : 5
        interval : 10
      }
      port_mappings = [
        {
          name          = local.app
          containerPort = local.app_port
          hostPort      = local.app_port
          protocol      = "http"
        }
      ]
      environment = local.env_vars
      secrets = local.secret_vars
      docker_labels = {
        "com.datadoghq.tags.service" : local.app,
        "com.datadoghq.tags.env" : local.env,
        "com.datadoghq.ad.logs" : "[{\"source\": \"${local.dd_source}\", \"service\": \"${local.app}\"}]"
      }
    },
    datadog-agent = {
      image     = "public.ecr.aws/datadog/agent:latest"
      cpu       = 256
      memory    = 512
      essential = true
      readonly_root_filesystem = false
      environment = [
        { name = "ECS_FARGATE", value = "true" },
        { name = "DD_APM_ENABLED", value = local.dd_apm },
        { name = "DD_LOGS_ENABLED", value = "true" },
        { name = "DD_LOGS_CONFIG_CONTAINER_COLLECT_ALL", value = "true" },
        { name = "DD_CONTAINER_ENV_AS_TAGS", value = "" }
      ]
      secrets = [
        { name = "DD_API_KEY", valueFrom = data.aws_secretsmanager_secret.selected.arn }
      ]
      port_mappings = [
        {
          protocol      = "tcp"
          containerPort = 8126
        }
      ]
    },
    log-router = {
      image     = "amazon/aws-for-fluent-bit:stable"
      essential = true
      firelens_configuration = {
        type = "fluentbit"
        options = {
          enable-ecs-log-metadata = "true"
          config-file-type        = "file"
          config-file-value       = "/fluent-bit/configs/parse-json.conf"
        }
      }
      memory_reservation = 50
    }
  }

Common tasks

Restart

aws ecs update-service --cluster test-ecs-cluster --service test-app --region region-1 --force-new-deployment

Resources

Footnotes

  1. https://docs.aws.amazon.com/AmazonECS/latest/developerguide/task-cpu-memory-error.html ↩