Build a Secure ECS Status Service

Deploy private ECS tasks with monitoring and secure GitHub Actions.

Introduction

30 Second Summary

A cloud engineering portfolio can look polished while leaving the hardest questions unanswered. Hiring managers still need to see what happens when traffic arrives, a deployment runs, or a task fails.

In this project, you will build a secure venue status service on AWS and deploy it through GitHub Actions. You'll demonstrate private Linux containers, automated delivery, live monitoring, and recovery from failure.

What You'll Build

Picture a hiring manager opening your load-balancer URL to see a venue status page respond from private cloud infrastructure you defined yourself.

By the end of this project, you'll have:

  • A live venue status page that a hiring manager can open through an Application Load Balancer. Two private Amazon ECS tasks on AWS Fargate serve its traffic across Availability Zones.
  • A reviewable deployment path where AWS CDK defines the network and security controls as code. A successful workflow proves that OIDC temporary credentials can deploy the service without stored AWS access keys.
  • Clear operational evidence through Amazon CloudWatch logs, utilization graphs, target-health graphs, and an unhealthy-target alarm.
  • Secret Mission: Stop one of the two running tasks while continuously checking the live URL. You'll prove that traffic stays available while ECS restores the missing capacity.

Are there any prerequisites?

You'll need AWS and GitHub accounts plus a Windows workstation with Node.js, Git, Docker Desktop, AWS CLI, and Visual Studio Code installed.

Budget under $1 for this short lab, then delete the AWS resources in the same session.

Before We Start

Before We Start...

Before the hands-on work begins, take a moment to define the secure AWS service you are building and why it matters. This gives every technical decision a clear purpose.

Set Up and Validate the Toolchain

Your workstation needs a trusted path to AWS before the service can leave your computer. A missing tool or expired login would make every later deployment check unreliable.

This step verifies the local toolchain, creates the GitHub repository, and prepares an empty project workspace. You finish with the dependencies and identifiers needed for the build.

In this step, get ready to:
  • Verify the required command-line tools.
  • Authenticate to AWS and record the repository identity.
  • Create the local project workspace and install its dependencies.
Verify the workstation tools

The project uses Node.js 24.21.0, Git, Docker, and AWS CLI v2. Checking all four now separates a ready workstation from a version or installation problem.

  • Press the Windows key to open Search.
  • Type PowerShell into the search field.
  • Press Enter to open PowerShell.
  • Display each installed version by running these commands:
node --version
git --version
docker --version
aws --version

✔️ All four tools are ready

Your workstation is ready when Node.js reports v24.21.0, Git and Docker each report a version, and AWS CLI reports version 2.

ⓧ A pinned tool is outdated

Use the pinned Node.js release when your installed version differs from v24.21.0. Download it from the official Node.js release page.

  • Upgrade AWS CLI to version 2 by running this official Windows installer command:
irm https://awscli.amazonaws.com/v2/install.ps1 | iex
  • Restart PowerShell after each installation.
  • Run the four version checks again.

ⓧ A command is not found

Install each missing product from its official source. Restart PowerShell before checking the commands again.

irm https://awscli.amazonaws.com/v2/install.ps1 | iex
  • Restart PowerShell after the installations finish.
  • Run the four version checks again.

Still missing a version?

Confirm that PowerShell was restarted after installation. Check that Docker Desktop finished starting before Step 2.

Help me diagnose the toolchain check.

Authenticate and identify the repository

AWS login opens a browser-based sign-in flow and returns short-lived credentials to PowerShell. You do not store an access key in this project.

  • Start the AWS sign-in flow by running:
aws login
  • Complete the browser sign-in for your AWS account.
  • Return to PowerShell after authentication succeeds.
  • Set the project Region by running:
aws configure set region us-west-2

Before you check the identity, which account and principal do you expect the signed-in session to return?

  • Confirm the authenticated identity by running:
aws sts get-caller-identity

You should see your AWS account identifier and current principal ARN. This proves the CLI can make authenticated requests without exposing a secret.

Identity check failed?

Repeat aws login if the session expired. Confirm the signed-in identity can call AWS Security Token Service.

Help me troubleshoot AWS authentication.

  • Open GitHub in your browser.
  • Select New repository from the GitHub creation menu.
  • Enter secure-venue-status-service in the Repository name field.
  • Select Public visibility.
  • Click Create repository.
  • Record your GitHub owner name: your GitHub owner name.
  • Load the public repository record by running:
$repo = Invoke-RestMethod -Uri "https://api.github.com/repos/[[GITHUB_OWNER="your GitHub owner name"]]/secure-venue-status-service"
$repo.owner.id
$repo.id
  • Record the first number as your GitHub owner ID.
  • Record the second number as your GitHub repository ID.
Prepare the local workspace

The empty workspace keeps setup separate from application code. The installed packages provide the CDK and TypeScript toolchain used in later steps.

  • Move to your Desktop by running:
cd ~/Desktop
  • Create the project folder and move into it by running:
mkdir secure-venue-status-service
cd secure-venue-status-service
  • Create the Node.js package manifest and install the pinned dependencies by running:
npm init -y
npm install aws-cdk-lib@2.273.0 constructs@10.8.1
npm install --save-dev aws-cdk@2.1145.0 typescript@5.9.3 ts-node@10.9.2 @types/node@24.10.1

The install can take a minute while npm downloads the packages. The terminal returns to the prompt when the workspace is ready.

  • Confirm the workspace files by running:
Get-ChildItem

You should see node_modules, package.json, and package-lock.json inside secure-venue-status-service.

Workspace files missing?

Confirm PowerShell is inside ~/Desktop/secure-venue-status-service. A network or permissions error can interrupt npm before all packages are installed.

Help me repair the workspace setup.

Your toolchain, cloud session, repository identity, and local workspace are ready. Next, you'll turn that empty folder into a Linux container you can test on your workstation.

Run the Linux Container Locally

Your workstation and empty project workspace are ready. Now you'll build the service and prove that Docker can run it as a healthy, non-root Linux container.

A local test keeps application faults separate from AWS networking problems. The browser, health route, logs, and runtime identity give you evidence before deployment.

In this step, get ready to:
  • Create and run the venue status service.
  • Package the service as a Linux container.
  • Verify its page, health response, logs, and non-root runtime.
Create the local service

The service uses the built-in Node.js HTTP server. It exposes a human-facing page plus the machine-readable /health route used by Docker and the load balancer.

  • Create the app folder by running:
mkdir app
Get-ChildItem

You should see app beside package.json.

  • Press the Windows key to open Search.
  • Type Visual Studio Code.
  • Press Enter to open Visual Studio Code.
  • Select File from the top menu.
  • Select Open Folder.
  • Open Desktop/secure-venue-status-service.
  • Create server.js inside the app folder.
  • Paste this service into app/server.js.
'use strict';
const http = require('node:http');

// Keep the service port aligned with Docker and ECS.
const port = 8080;

// Show a visible marker before automated deployment is added.
const page = `<!doctype html><html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Venue Operations Status</title></head><body><main><span>Operational</span><h1>Venue systems are ready.</h1><p>Local container validated</p></main></body></html>`;

// Send every response with an explicit status and content type.
function send(response, statusCode, contentType, body) {
  response.writeHead(statusCode, { 'Content-Type': contentType });
  response.end(body);
}

// Route health checks and browser traffic, then log the result.
const server = http.createServer((request, response) => {
  const startedAt = Date.now();
  const isHealth = request.url === '/health';
  const statusCode = request.url === '/' || isHealth ? 200 : 404;
  const body = isHealth ? JSON.stringify({ status: 'ok' }) : request.url === '/' ? page : JSON.stringify({ error: 'not_found' });
  send(response, statusCode, isHealth || statusCode === 404 ? 'application/json' : 'text/html; charset=utf-8', body);
  console.log(JSON.stringify({ event: 'request', path: request.url, statusCode, durationMs: Date.now() - startedAt }));
});

server.listen(port, '0.0.0.0', () => console.log(JSON.stringify({ event: 'startup', port })));

What does this service do?

  • The page constant provides the visible status page.
  • The /health route returns {"status":"ok"} for automated checks.
  • The request log records the path, status code, and response time.
  • Save app/server.js.
  • Start the service from PowerShell by running:
node .\app\server.js

You should see a startup record with port 8080. Keep this PowerShell window running.

  • Open a second PowerShell window through Windows Search.
  • Check the running health route by running:
Invoke-RestMethod -Uri "http://localhost:8080/health"

You should see status set to ok.

Health route unavailable?

Check that the first PowerShell window is still running the server. Confirm the file is named app/server.js and the request uses port 8080.

Help me debug the local service.

✔️ Awesome, I've got everything!

Your saved service starts and returns the expected health response. Double-check that you saved app/server.js before continuing.

ⓧ I'd like to double check the full code

Compare your file with the complete app/server.js below.

'use strict';
const http = require('node:http');

// Keep the service port aligned with Docker and ECS.
const port = 8080;

// Show a visible marker before automated deployment is added.
const page = `<!doctype html><html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Venue Operations Status</title></head><body><main><span>Operational</span><h1>Venue systems are ready.</h1><p>Local container validated</p></main></body></html>`;

// Send every response with an explicit status and content type.
function send(response, statusCode, contentType, body) {
  response.writeHead(statusCode, { 'Content-Type': contentType });
  response.end(body);
}

// Route health checks and browser traffic, then log the result.
const server = http.createServer((request, response) => {
  const startedAt = Date.now();
  const isHealth = request.url === '/health';
  const statusCode = request.url === '/' || isHealth ? 200 : 404;
  const body = isHealth ? JSON.stringify({ status: 'ok' }) : request.url === '/' ? page : JSON.stringify({ error: 'not_found' });
  send(response, statusCode, isHealth || statusCode === 404 ? 'application/json' : 'text/html; charset=utf-8', body);
  console.log(JSON.stringify({ event: 'request', path: request.url, statusCode, durationMs: Date.now() - startedAt }));
});

server.listen(port, '0.0.0.0', () => console.log(JSON.stringify({ event: 'startup', port })));
Package the container

The health diagnostic gives Docker a command it can run inside the image. A successful check proves the server responds from the container's own network namespace.

  • Create healthcheck.js inside the app folder.
  • Paste this diagnostic into app/healthcheck.js.
'use strict';
const http = require('node:http');

// Fail the container check unless the local health route answers quickly.
const request = http.get('http://127.0.0.1:8080/health', (response) => {
  if (response.statusCode !== 200) process.exit(1);
  response.resume();
  response.on('end', () => {
    console.log('Health check passed.');
    process.exit(0);
  });
});

request.setTimeout(2000, () => request.destroy());
request.on('error', () => process.exit(1));

How does the diagnostic work?

The script requests the loopback health route and exits successfully only after a 200 response. Docker treats any other exit code as a failed check.

  • Save app/healthcheck.js.
  • Run the diagnostic while the local server is still active:
node .\app\healthcheck.js

You should see Health check passed. in the second PowerShell window.

  • Press Ctrl+C in the first PowerShell window to stop the direct Node.js server.
  • Create Dockerfile inside the app folder.
  • Paste this image definition into app/Dockerfile.
FROM node:24.21.0-bookworm

# Copy only the runtime files into the image.
WORKDIR /app
COPY server.js healthcheck.js ./

# Run the application without root privileges.
USER node
EXPOSE 8080
HEALTHCHECK --interval=10s --timeout=3s --retries=3 CMD ["node", "healthcheck.js"]
CMD ["node", "server.js"]

What does the image enforce?

  • The pinned base image supplies Node.js 24.21.0 on Debian Bookworm.
  • The USER node instruction removes root privileges before startup.
  • The health check runs the diagnostic inside the container.
  • Save app/Dockerfile.
  • Build the local image by running:
docker build -t venue-status:local .\app

The first build can take a few minutes while Docker downloads the base layers. You should eventually see venue-status:local tagged successfully.

Image build failed?

Confirm Docker Desktop is running with Linux containers. Check that server.js, healthcheck.js, and Dockerfile are all inside app.

Help me troubleshoot the image build.

Prove and clean up the runtime
  • Start the container by running:
docker run --name venue-status -d -p 8080:8080 venue-status:local

PowerShell prints a container ID. That confirms the process started in the background.

  • Open your browser through Windows Search.
  • Enter http://localhost:8080 in the address bar.

You should see the venue page with Operational and Local container validated.

  • Query the container health route by running:
Invoke-RestMethod -Uri "http://localhost:8080/health"

You should see status set to ok again.

  • Read the structured application logs by running:
  • Inspect the runtime user and health state by running the second command:
docker logs venue-status
docker inspect --format "user={{.Config.User}} health={{.State.Health.Status}}" venue-status

You should see startup and request records followed by user=node and health=healthy after the first health check completes.

Runtime evidence incomplete?

Wait ten seconds if the health value still says starting. A missing user value means the Dockerfile did not apply USER node.

Help me interpret the container checks.

Your page, health response, request logs, and runtime identity now support the same result. The Linux image is ready for AWS.

  • Remove the temporary container by running:
docker rm --force venue-status

You should see venue-status after removal. The reusable venue-status:local image remains on your workstation.

Your local container has passed its application, health, logging, and runtime checks. Next, you'll define the private AWS network and deploy one task behind a load balancer.

Deploy One Private Task

Your local image now serves a healthy page as the non-root node user. The next goal is to place that image on Amazon ECS without giving the task a public IP address.

You'll define the network and service with AWS CDK, then deploy one AWS Fargate task behind an Application Load Balancer. Scaling that task away exposes the availability limit you fix next.

Why ECS for this project?

AWS App Runner gives you a shorter path to a public web service. However, it hides the network and security controls this project is designed to expose.

Amazon EKS adds Kubernetes administration that this service does not need. Amazon ECS keeps task placement, subnets, security groups, and load balancing visible.

In this step, get ready to:
  • Configure the TypeScript CDK application.
  • Define the private network, service, and deployment identity.
  • Deploy one task and expose its availability gap.
Configure the CDK application

The configuration connects the TypeScript compiler to the CDK entry point. Keeping the Region fixed at us-west-2 makes every deployment and cleanup command target the same environment.

  • Open the existing project in Visual Studio Code.
  • Replace package.json with this dependency and build configuration:
{
  "name": "secure-venue-status-service",
  "version": "1.0.0",
  "private": true,
  "scripts": {
    "build": "tsc"
  },
  "dependencies": {
    "aws-cdk-lib": "2.273.0",
    "constructs": "10.8.1"
  },
  "devDependencies": {
    "@types/node": "24.10.1",
    "aws-cdk": "2.1145.0",
    "ts-node": "10.9.2",
    "typescript": "5.9.3"
  }
}
  • Create tsconfig.json beside package.json.
  • Paste this compiler configuration into tsconfig.json.
{
  "compilerOptions": {
    "target": "ES2022",
    "module": "commonjs",
    "strict": true,
    "esModuleInterop": true,
    "skipLibCheck": true,
    "outDir": "dist"
  },
  "include": ["bin/**/*.ts", "lib/**/*.ts"]
}
  • Create cdk.json beside package.json.
  • Paste this CDK entry-point configuration into cdk.json.
{
  "app": "npx ts-node --prefer-ts-exts bin/app.ts"
}
  • Save all three configuration files.
  • Confirm they exist by running:
Get-ChildItem package.json, tsconfig.json, cdk.json

PowerShell should list all three files from secure-venue-status-service.

  • Create bin and lib folders by running:
mkdir bin, lib
  • Create bin/app.ts.
  • Paste this CDK application entry point into bin/app.ts.
import * as cdk from 'aws-cdk-lib';
import { VenueStatusStack } from '../lib/venue-status-stack';

// Bind this deployment to the authenticated account and chosen Region.
const app = new cdk.App();
new VenueStatusStack(app, 'VenueStatusStack', {
  env: {
    account: process.env.CDK_DEFAULT_ACCOUNT,
    region: 'us-west-2',
  },
  githubOwnerId: '[[GITHUB_OWNER_ID="your GitHub owner ID"]]',
  githubRepositoryId: '[[GITHUB_REPOSITORY_ID="your GitHub repository ID"]]',
});
  • Create lib/venue-status-stack.ts.
  • Paste this minimal stack into lib/venue-status-stack.ts.
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';

export interface VenueStatusStackProps extends cdk.StackProps {
  readonly githubOwnerId: string;
  readonly githubRepositoryId: string;
}

// Start with a compilable stack before adding infrastructure helpers.
export class VenueStatusStack extends cdk.Stack {
  constructor(scope: Construct, id: string, props: VenueStatusStackProps) {
    super(scope, id, props);
  }
}
  • Save both TypeScript files.
  • Compile the minimal application by running:
npm run build

A clean return to the PowerShell prompt confirms the entry point, stack interface, and compiler configuration agree.

Initial build failed?

Run npm install if PowerShell reports a missing package. Use the reported filename and line number for TypeScript syntax problems.

Help me fix the initial CDK build.

Define the private infrastructure

The network uses public subnets only for the load balancer. Isolated application subnets reach Amazon ECR and Amazon CloudWatch through private endpoints.

  • Create lib/network.ts.
  • Paste this network helper into lib/network.ts.
import { Construct } from 'constructs';
import * as ec2 from 'aws-cdk-lib/aws-ec2';

// Build a two-AZ VPC with no NAT gateway or public task route.
export function createNetwork(scope: Construct) {
  const vpc = new ec2.Vpc(scope, 'VenueVpc', {
    ipAddresses: ec2.IpAddresses.cidr('10.20.0.0/16'),
    maxAzs: 2,
    natGateways: 0,
    subnetConfiguration: [
      { cidrMask: 24, name: 'Public', subnetType: ec2.SubnetType.PUBLIC },
      { cidrMask: 24, name: 'Application', subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
    ],
  });
  const albSecurityGroup = new ec2.SecurityGroup(scope, 'AlbSecurityGroup', { vpc });
  albSecurityGroup.addIngressRule(ec2.Peer.anyIpv4(), ec2.Port.tcp(80));
  const taskSecurityGroup = new ec2.SecurityGroup(scope, 'TaskSecurityGroup', { vpc });
  taskSecurityGroup.addIngressRule(albSecurityGroup, ec2.Port.tcp(8080));
  const endpointSecurityGroup = new ec2.SecurityGroup(scope, 'EndpointSecurityGroup', { vpc });
  endpointSecurityGroup.addIngressRule(taskSecurityGroup, ec2.Port.tcp(443));
  const endpointOptions = { open: false, securityGroups: [endpointSecurityGroup], subnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED } };
  vpc.addInterfaceEndpoint('EcrApiEndpoint', { service: ec2.InterfaceVpcEndpointAwsService.ECR, ...endpointOptions });
  vpc.addInterfaceEndpoint('EcrDockerEndpoint', { service: ec2.InterfaceVpcEndpointAwsService.ECR_DOCKER, ...endpointOptions });
  vpc.addInterfaceEndpoint('LogsEndpoint', { service: ec2.InterfaceVpcEndpointAwsService.CLOUDWATCH_LOGS, ...endpointOptions });
  vpc.addGatewayEndpoint('S3Endpoint', { service: ec2.GatewayVpcEndpointAwsService.S3 });
  return { vpc, albSecurityGroup, taskSecurityGroup };
}
  • Create lib/service.ts.
  • Paste this one-task service helper into lib/service.ts.
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';
import * as ec2 from 'aws-cdk-lib/aws-ec2';
import * as ecrAssets from 'aws-cdk-lib/aws-ecr-assets';
import * as ecs from 'aws-cdk-lib/aws-ecs';
import * as patterns from 'aws-cdk-lib/aws-ecs-patterns';
import * as elbv2 from 'aws-cdk-lib/aws-elasticloadbalancingv2';
import * as logs from 'aws-cdk-lib/aws-logs';

// Place one private Fargate task behind the public load balancer.
export function createService(scope: Construct, vpc: ec2.Vpc, albSecurityGroup: ec2.SecurityGroup, taskSecurityGroup: ec2.SecurityGroup) {
  const cluster = new ecs.Cluster(scope, 'VenueCluster', { vpc });
  const logGroup = new logs.LogGroup(scope, 'ApplicationLogGroup', { removalPolicy: cdk.RemovalPolicy.DESTROY });
  const image = new ecrAssets.DockerImageAsset(scope, 'VenueImage', { directory: 'app', platform: ecrAssets.Platform.LINUX_AMD64 });
  const loadBalancer = new elbv2.ApplicationLoadBalancer(scope, 'VenueLoadBalancer', { vpc, internetFacing: true, securityGroup: albSecurityGroup, vpcSubnets: { subnetType: ec2.SubnetType.PUBLIC } });
  const venueService = new patterns.ApplicationLoadBalancedFargateService(scope, 'VenueService', {
    cluster, loadBalancer, openListener: false, listenerPort: 80,
    assignPublicIp: false, taskSubnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
    securityGroups: [taskSecurityGroup], cpu: 256, memoryLimitMiB: 512, desiredCount: 1,
    taskImageOptions: { image: ecs.ContainerImage.fromDockerImageAsset(image), containerPort: 8080, logDriver: ecs.LogDrivers.awsLogs({ streamPrefix: 'venue-status', logGroup }) },
  });
  venueService.targetGroup.configureHealthCheck({ path: '/health', healthyHttpCodes: '200' });
  return { cluster, logGroup, venueService };
}
  • Save both helper files.
  • Confirm the helper files exist by running:
Get-ChildItem .\lib

You should see network.ts, service.ts, and venue-status-stack.ts.

  • Create lib/deployment-role.ts.
  • Paste this repository-scoped OIDC role helper into lib/deployment-role.ts.
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';
import * as iam from 'aws-cdk-lib/aws-iam';

// Trust only the recorded repository IDs and its main branch.
export function createDeploymentRole(scope: Construct, ownerId: string, repositoryId: string) {
  const provider = new iam.OpenIdConnectProvider(scope, 'GitHubOidcProvider', {
    url: 'https://token.actions.githubusercontent.com',
    clientIds: ['sts.amazonaws.com'],
  });
  const principal = new iam.OpenIdConnectPrincipal(provider, {
    StringEquals: {
      'token.actions.githubusercontent.com:aud': 'sts.amazonaws.com',
      'token.actions.githubusercontent.com:repository_owner_id': ownerId,
      'token.actions.githubusercontent.com:repository_id': repositoryId,
      'token.actions.githubusercontent.com:ref': 'refs/heads/main',
    },
    StringLike: { 'token.actions.githubusercontent.com:sub': 'repo:*' },
  });
  const role = new iam.Role(scope, 'GitHubDeploymentRole', { roleName: 'VenueStatusDeployRole', assumedBy: principal });
  role.addToPolicy(new iam.PolicyStatement({ actions: ['sts:GetCallerIdentity'], resources: ['*'] }));
  role.addToPolicy(new iam.PolicyStatement({ actions: ['sts:AssumeRole'], resources: [`arn:${cdk.Aws.PARTITION}:iam::${cdk.Aws.ACCOUNT_ID}:role/cdk-hnb659fds-*-${cdk.Aws.ACCOUNT_ID}-${cdk.Aws.REGION}`] }));
  return role;
}
  • Replace lib/venue-status-stack.ts with this assembled stack:
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';
import { createNetwork } from './network';
import { createService } from './service';
import { createDeploymentRole } from './deployment-role';

export interface VenueStatusStackProps extends cdk.StackProps {
  readonly githubOwnerId: string;
  readonly githubRepositoryId: string;
}

// Assemble the network, service, identity, and reusable outputs.
export class VenueStatusStack extends cdk.Stack {
  constructor(scope: Construct, id: string, props: VenueStatusStackProps) {
    super(scope, id, props);
    const network = createNetwork(this);
    const service = createService(this, network.vpc, network.albSecurityGroup, network.taskSecurityGroup);
    const role = createDeploymentRole(this, props.githubOwnerId, props.githubRepositoryId);
    new cdk.CfnOutput(this, 'ServiceUrl', { value: `http://${service.venueService.loadBalancer.loadBalancerDnsName}` });
    new cdk.CfnOutput(this, 'ClusterName', { value: service.cluster.clusterName });
    new cdk.CfnOutput(this, 'ServiceName', { value: service.venueService.service.serviceName });
    new cdk.CfnOutput(this, 'LogGroupName', { value: service.logGroup.logGroupName });
    new cdk.CfnOutput(this, 'DeploymentRoleArn', { value: role.roleArn });
  }
}
  • Save the deployment role and assembled stack files.
  • Compile the complete infrastructure definition by running:
npm run build

A clean return confirms that all helper imports, resource references, and output names compile.

  • Synthesize the CloudFormation template by running:
npx aws-cdk synth

You should see the VenueStatusStack template. Confirm it contains one desired task, isolated application subnets, three interface endpoints, and an S3 gateway endpoint.

Build or synthesis failed?

Use the first TypeScript error to find the affected helper. An existing GitHub OIDC provider can also block synthesis or deployment because AWS allows one provider for the shared GitHub URL.

Help me troubleshoot the infrastructure definition.

Deploy and expose the capacity gap

Keep the lab short

This deployment starts billable resources, including a load balancer, a Fargate task, and interface endpoints. Keep the lab to this session and delete the stack afterward to stay near the estimated cost.

  • Bootstrap the authenticated AWS environment by running:
npx aws-cdk bootstrap

Bootstrapping can take several minutes while AWS creates the CDKToolkit support stack. You should see a success message for the target environment.

  • Deploy the application stack and save its outputs by running:
npx aws-cdk deploy --require-approval never --outputs-file outputs.json

The first deployment can take several minutes while CDK publishes the image and CloudFormation builds the network. You should see VenueStatusStack complete successfully.

Deployment failed?

Confirm Docker Desktop is running, the AWS session is current, and the Region is us-west-2. Check for an existing GitHub OIDC provider if identity creation fails.

Help me diagnose the deployment.

  • Load the stack outputs into PowerShell by running:
$outputs = Get-Content .\outputs.json | ConvertFrom-Json
$stack = $outputs.VenueStatusStack
$cluster = $stack.ClusterName
$service = $stack.ServiceName
$url = $stack.ServiceUrl
$stack

PowerShell should display ServiceUrl, ClusterName, ServiceName, LogGroupName, and DeploymentRoleArn.

  • Record the value shown for ServiceUrl: your ServiceUrl.
  • Open your browser through Windows Search.
  • Paste your ServiceUrl into the address bar.

You should see the live venue status page with Operational and Local container validated.

One task can serve traffic, but it cannot survive a desired capacity of zero. Before you run the check, do you expect the load balancer to keep returning the venue page?

  • Scale the service to zero and wait for it to stabilize by running:
aws ecs update-service --cluster $cluster --service $service --desired-count 0
aws ecs wait services-stable --cluster $cluster --services $service
  • Test the load-balancer URL by running:
try { Invoke-WebRequest -Uri $url -UseBasicParsing -TimeoutSec 10 } catch { $_.Exception.Message }

You should see an unsuccessful load-balancer response after the only target disappears. That shortfall is intentional: the architecture has no healthy capacity left.

You proved the single-task limit by making the live service unavailable. Next, you'll restore capacity with two private tasks and add the monitoring needed to see their health.

Add High Availability and Monitoring

The previous step proved the weakness of a single Amazon ECS task. Scaling the service to zero left the Application Load Balancer with no healthy target.

Now you'll use AWS Fargate to maintain two private tasks across Availability Zones. Amazon CloudWatch will turn their utilization and target health into visible operational evidence.

In this step, get ready to:
  • Restore the service with two private tasks.
  • Add logs, metrics, a dashboard, and an unhealthy-target alarm.
  • Deploy and verify the high-availability result.
Restore redundant capacity

A desired count of two gives the load balancer another healthy target when one task becomes unavailable. A one-week log retention period preserves enough evidence for this lab without keeping logs indefinitely.

  • Open lib/service.ts in Visual Studio Code.
  • Find these existing declarations:
const logGroup = new logs.LogGroup(scope, 'ApplicationLogGroup', { removalPolicy: cdk.RemovalPolicy.DESTROY });
securityGroups: [taskSecurityGroup], cpu: 256, memoryLimitMiB: 512, desiredCount: 1,
  • Replace the log-group declaration with the version below.
  • Replace desiredCount: 1 with the complete capacity line below.
const logGroup = new logs.LogGroup(scope, 'ApplicationLogGroup', { retention: logs.RetentionDays.ONE_WEEK, removalPolicy: cdk.RemovalPolicy.DESTROY });
securityGroups: [taskSecurityGroup], cpu: 256, memoryLimitMiB: 512, desiredCount: 2, minHealthyPercent: 100,
  • Save lib/service.ts.
  • Compile the capacity and retention edits by running:
npm run build

A clean return confirms the service now requests two tasks and retains application logs for one week.

Capacity edit not compiling?

Use the reported line number to check the replaced declarations. Confirm the capacity line ends with a comma inside the service properties.

Help me fix the capacity edit.

Add the monitoring signals

Metrics turn service activity into values that CloudWatch can graph. The unhealthy-target count also provides a signal when the load balancer loses a target.

  • Create lib/monitoring.ts.
  • Paste this monitoring helper into lib/monitoring.ts.
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
import * as patterns from 'aws-cdk-lib/aws-ecs-patterns';

// Graph service pressure and load-balancer target health together.
export function addMonitoring(scope: Construct, service: patterns.ApplicationLoadBalancedFargateService) {
  const cpu = service.service.metricCpuUtilization({ period: cdk.Duration.minutes(1) });
  const memory = service.service.metricMemoryUtilization({ period: cdk.Duration.minutes(1) });
  const healthy = service.targetGroup.metrics.healthyHostCount({ period: cdk.Duration.minutes(1) });
  const unhealthy = service.targetGroup.metrics.unhealthyHostCount({ period: cdk.Duration.minutes(1), statistic: 'Minimum' });

  // Require two consecutive unhealthy data points before alarming.
  const alarm = unhealthy.createAlarm(scope, 'UnhealthyTargetAlarm', {
    threshold: 1,
    evaluationPeriods: 2,
    datapointsToAlarm: 2,
  });
  const dashboard = new cloudwatch.Dashboard(scope, 'VenueDashboard');
  dashboard.addWidgets(
    new cloudwatch.GraphWidget({ title: 'ECS CPU and memory utilization', left: [cpu, memory], width: 12 }),
    new cloudwatch.GraphWidget({ title: 'Application target health', left: [healthy, unhealthy], width: 12 }),
  );
  return { dashboard, alarm };
}

What does the monitoring helper measure?

  • The CPU and memory metrics show Fargate service utilization.
  • The healthy and unhealthy metrics show load-balancer target counts.
  • The alarm requires two consecutive periods with at least one unhealthy target.
  • Replace lib/venue-status-stack.ts with this monitored stack:
import * as cdk from 'aws-cdk-lib';
import { Construct } from 'constructs';
import { createNetwork } from './network';
import { createService } from './service';
import { createDeploymentRole } from './deployment-role';
import { addMonitoring } from './monitoring';

export interface VenueStatusStackProps extends cdk.StackProps {
  readonly githubOwnerId: string;
  readonly githubRepositoryId: string;
}

// Assemble the service and expose every operational resource name.
export class VenueStatusStack extends cdk.Stack {
  constructor(scope: Construct, id: string, props: VenueStatusStackProps) {
    super(scope, id, props);
    const network = createNetwork(this);
    const service = createService(this, network.vpc, network.albSecurityGroup, network.taskSecurityGroup);
    const monitoring = addMonitoring(this, service.venueService);
    const role = createDeploymentRole(this, props.githubOwnerId, props.githubRepositoryId);
    new cdk.CfnOutput(this, 'ServiceUrl', { value: `http://${service.venueService.loadBalancer.loadBalancerDnsName}` });
    new cdk.CfnOutput(this, 'ClusterName', { value: service.cluster.clusterName });
    new cdk.CfnOutput(this, 'ServiceName', { value: service.venueService.service.serviceName });
    new cdk.CfnOutput(this, 'LogGroupName', { value: service.logGroup.logGroupName });
    new cdk.CfnOutput(this, 'DashboardName', { value: monitoring.dashboard.dashboardName });
    new cdk.CfnOutput(this, 'AlarmName', { value: monitoring.alarm.alarmName });
    new cdk.CfnOutput(this, 'DeploymentRoleArn', { value: role.roleArn });
  }
}
  • Save both monitoring files.
  • Compile the monitored stack by running:
npm run build

A clean return confirms the service metrics feed the dashboard, alarm, and output declarations.

  • Synthesize the updated template by running:
npx aws-cdk synth

You should see the updated template with a desired count of two plus dashboard and alarm resources.

Monitoring definition failing?

Confirm monitoring.ts exports addMonitoring and the stack imports it from ./monitoring.

Help me fix the monitoring definition.

Deploy and verify high availability

This deployment resumes paid Fargate capacity. Keep the service only as long as you need it, then use the cleanup section to stop ongoing charges.

Expect the update to take several minutes while CloudFormation launches two tasks and the load balancer checks their health. A quiet terminal during stabilization is normal.

  • Deploy the updated stack and refresh outputs.json by running:
npx aws-cdk deploy --require-approval never --outputs-file outputs.json

You should see VenueStatusStack complete successfully with the service restored.

  • Reload the current stack outputs by running:
$outputs = Get-Content .\outputs.json | ConvertFrom-Json
$stack = $outputs.VenueStatusStack
$cluster = $stack.ClusterName
$service = $stack.ServiceName
$url = $stack.ServiceUrl
$stack

PowerShell should now include DashboardName and AlarmName alongside the earlier outputs.

Before you list the tasks, how many running task ARNs do you expect after the high-availability deployment?

  • List the running service tasks by running:
aws ecs list-tasks --cluster $cluster --service-name $service --desired-status RUNNING

You should see two task ARNs. The service has restored its desired capacity across the private subnets.

  • Open your ServiceUrl in your browser.
  • Confirm the venue page loads with the Operational badge.
  • Open the CloudWatch console.
  • Select Dashboards from the left navigation.
  • Open the dashboard named by DashboardName in outputs.json.

You should see CPU and memory utilization beside healthy and unhealthy target counts. The healthy-target graph should show two targets after metrics arrive.

High-availability result missing?

Refresh the dashboard after several minutes if the first graph is empty. Confirm the deployment completed and the task list contains two running tasks.

Help me verify the high-availability deployment.

Your service now maintains two private tasks and exposes their health through CloudWatch. Next, you'll move deployment off your workstation and into a repository-scoped GitHub Actions workflow.

Automate Deployment with GitHub Actions

Your AWS service now stays available across two healthy Fargate tasks and exposes its health through CloudWatch. A deployment that depends on your workstation still does not prove repeatable delivery.

This step moves deployment into GitHub Actions as a CI/CD path. OIDC gives each workflow run temporary AWS credentials scoped to your repository and main branch.

In this step, get ready to:
  • Create the repository-scoped deployment workflow.
  • Commit a visible application change.
  • Verify the automated deployment in GitHub and on the live page.
Connect GitHub Actions to AWS

The workflow assumes the role exposed by DeploymentRoleArn. GitHub requests a short-lived identity token for each job, so the repository never stores a long-lived AWS access key.

  • Open outputs.json in Visual Studio Code.
  • Locate the DeploymentRoleArn value under VenueStatusStack.
  • Record the value here: your deployment role ARN.
  • Create the workflow folders by running:
mkdir .github\workflows
  • Create .github/workflows/deploy.yml in Visual Studio Code.
  • Paste this workflow into .github/workflows/deploy.yml.
name: Deploy venue status service
on:
  push:
    branches: [main]
permissions:
  contents: read
  id-token: write
jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      # Check out the exact commit without preserving Git credentials.
      - uses: actions/checkout@v7.0.1
        with:
          persist-credentials: false
      # Match the local Node.js runtime used by the project.
      - uses: actions/setup-node@v7.1.0
        with:
          node-version: '24.21.0'
          package-manager-cache: false
      - run: npm install
      - run: npm run build
      # Exchange the GitHub OIDC token for temporary AWS credentials.
      - uses: aws-actions/configure-aws-credentials@v6.3.0
        with:
          role-to-assume: [[DEPLOYMENT_ROLE_ARN="your deployment role ARN"]]
          role-session-name: VenueStatusDeploy
          aws-region: us-west-2
      - run: aws sts get-caller-identity
      - run: npx aws-cdk synth
      - run: npx aws-cdk deploy --require-approval never --outputs-file outputs.json

How does the workflow deploy securely?

  • The permissions block allows repository checkout and OIDC token creation.
  • The credential action exchanges the GitHub token for temporary AWS role credentials.
  • The final steps build, synthesize, and deploy the same CDK application you tested locally.
  • Save .github/workflows/deploy.yml.
  • Search the saved file for your deployment role ARN to confirm the rendered value is present.
  • Create .gitignore beside package.json.
  • Paste these generated paths into .gitignore.
node_modules/
dist/
cdk.out/
outputs.json
  • Save .gitignore.
  • Confirm the workflow and ignore file exist by running:
Get-ChildItem .github\workflows, .gitignore

PowerShell should list deploy.yml and .gitignore.

Workflow file incomplete?

Confirm the workflow is inside .github/workflows and the ARN came from the current outputs.json file.

Help me check the workflow setup.

Commit a visible deployment change

A visible banner connects the new container image to the automated deployment. The live page becomes the final proof that GitHub Actions updated the running service.

  • Open app/server.js.
  • Find the page template containing this text:
<p>Local container validated</p>
  • Replace the existing paragraph with this deployment marker:
<p>Deployment verified by GitHub Actions</p>
  • Save app/server.js.
  • Search the saved file for Deployment verified by GitHub Actions.

You should find the new marker once and no remaining Local container validated text.

  • Enter the public repository URL here: https://github.com/your-owner/secure-venue-status-service.git.

Before you push, do you expect AWS authentication to use a stored access key or a temporary OIDC credential?

  • Initialize Git, commit the project, and push the main branch by running:
git init
git add .
git commit -m "Deploy secure venue status service"
git branch -M main
git remote add origin [[YOUR_REPOSITORY_URL="https://github.com/your-owner/secure-venue-status-service.git"]]
git push -u origin main

PowerShell should report that main now tracks the remote branch. The push starts the deployment workflow.

Commit did not reach GitHub?

Confirm the URL belongs to the empty public repository from Step 1. Configure the Git author name and email requested by Git if the commit stops before the push.

Help me diagnose the Git push.

Inspect the automated deployment

The workflow logs prove deployment moved off your workstation. The run can remain on its deployment step for several minutes while Docker publishes the image and AWS replaces both tasks.

  • Open the main page of your GitHub repository.
  • Click Actions under the repository name.
  • Select Deploy venue status service from the workflow list.
  • Open the run triggered by your latest push.
  • Expand the deploy job.

You should see successful checkout, build, identity, synthesis, and deployment steps. The identity step proves the workflow received temporary AWS credentials.

Workflow run failed?

Open the first failed workflow step and read its final lines. An OIDC failure points to the role ARN or repository trust, while a TypeScript failure points to the committed source.

Help me troubleshoot the workflow run.

  • Open your ServiceUrl after the workflow turns green.
  • Refresh the browser until the new task deployment finishes.

You should see Deployment verified by GitHub Actions beneath the venue status heading. That is the automated delivery path working end to end.

Your repository now builds and deploys the monitored service with temporary AWS credentials. The live banner gives you visible proof that the workflow updated both private tasks.

Secret mission

Prove the Service Survives a Task Failure

A diagram can claim high availability. This mission tests that claim by stopping one task while thirty requests hit the live service. You will collect evidence that traffic remains available while capacity recovers.

Clean Up Your Resources

Clean Up Your Resources

Your AWS deployment includes paid resources, so your cleanup choice affects ongoing charges. Decide whether to keep the service running, pause its compute capacity, or delete the application stack.

Cost warning

The Application Load Balancer, AWS Fargate tasks, interface VPC endpoints, Amazon ECR storage, and Amazon CloudWatch usage continue generating charges while retained.

The lab is estimated to cost under $1 only when it stays short and low traffic. Destroy the application stack in the same session to keep costs near that estimate.

Deleting VenueStatusStack leaves the separate CDKToolkit bootstrap stack in place. That stack can retain a small asset bucket and ECR repository.

Resources you used:

  • Application stack: VenueStatusStack, including the VPC, subnets, endpoints, load balancer, ECS service, task definition, security groups, log group, dashboard, alarm, OIDC provider, and deployment role.
  • Bootstrap support: CDKToolkit, including its S3 asset bucket, ECR repository, and deployment roles.
  • Local assets: the venue-status:local Docker image and ~/Desktop/secure-venue-status-service project folder.
  • Repository: the public GitHub repository named secure-venue-status-service.

Keep everything running

No action is needed. Choose this if you are still testing the live service or plan to demonstrate it soon.

  • Keep VenueStatusStack deployed with two running tasks.
  • Monitor AWS charges while the load balancer, interface endpoints, tasks, and stored assets remain provisioned.
  • Keep the local image, project folder, and GitHub repository for continued development.
  • Retain CDKToolkit if this AWS account and Region will host more CDK projects.

Pause - I'll come back to this later

Shut down the Fargate tasks to reduce compute charges while keeping the network, load balancer, endpoints, files, and repository available. The retained AWS resources continue generating some charges.

  • Load the current cluster and service names from outputs.json by running:
$outputs = Get-Content .\outputs.json | ConvertFrom-Json
$stack = $outputs.VenueStatusStack
$cluster = $stack.ClusterName
$service = $stack.ServiceName
  • Scale the service to zero by running:
aws ecs update-service --cluster $cluster --service $service --desired-count 0
  • Wait for the service update to stabilize by running:
aws ecs wait services-stable --cluster $cluster --services $service
  • Confirm that no tasks remain by running:
aws ecs list-tasks --cluster $cluster --service-name $service --desired-status RUNNING

You should see an empty task list. The load balancer and interface endpoints remain provisioned, so return to the Delete tab if you want to stop their charges.

Service did not pause?

Confirm the PowerShell session is authenticated to the deployment account in us-west-2 and that outputs.json belongs to the current stack.

Help me pause the ECS service.

Delete - I don't want to use this again

Remove all project resources and start fresh. This path deletes the application stack, optional bootstrap support, local image, local project folder, and GitHub repository.

Check each target before deletion

These actions are destructive. Confirm that VenueStatusStack is the application stack and save any evidence you want to keep before continuing.

  • Delete the application stack from ~/Desktop/secure-venue-status-service by running:
npx aws-cdk destroy VenueStatusStack
  • Approve the prompt only when it names VenueStatusStack.
  • Open the CloudFormation console in us-west-2.
  • Confirm VenueStatusStack no longer appears among active stacks.

The application URL stops loading after the load balancer finishes deleting. This closes the main application cost path.

Does another CDK project use this environment?

Keep CDKToolkit if any other CDK project uses this AWS account and Region. AWS warns that deleting the bootstrap stack removes the resources required by CDK deployments.

Record the bootstrap resources
  • Skip the bootstrap deletion steps if another CDK project uses CDKToolkit.
  • Open CDKToolkit in the CloudFormation console if this environment will no longer host CDK projects.
  • Select the Resources tab.
  • Record the bootstrap S3 bucket name here: your bootstrap S3 bucket name.
  • Record the bootstrap ECR repository name here: your bootstrap ECR repository name.
Empty the bootstrap assets
  • Open the S3 console.
  • Select your bootstrap S3 bucket name.
  • Click Empty.
  • Complete the confirmation requested by the S3 console.
  • Open the Amazon ECR console.
  • Open your bootstrap ECR repository name.
  • Select every remaining image.
  • Click Delete.
Delete the bootstrap stack
  • Return to the CloudFormation stack list.
  • Select CDKToolkit.
  • Click Delete.
  • Confirm the bootstrap stack no longer appears among active stacks.
  • Remove the local Docker image by running:
docker image rm venue-status:local

Docker should report that the venue-status:local tag and image were removed from the workstation.

  • Open the public GitHub repository.
  • Click Settings under the repository name.
  • Scroll to the Danger Zone on the General page.
  • Click Delete this repository.
  • Click I want to delete this repository.
  • Click I have read and understand these effects.
  • Enter secure-venue-status-service in the repository confirmation field.
  • Click Delete this repository.
  • Move to your Desktop by running:
Set-Location ~/Desktop
  • Delete the local project folder by running:
Remove-Item -Recurse -Force .\secure-venue-status-service
  • Confirm the project folder is gone by running:
Get-ChildItem

You should no longer see secure-venue-status-service on your Desktop. Every resource listed above now has a completed teardown path.

AWS cleanup did not finish?

For a failed application deletion, confirm the current directory still contains cdk.json and the AWS session targets us-west-2. A bootstrap deletion can fail until its S3 bucket and ECR repository are empty.

Help me finish the AWS cleanup.

Nice Work!

Nice Work!

You did it! Your venue status service now runs on two private AWS Fargate tasks across separate Availability Zones behind an internet-facing Application Load Balancer.

You've learned how to:

  • Build a non-root Linux container with Docker. Prove runtime health with a Bash diagnostic. Trace requests through structured logs.
  • Define a monitored multi-AZ service in AWS CDK using TypeScript. Keep its Amazon ECS tasks private behind the load balancer. Surface service health in Amazon CloudWatch.
  • Automate repeatable deployments through GitHub Actions using OIDC temporary credentials. Deploy from the main branch without storing long-lived AWS access keys.
  • Secret Mission: Stop one running task during a controlled resilience test. Prove continuous traffic through the surviving task while Amazon ECS restores the missing capacity.

Ready to quiz yourself?