Recover an Azure Container App Outage

Diagnose and recover an Azure Container Apps ingress outage.

Introduction

30 Second Summary

When a public service goes down, the error page tells users almost nothing about the configuration that failed. Restoring it quickly depends on finding evidence before changing anything.

In this project, you will complete an Azure Container Apps incident response drill. Your final evidence connects a deliberate ingress outage to its target-port fix in incident-report.md.

What You'll Build

You refresh a public URL in your browser to see the Microsoft quickstart page return after your evidence identifies the broken port.

By the end of this project, you'll have:

  • A healthy public baseline you can open in your browser to see the Microsoft quickstart page.
  • A reproducible outage you can trigger by switching the target port from 80 to 81. The unavailable endpoint makes the routing failure visible.
  • An evidence-backed recovery where system logs reveal platform behavior. Live ingress configuration exposes the mismatch. Your incident-report.md preserves the diagnosis plus remediation.
  • Secret Mission: Deploy an intentionally unreachable image to prove that single revision mode keeps the healthy revision serving traffic.

Are there any prerequisites?

You'll need an active Azure subscription with permission to create Azure resource groups plus Azure Container Apps resources. Confirm that the subscription's monthly Container Apps grants remain available before you begin.

Before We Start

This is your moment to commit to a timed Azure Container Apps support incident before the hands-on work begins. You will treat the outage as an operations exercise that produces current troubleshooting evidence.

Prepare Your Azure Operations Console

A production-style Azure Container Apps incident becomes harder to investigate when your administration tools point at the wrong subscription. The lab also needs a clean command environment before you touch the service.

The official Azure CLI image gives you an isolated console inside Docker Desktop. This keeps the host-level setup on your computer unchanged.

Your goal is an authenticated console for the intended Azure subscription. The Microsoft.App provider also needs to report Registered before you create any lab resources.

In this step, get ready to:
  • Confirm the intended subscription in the Azure portal.
  • Authenticate the containerized Azure CLI.
  • Register Microsoft.App for Azure Container Apps.
Confirm portal access and launch Azure CLI

The Azure portal gives you a visual check of the subscription you intend to use. Keep the portal open in your preferred browser during the lab.

The browser handles your Azure account sign-in. Your password never needs to be typed into the CLI container.

  • Press Cmd+Space (macOS) or the Windows key (Windows) to open system search.
  • Type the name of your preferred browser, then press Enter to open it.
  • Navigate to https://portal.azure.com.
  • Sign in with the Azure account for this lab.
  • Confirm the intended subscription is visible in the portal.
  • Keep the portal tab available for later verification.

Good. Your browser now confirms the account context for the incident lab.

The temporary CLI runs in Terminal. Docker Desktop provides the isolated environment that contains it.

  • Press Cmd+Space (macOS) or the Windows key (Windows) to open system search.
  • Type Docker Desktop and press Enter to open it.
  • Wait until Docker Desktop reports that it is ready.
  • Open system search again.
  • Type Terminal (macOS) or Windows Terminal (Windows), then press Enter to open it.

The first launch can take a few minutes while Docker downloads the image. A quiet terminal during that download is normal.

  • Launch the temporary Azure CLI container by running this command:
docker run --rm -it mcr.microsoft.com/azure-cli:azurelinux3.0

What does this command do?

  • The --rm option removes the temporary container when it exits.
  • The -it options provide the interactive terminal used for Azure commands.
  • The azurelinux3.0 tag selects Microsoft's maintained Azure Linux 3.0 image channel.
  • Verify Azure CLI inside the container by running this command at its prompt:
az version

What does this check prove?

The command prints the installed Azure CLI version information. It also prints any installed extension information.

Seeing that output proves the az command is available inside the isolated image.

You should see Azure CLI information printed in your terminal. The container prompt remains available for your next command.

CLI information not appearing?

Confirm Docker Desktop reached its ready state before you launched the container. If the container exited, run the Docker command above again.

If the image download is still active, let it finish before entering an Azure command.

Get targeted help to troubleshoot the Azure CLI container.

Authenticate to the intended subscription

A device-code flow links the browser session to the CLI without opening a browser inside the container. The CLI prints the sign-in details you need.

  • Begin the device-code sign-in from the container prompt by running:
az login --use-device-code

How does device login work?

The --use-device-code option sends the interactive sign-in to your browser. This works when the CLI environment cannot open its own browser window.

The command pauses while you approve access. Terminal continues automatically after the browser flow succeeds.

  • Copy the device code printed in your terminal.
  • Return to your browser.
  • Open the sign-in address printed by Azure CLI.
  • Enter the device code on the sign-in page.
  • Complete the Azure authorization prompts.
  • Return to your terminal after the browser confirms the sign-in.
  • Inspect the active Azure subscription by running:
az account show --output table

What does this account check show?

The command returns details for the default Azure subscription. The table format makes the active subscription easier to identify.

You should see one subscription in the table. Compare it with the subscription you confirmed in the Azure portal.

✔️ I see the intended subscription

Your browser session matches the CLI account context. You can continue with provider registration.

ⓧ I see a different subscription

Azure CLI can access more than one enabled subscription. You need to select the subscription reserved for this lab.

  • List the enabled subscriptions by running:
az account list --output table

What does this list show?

The table shows the enabled subscriptions available to your signed-in account. Use it to identify the required subscription name or ID.

  • Select the required subscription by replacing SUBSCRIPTION_NAME_OR_ID in this command:
az account set --subscription "SUBSCRIPTION_NAME_OR_ID"

What does this command change?

The command changes the active Azure CLI subscription. Future Azure commands target that subscription until you change it again.

  • Confirm the subscription change by running:
az account show --output table

What should the table confirm?

The table should now show the subscription you selected. Its name should match the intended subscription in the Azure portal.

Device login not completing?

Confirm your browser is signed in with the same Azure account you intend to use for the lab. Start the device-code flow again if the earlier code expired.

If the expected subscription is missing, confirm that your Azure account has access to it.

Get help to resolve the Azure device login or subscription mismatch.

Register Microsoft.App and verify readiness

A resource provider connects an Azure subscription to a service namespace. Azure Container Apps relies on the Microsoft.App namespace.

Provider registration can take a few minutes. The command waits for Azure to finish before it returns.

  • Register the Azure Container Apps resource provider by running:
az provider register --namespace Microsoft.App --wait

What does provider registration do?

The command registers the Microsoft.App namespace with the active subscription. The --wait option keeps the command running until Azure finishes processing the request.

This prepares the subscription for Azure Container Apps resources in the next step.

Before you run the final check, which subscription should appear in the table? What provider state should the second command print?

  • Verify the active subscription and provider state by running:
az account show --output table
az provider show --namespace Microsoft.App --query registrationState --output tsv

What does this final check prove?

  • The first command prints the active subscription as a table.
  • The second command filters the provider response to its registration state.
  • The tsv output format prints that state without extra JSON structure.

You should see the intended subscription in the table. The final line should print Registered.

That's your operations console ready. The official Azure CLI container is authenticated to the same subscription you can see in the Azure portal.

Provider not registered?

If the provider state differs from Registered, confirm the registration command completed before running the final check again.

If the account table shows the wrong subscription, use the subscription selection path above before repeating registration.

Get help to diagnose Microsoft.App provider registration.

Operations Console Checkpoint

Your interactive official Azure Linux 3.0 Azure CLI Docker container is running. It is authenticated to the intended subscription.

The Microsoft.App provider registration state is Registered. No Azure lab resources exist yet.

Safari is signed in at https://portal.azure.com. The intended Azure subscription is visible.

Your Azure operations console is ready. Next up, you will put a healthy public container endpoint online.

Deploy a Healthy Container App

Your isolated Azure administration console is ready. The incident drill now needs a healthy service baseline.

Azure Container Apps gives the lab a public endpoint without cluster administration. Once the quickstart page loads, every later test has a known-good comparison.

In this step, get ready to:
  • Create the dedicated Azure resource group.
  • Provision a Container Apps environment without persistent log storage.
  • Deploy the quickstart workload as a public scale-to-zero endpoint.
Create the lab resource group

A resource group gives the lab one boundary for creation and cleanup. Shell variables keep the resource names consistent across every command.

  • Define the lab names and create their resource group by running:
export RESOURCE_GROUP=aca-support-lab
export LOCATION=centralus
export ENVIRONMENT=aca-support-env
export APP_NAME=aca-support-app

az group create --location "$LOCATION" --name "$RESOURCE_GROUP"

What does this command block do?

  • Each export stores a reusable lab value in the current CLI session.
  • The final command creates the dedicated resource group in centralus.
  • Keeping the lab in one resource group makes the final cleanup a single operation.
  • Review the resource group details returned by the command.

You should see aca-support-lab in centralus. Good progress. Your incident lab now has a dedicated Azure boundary.

Resource group not created?

  • Confirm the CLI container from the previous step still shows an active prompt.
  • Return to the subscription check from the previous step if Azure reports an authentication problem.
  • Confirm that your account can create resources in the intended subscription.

Help me troubleshoot the resource group creation.

Provision the Container Apps environment

The environment provides the managed boundary where your container app runs. Selecting no log destination avoids creating a persistent Log Analytics workspace.

Real-time logs remain available for the investigation. This keeps the evidence you need without adding persistent log storage.

A few minutes of provisioning time is normal here.

  • Create the environment with log storage disabled by running:
az containerapp env create \
  --name "$ENVIRONMENT" \
  --resource-group "$RESOURCE_GROUP" \
  --location "$LOCATION" \
  --logs-destination none

How is the environment configured?

  • The environment is named aca-support-env.
  • The environment belongs to aca-support-lab.
  • The --logs-destination none setting disables persistent log storage.
  • Real-time system logs remain available for the later diagnosis.
  • Wait for the CLI prompt to return.

You should see details for aca-support-env after provisioning completes. That environment can now host the incident workload without a persistent log destination.

Environment provisioning failed?

  • Confirm that Docker Desktop remains running on your computer.
  • Check that the CLI container remains authenticated to the intended subscription.
  • Confirm that the Microsoft.App provider registration from the previous step remains complete.

Help me debug the environment deployment.

Deploy the public quickstart app

External ingress gives the app a public HTTPS address. The target port must match the port where the image listens.

A minimum of 0 replicas lets the app scale down when idle. A maximum of 1 replica keeps this support lab small.

Provisioning the first app can take a few minutes.

  • Deploy the quickstart image and capture its public address by running:
export FQDN=$(az containerapp create \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --environment "$ENVIRONMENT" \
  --image mcr.microsoft.com/k8se/quickstart:latest \
  --ingress external \
  --target-port 80 \
  --min-replicas 0 \
  --max-replicas 1 \
  --query properties.configuration.ingress.fqdn \
  --output tsv)

echo "https://$FQDN"

What does this deployment create?

  • The app uses Microsoft's mcr.microsoft.com/k8se/quickstart:latest image.
  • External ingress routes public traffic to port 80.
  • The ingress uses auto transport selection.
  • The replica range permits the workload to scale between 0 and 1.
  • The FQDN variable stores the public hostname returned by Azure.
  • The final line prints the complete HTTPS address for the browser check.

Before you load the URL, what do you expect the official quickstart image to show?

  • Switch back to your browser.
  • Create a new browser tab.
  • Paste the printed HTTPS URL into the address bar.
  • Press Enter.

The first request can take a short moment while the app scales from zero. You should then see the Microsoft quickstart page.

Quickstart page not loading?

  • Confirm that the address starts with https://.
  • Wait for the app to finish starting.
  • Refresh the browser tab.
  • Check the deployment command in your terminal for a provisioning failure.

Help me get the public endpoint working.

✔️ Awesome, I've got everything!

Your authenticated CLI session now contains every lab variable. Your browser tab proves that the public endpoint is healthy.

ⓧ I'd like to double check the full code

  • Compare the cumulative commands below with the commands you ran:
export RESOURCE_GROUP=aca-support-lab
export LOCATION=centralus
export ENVIRONMENT=aca-support-env
export APP_NAME=aca-support-app

az group create --location "$LOCATION" --name "$RESOURCE_GROUP"

az containerapp env create \
  --name "$ENVIRONMENT" \
  --resource-group "$RESOURCE_GROUP" \
  --location "$LOCATION" \
  --logs-destination none

export FQDN=$(az containerapp create \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --environment "$ENVIRONMENT" \
  --image mcr.microsoft.com/k8se/quickstart:latest \
  --ingress external \
  --target-port 80 \
  --min-replicas 0 \
  --max-replicas 1 \
  --query properties.configuration.ingress.fqdn \
  --output tsv)

echo "https://$FQDN"

What should this session contain?

  • The first four exports define the names reused by every Azure command.
  • The resource commands create the group, environment, and public app.
  • The FQDN variable stores the app's public hostname.
  • The final line prints the HTTPS address that loads in your browser.

Your healthy baseline is live. Next, you will introduce a controlled ingress fault and observe how the endpoint behaves during an outage.

Trigger an Ingress Outage

The quickstart page in your browser proved that your public Azure Container Apps endpoint could reach the container on port 80. You now have a known-good baseline for the timed incident.

A healthy deployment hides how a small ingress error appears to users. This step introduces one live configuration fault so you can experience the outage before examining its evidence.

In this step, get ready to:
  • Capture the healthy ingress configuration as a baseline.
  • Change the live target port from 80 to 81.
  • Request the public endpoint until the outage is visible.
Capture the healthy ingress baseline

The Azure CLI can read the live ingress object before you change it. This gives you a before-state that separates the known-good route from the incident state.

  • Read the live ingress object for aca-support-app by running this command:
az containerapp ingress show \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP"

What does this command show?

  • The command reads the current ingress configuration without changing the application.
  • The targetPort property identifies the container port that receives incoming traffic.
  • The output becomes your known-good configuration baseline.

In the output, you'll see targetPort set to 80. Good call capturing this now. Your healthy route is recorded before the incident begins.

Can't read the ingress configuration?

  • Return to the authenticated Azure CLI container from earlier.
  • Confirm that the active subscription contains aca-support-app in aca-support-lab.

Help me inspect the existing ingress configuration.

Inject the target-port fault

A target port directs incoming traffic to the port used by the containerized application. This fault changes only that application-scoped setting.

  • Change the live ingress target port to 81 by running this command:
az containerapp ingress update \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --target-port 81

What does this update change?

  • The update changes targetPort from 80 to 81.
  • The container image remains mcr.microsoft.com/k8se/quickstart:latest.
  • The replica range remains 0 to 1.
  • Ingress settings apply across the application without creating a new revision.

The CLI output should show targetPort as 81. The new route setting is now live.

Ingress update did not complete?

  • Confirm that the Azure CLI container remains authenticated.
  • Verify that APP_NAME still represents aca-support-app.
  • Verify that RESOURCE_GROUP still represents aca-support-lab.

Help me apply the ingress update.

Observe the public outage

Browser requests reveal how users experience the altered route. Repeated refreshes also send traffic to the app after scale to zero.

Before you refresh, predict whether the quickstart page can still load through the changed route.

  • Switch back to the public application tab in your browser.
  • Refresh the page several times.

You'll see an HTTP error or an unavailable response instead of the Microsoft quickstart page. This failed response is intentional. You now have a reproducible ingress outage.

Still seeing the quickstart page?

  • Confirm that the ingress update output shows targetPort set to 81.
  • Refresh the same public application tab that displayed the healthy quickstart page earlier.

Help me verify the outage injection.

  • Leave the ingress target port at 81 for the investigation.
  • Keep the failed browser tab available for the investigation.

The outage is now reproducible while the faulty configuration remains intact. Next, you'll investigate the failure without destroying the evidence.

Diagnose the Port Mismatch

Your controlled ingress change has turned a healthy Azure Container Apps endpoint into an outage. The failed URL now gives you a realistic incident to investigate while the evidence is fresh.

Azure CLI gives you current platform events. The live configuration gives you the decisive routing value.

The Azure portal adds a purpose-built diagnostic check.

Guessing or redeploying would erase useful evidence. The incorrect target port stays in place throughout this step.

In this step, get ready to:
  • Review real-time system logs without creating persistent log storage.
  • Compare the live ingress target port with the image's listening port.
  • Run the Azure portal ingress detector to confirm the fault.
Review the real-time system logs

The environment has no persistent Log Analytics destination. Real-time system logs still expose current platform activity without changing the lab's logging configuration.

  • Switch back to the authenticated Azure CLI container from earlier.
  • Inspect the latest 50 platform events by running:
az containerapp logs show \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --type system \
  --tail 50

What does this command show?

  • The --type system option requests platform events for the container app.
  • The --tail 50 option limits the result to the latest 50 entries.
  • Real-time logs remain available when the logs destination is set to none.
  • Review the output for events related to the recent failed requests.

The output may contain scaling or revision events. System events provide supporting evidence for your investigation.

Seeing no recent system events?

The app can scale to zero, so the event stream may be sparse between requests.

  • Return to the broken application tab in your browser.
  • Refresh the unavailable endpoint once.
  • Rerun the system log command from the Azure CLI container.

help me investigate an empty Azure Container Apps system log stream

Compare the live ingress port

External ingress forwards each request to the configured targetPort. That value must match the port where the containerized application listens.

Before you inspect it, predict whether the live configuration still points to port 80 or port 81.

  • Inspect the live ingress object by running:
az containerapp ingress show \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --output json

What does this command show?

  • The --output json option returns the ingress object as JSON.
  • The response reflects the app's current ingress configuration.
  • Ingress settings apply across the app's revisions. This inspection leaves the incident state unchanged.
  • Locate targetPort in the JSON output.
  • Compare its value with port 80 used by the quickstart image.

You'll see targetPort set to 81. The image mcr.microsoft.com/k8se/quickstart:latest listens on port 80.

You have found the decisive mismatch. Ingress sends traffic to a port with no application listener.

Cannot find the target port?

  • Scroll through the ingress output until you find targetPort.
  • Confirm the returned app name is aca-support-app.
  • Confirm your Azure CLI session still uses the intended subscription.

help me locate the target port in my Azure Container Apps ingress JSON

Run the portal ingress detector

The Ingress Port Settings Check evaluates the app's ingress port configuration. Its findings are most useful while the app is starting or scaling.

This diagnostic path is easy to miss.

  • Switch back to the Azure portal tab in your browser.
  • Use the portal search at the top of the page to find aca-support-app.
  • Select aca-support-app from the search results.
  • Select Diagnose and solve problems from the container app resource menu.
  • Select Availability and Performance.
  • Choose Ingress Port Settings Check.
  • Run the Ingress Port Settings Check.

The detector may immediately highlight the ingress port configuration. A recent failed request gives it a stronger signal to evaluate.

Before the final detector run, predict whether its evidence will point toward the container image or the ingress route.

  • Switch to the public application tab in your browser.
  • Refresh the unavailable endpoint to generate a current request.

You'll see the same HTTP error or unavailable response. The failed request preserves the incident while generating fresh diagnostic activity.

  • Return to the diagnostics tab in the Azure portal.
  • Run the Ingress Port Settings Check again.

The detector identifies the ingress port configuration or directs you toward that area. Your CLI evidence confirms that ingress routes to 81 while the image listens on 80.

  • Leave the target port set to 81 for the recovery step.

The public URL remains unavailable. You have preserved the faulty state with evidence from system logs, live configuration, and the portal detector.

Detector has no recent finding?

The detector raises findings while the app is trying to start or scale.

  • Refresh the broken public endpoint again.
  • Wait until the failed request finishes.
  • Run the Ingress Port Settings Check again.

help me run the Azure ingress port detector during the failed request

That is a solid incident diagnosis. Your evidence now points to the ingress target-port mismatch.

Next, you'll restore port 80. You'll turn this evidence into an incident report.

Recover Service and Document the Incident

Your evidence has isolated the outage to one configuration mismatch in Azure Container Apps. External ingress sends requests to port 81 even though the quickstart container listens on port 80.

Recovery turns that evidence into a verified fix. You will restore target port 80 with the Azure CLI. You will verify the public endpoint before documenting the response in Visual Studio Code.

In this step, get ready to:
  • Restore the ingress target port to 80.
  • Verify the repaired configuration and public endpoint.
  • Create an incident report from the evidence you collected.
Restore the healthy target port

The confirmed recovery action is to align the ingress target port with the container's listening port. This removes the routing mismatch without replacing the image or environment.

  • Apply the confirmed recovery change by running this command:
az containerapp ingress update \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --target-port 80

What Does This Recovery Change?

  • The ingress update changes targetPort from 81 to 80.
  • The app keeps the same image.
  • The app keeps its external endpoint.
  • Requests can now reach the port where the quickstart container listens.
  • Confirm the command returns the updated app details without reporting a failure.

You have cleared the configuration fault. The ingress now points at the port the container actually serves.

Recovery Command Failed?

Confirm that the authenticated Azure CLI container from earlier is still running. Check that APP_NAME remains set to aca-support-app.

Check that RESOURCE_GROUP remains set to aca-support-lab. A new terminal session does not inherit variables from the existing container.

Ask for help with the exact output you received: help me troubleshoot the ingress recovery command

Verify the repaired service

A reliable recovery check uses two signals. The live configuration proves the change persisted. The public page proves requests reach the container again.

Before you inspect the app, which target port do you expect Azure to report now?

  • Inspect the live ingress configuration by running this command:
az containerapp ingress show \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --output json

What Does This Verification Prove?

The command reads the current ingress object from Azure. A targetPort value of 80 proves the recovery setting persisted.

The browser check tests the complete request path. Together, these checks connect the configuration change to the user-visible recovery.

You should see targetPort set to 80 in the live ingress JSON.

Before you refresh the endpoint, do you expect the quickstart page to return immediately or after the app wakes from zero replicas?

  • Switch back to the browser tab for the public https:// URL from earlier.
  • Refresh the page until it responds.

The first request may take a moment while the app scales from zero. You should then see the Microsoft quickstart page load successfully.

That is the service restored. Your public endpoint is healthy again after the confirmed port correction.

Quickstart Page Still Unavailable?

Refresh the same public URL a few more times so the scale-to-zero app has a chance to start. Confirm that the address begins with https://.

Return to the live ingress JSON if the page remains unavailable. Verify that targetPort is 80 before changing anything else.

Share the current ingress JSON without credentials when you ask for help with the restored endpoint.

Document the incident response

A root-cause analysis turns a successful recovery into a reusable operations record. A structured Markdown report connects the user impact to evidence.

  • Press Cmd+Space on macOS or the Windows key on Windows to open system search.
  • Type Visual Studio Code and press Enter to open it.

You should see the Visual Studio Code window ready for your local incident report.

  • Select File from the top menu.
  • Select Open Folder....

You will see the folder picker for your local workspace.

  • Select your Desktop folder.
  • Select Open on macOS or Select Folder on Windows.

About Folder Trust

Visual Studio Code may ask you to confirm that you trust the folder. Approve the prompt only because this is your own Desktop folder.

The Explorer view should now show your Desktop as the open workspace.

  • Select the Explorer view in the left Activity Bar.
  • Select the New File... button in the Explorer view.

A filename field opens inside the Desktop workspace.

  • Enter incident-report.md as the filename.
  • Press Enter to create the file.

You should see incident-report.md in the Explorer. The file opens in the editor.

  • Add a # Summary section describing the public endpoint outage caused by the ingress target-port mismatch.
  • Add a # User Impact section containing the affected URL from your browser. Record that the Microsoft quickstart page was unavailable.
  • Open a side-by-side Markdown preview by pressing Cmd+K V on macOS or Ctrl+K V on Windows.

The preview should render the first two headings. It updates as you add the remaining evidence.

  • Add a # Timeline section covering the healthy baseline. Include the change from port 80 to port 81. Include the investigation. Finish with the restoration to port 80.
  • Add a # Detection section describing the browser error or unavailable response that exposed the outage.

The preview should now show four report sections in the same order as your response timeline.

  • Add an # Evidence section summarizing the live ingress JSON. Include relevant system-log evidence or the Ingress Port Settings Check result.
  • Add a # Root Cause section stating that ingress routed to port 81 while the image listened on port 80.

The report now connects the observed outage to the confirmed configuration mismatch.

  • Add a # Remediation section recording the Azure CLI update that restored targetPort to 80.
  • Add a # Recovery Validation section recording the live port value and the restored quickstart page.

The preview should now show a complete chain from diagnosis through verified recovery.

  • Add a # Prevention section requiring listening-port validation before ingress changes. Include a public endpoint check after each change.
  • Save incident-report.md.

Before you review the finished report, would another responder be able to connect the symptom to the evidence and recovery without replaying the lab?

✔️ Awesome, I've got everything!

Your report has the incident timeline, decisive evidence, confirmed root cause, recovery action, and prevention steps.

ⓧ I'd like to double check the report

  • Confirm the report contains all nine required section headings.
  • Confirm the report names the affected URL and ports 80 and 81.
  • Confirm the evidence references system logs or the Ingress Port Settings Check.
  • Confirm the report records restoration of targetPort to 80 and the successful quickstart page refresh.

That completes the support loop. You now have a healthy public endpoint and a written incident record grounded in the evidence you collected.

Secret mission

Prove Single-Revision Safety

An image release can fail before a new revision is ready. You will attempt an unreachable image update. You will inspect the failure while the healthy revision continues serving users.

Clean Up Your Resources

Clean Up Your Resources

Requests to your scale-to-zero Azure Container Apps app consume shared monthly Consumption plan grants when they wake it. Decide whether to keep your resources running, pause access until later, or delete them entirely.

Cost warning

The monthly grants are shared across your Azure subscription. Existing usage can exhaust them before this lab runs.

Pausing ingress prevents external requests from waking the app. Deleting the resource group removes the cloud resources that could consume additional grant allowance.

Resources you used:

  • Azure resource group aca-support-lab in centralus.
  • Container Apps environment aca-support-env with no persistent Log Analytics workspace.
  • Container app aca-support-app with external ingress at target port 80. Its revisions include the healthy quickstart revision plus the unsuccessful image revision.
  • Temporary Docker container running the official Azure Linux 3.0 Azure CLI image.
  • Local incident-report.md file created in Visual Studio Code.

Keep everything running

No Azure resource deletion is needed. Choose this option while you are actively using the incident lab.

  • Restore the documented quickstart image after the Secret Mission by running this command:
az containerapp update \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --image mcr.microsoft.com/k8se/quickstart:latest

What Does This Restore?

This update creates a valid revision from the official quickstart image. Single revision mode keeps traffic on the current healthy revision until the replacement is ready.

  • Leave external ingress enabled so the public endpoint remains available.
  • Keep the minimum replica count at 0 so the app can scale down between requests.
  • Keep incident-report.md as your record of the outage investigation.
  • Leave the authenticated Azure CLI container open if you plan to continue administering the lab now.

Pause - I'll come back to this later

Shut down public access to prevent external requests from waking the app. Your Azure resources plus incident-report.md remain available for later.

  • Restore the documented quickstart image before pausing the lab by running this command:
az containerapp update \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --image mcr.microsoft.com/k8se/quickstart:latest

Why Restore the Image?

The valid image leaves the retained lab in a known-good state. The app remains ready for your next troubleshooting session.

  • Disable external ingress by running this command:
az containerapp ingress disable \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP"

What Does Pausing Change?

Disabling ingress removes public access to the app. External requests can no longer wake it through its public endpoint.

  • Close the temporary Azure CLI session by running this command:
exit

What Happens to the Container?

The original Docker command used --rm. Docker automatically removes this temporary CLI container when the session exits.

When you return, use a fresh authenticated CLI container to restore access.

  • Start a new official Azure CLI container by running this command:
docker run --rm -it mcr.microsoft.com/azure-cli:azurelinux3.0

Why Start a Fresh Container?

The disposable container supplies the Azure CLI without changing your computer. It starts with a clean shell session.

  • Authenticate the new session through the device-code flow by running this command. The CLI prints the browser steps to complete.
az login --use-device-code

What Does Device Login Do?

The device-code flow authenticates the container through your browser. The new CLI session can then access your intended subscription.

  • Restore the shell variables needed for the ingress command by running:
export RESOURCE_GROUP=aca-support-lab
export APP_NAME=aca-support-app

Why Restore These Variables?

Shell variables belong to one container session. These values point the new session back to the existing lab resources.

  • Re-enable external ingress on target port 80 by running this command:
az containerapp ingress enable \
  --name "$APP_NAME" \
  --resource-group "$RESOURCE_GROUP" \
  --type external \
  --target-port 80 \
  --transport auto

What Does Resuming Restore?

This reconnects external HTTP traffic to the container's listening port. The public quickstart endpoint becomes available again.

Delete - I don't want to use this again

Remove all cloud resources from the dedicated lab resource group. This deletion is final, while your local incident report remains available until you remove it separately.

  • Delete aca-support-lab plus every Azure resource inside it by running this command:
az group delete \
  --name "$RESOURCE_GROUP" \
  --yes \
  --no-wait

What Does This Delete?

Deleting the resource group removes aca-support-env plus aca-support-app. It also removes the healthy revision plus the unsuccessful image revision.

The --no-wait flag returns control while Azure continues deleting the resources.

  • Close the temporary Azure CLI container by running this command:
exit

What Does Exiting Remove?

Exiting closes the authenticated shell. Docker removes the temporary container because it started with --rm.

  • Return to an authenticated Azure CLI session later by repeating the container workflow from Step 1.
  • Confirm that the resource group is gone by running this command:
az group exists --name "$RESOURCE_GROUP"

What Does This Check?

This command checks whether Azure can still find the dedicated resource group. It does not recreate or modify anything.

You will see false after deletion finishes. That result confirms the environment plus the container app are gone.

  • Switch back to Visual Studio Code.
  • Locate incident-report.md in the Explorer view.
  • Right-click incident-report.md.
  • Select Delete.
  • Confirm the deletion if Visual Studio Code prompts you.

You should no longer see incident-report.md in the VS Code file list. Your cloud lab plus its local report are now removed.

Nice Work!

Nice Work!

You did it! You completed a production-style Azure Container Apps support incident from healthy baseline through verified recovery.

You've learned how to:

  • Establish a healthy service baseline with a public quickstart endpoint. Configure external ingress on target port 80 with scale-to-zero settings.
  • Trigger a repeatable ingress outage by changing the target port from 80 to 81. Correlate real-time system logs with the live ingress configuration. Confirm the fault through Ingress Port Settings Check.
  • Restore the public endpoint by returning ingress to target port 80. Turn the response into an interview-ready report in incident-report.md.
  • Secret Mission: Prove single revision mode preserves live traffic during a failed image update. Show that the healthy revision keeps serving while the unreachable image produces a failed or degraded revision.

Ready to quiz yourself?