Back to insights

Application Gateway for Containers in production: lessons from an ingress-nginx migration

We migrated a fintech platform from ingress-nginx to Application Gateway for Containers. Azure prerequisites, provisioning time and feature availability shaped the release as much as the routing did. Here is what we would check before the next migration.

A FlowFactor engineer walks colleagues through code on a screen

When you replace ingress-nginx on a production platform, the migration plan needs to cover more than routing configuration. Azure prerequisites, provisioning time and feature availability all affect when you can safely switch traffic.

‍

We recently migrated a fintech platform we manage to Gateway API on Application Gateway for Containers. The platform runs software used by banks, so the ingress-nginx retirement announcement gave us a clear reason to move quickly. We started the migration early.

‍

Application Gateway for Containers uses a controller in your AKS cluster to keep Gateway API resources in sync with Azure-managed load balancers. The setup looked straightforward. In practice, Azure prerequisites, longer provisioning times and feature availability shaped how we prepared and carried out the migration.

‍

These are the lessons we would take into the next migration.

‍

Healthy Kubernetes resources do not tell the whole story

A healthy-looking Kubernetes configuration does not always mean Azure has everything it needs to create the underlying resources. In our setup, three prerequisites needed particular attention:

  • Register and activate the resource provider at subscription level.
  • Give the controller a user-assigned identity with the required permissions on the load balancer and network.
  • Prepare a dedicated subnet with the delegation the service requires.

‍

Missing one of these can leave you with Kubernetes resources that appear healthy while the expected Azure resources never appear. In our case, the feedback did not clearly identify the missing prerequisite.

‍

If your Gateway reports success but you cannot find the load balancer in Azure, check these dependencies first. Otherwise, you can spend time debugging routing configuration that is already correct.

‍

Five minutes on paper, thirty to sixty in our environment

The documentation we followed suggested five to six minutes for provisioning. In our environment, traffic took thirty minutes to an hour to come through. The delay came from provisioning in the Azure layer, rather than DNS caching or TTL settings.

‍

That changes your release plan. During a cutover, you need to know whether you are waiting for Azure or troubleshooting a routing problem. If your schedule assumes five minutes, that uncertainty can consume much of your release window.

‍

During disaster recovery, the same wait becomes part of your recovery time. We are updating the platform's DR test with the aim of provisioning these resources within half an hour.

‍

Measure provisioning time in your own environment and build it into both plans.

‍

One subscription setting can block a production release

Some prerequisites sat outside our normal infrastructure-as-code release flow. They involved a one-time action in the Azure portal, which made them easy to overlook when moving between environments.

‍

We encountered this during a production release. A missing subscription-level registration prevented the deployment from working, and we had to revert. Completing that registration resolved the blocker, but investigating and correcting it did not fit within the release window.

‍

Document every manual prerequisite and check it for each target subscription before the release. If acceptance and production use different subscriptions, a successful acceptance deployment does not confirm that production is ready.

‍

Separate the core application without multiplying costs

A separate application load balancer for every environment would have added cost without enough benefit for this platform. We chose two:

  • A dedicated load balancer for the core application.
  • A shared load balancer for supporting workloads, including a demo application and internal tools.

‍

The boundary followed the impact of a failure. An issue on the shared load balancer could affect supporting workloads without taking down the core application. That gave the client the separation the platform needed while keeping costs contained.

‍

Inside the cluster, we used separate namespaces and Gateways for the backend and frontend to make ownership clear. A single controller synchronises the configuration with Azure. We also split the Terraform configuration per application load balancer, leaving room to divide it further as the platform grows.

‍

Capacity limits belong in that design discussion. At the time of writing, the limits covered in this migration were five application load balancers per controller, five Gateways per load balancer and 200 routing rules per Gateway. Check those boundaries against your expected growth before deciding how to split the platform.

‍

Prepare the new route before switching traffic

Azure provides a hostname that you can use as a CNAME target for each Gateway. This avoids tying your DNS configuration directly to the underlying IP address.

‍

We deployed the new controller, load balancers, Gateways and HTTPRoutes alongside the existing ingress-nginx setup. The new services and routing configuration were in place before we directed production traffic to them.

‍

We then replaced the existing A record with the CNAME. Preparing the Azure resources ahead of time kept their provisioning delay out of the traffic switch itself. The remaining transition depended on DNS propagation and cached records, with some clients continuing to use the old route for a while.

‍

If traffic does not reach the new setup as expected, inspect the logs and decide whether to switch back or continue within a maintenance window.

‍

Keep the old Ingress resources available until you have confirmed that the new route works as expected. A few days of overlap gives you time to observe the migration and preserves a route back before you decommission the old setup.

‍

Moving early meant depending on Azure's feature timeline

Mutual TLS was a requirement for this platform. In the Gateway API setup, a FrontendTLSPolicy configures the listener to verify client certificates against the appropriate certificate authority. With Ingress controllers, this typically relies on controller-specific annotations.

‍

When we started the migration, Azure had not yet fully implemented the mutual TLS support we needed. We could prepare the migration, but completing it depended on that support becoming available. It eventually arrived.

‍

The platform's banking context gave us a reason to move quickly. The availability of a required security feature determined how quickly we could finish.

‍

Before setting a migration date, confirm that your chosen implementation supports every feature your platform depends on. Support in the API alone does not settle that question; the implementation you deploy has to support it too.

‍

Before your own migration

Our preparation checklist now includes the following:

  • Confirm resource provider registration in every target subscription, including production.
  • Check the controller identity's permissions on both the load balancer and the network.
  • Prepare the subnet and its required delegation.
  • Budget for the provisioning time you observe in testing. In our case, that meant thirty minutes to an hour.
  • Compare capacity limits with your expected services and routing rules.
  • Deploy the new setup alongside the old one, then switch traffic through DNS.
  • Verify required features, including mutual TLS, before committing to a release date.

‍

For this migration, Azure setup and timing shaped the release as much as the routing configuration did. Those dependencies deserve a place in the migration plan from the start.

‍

Planning and carrying out migrations like this is part of our managed services. If you are preparing your own move from ingress-nginx, an audit can help you assess the dependencies and release plan for your environment.

Miguel Losa

Related items

Kubernetes & OpenShift

Gateway API after ingress-nginx: don't turn one migration into two

read more
Two FlowFactor engineers working together at one screen

CI/CD Pipelines

Building a CI/CD pipeline: the choices no vendor will tell you about

read more

Security & Secrets

Dynamic PostgreSQL Credentials with HashiCorp Vault

read more