Your Azure App Is Fine Because the Weather Is Fine
Every App Service instance gets 128 SNAT ports. That number is guaranteed. It is also, quietly, the only number you can count on.
Beyond that floor, the stamp your app happens to be running on maintains a shared pool of additional ports fronted by the scale unit’s outbound public IPs. When your app needs a 129th port, the platform hands one out from the pool. When the pool has capacity, everyone gets what they need and nobody notices the boundary exists. When it doesn’t, you get 128 and not one more, and your outbound calls start failing.
Most apps have no idea which mode they are operating in.
Your load test passed. Production has been fine for months. The 500-request-per-second workload you benchmarked was actually consuming 400 shared-pool ports the entire time, and it worked because the stamp had them to give. Then on a Tuesday afternoon a neighbor tenant deploys a chatty background job, the shared pool drains, and your app is suddenly capped at its guaranteed floor. Nothing changed in your code. Nothing changed in your traffic. Your app broke because someone else’s app got busy.
That is the core risk. You are capacity-planning against a number you cannot see and cannot control. The 128 is the only figure that will hold under load. Everything above that is weather.

Figure 1. SNAT allocation: 128 guaranteed ports plus a shared stamp pool whose availability depends on other tenants.
The failure mode does not reproduce on demand.
By the time you open a support case and start capturing data, the noisy neighbor has moved on and the pool is healthy again. Your dashboards go green. Your developers conclude the platform is flaky. The ticket gets closed as “no repro” and the same thing happens again three weeks later.
The failure mode does not appear in pre-prod.
Staging environments live on quieter stamps with more headroom. The anti-pattern that will kill you in production tests clean. You cannot performance-test your way out of this because you are not the variable being tested. The other tenants are.
The failure mode masquerades as other problems.
Outbound HTTPS calls fail with connection timeouts. Logs blame Cosmos. Or Storage. Or the partner API. Or DNS. The application team blames the network. The network team blames the application. Everyone chases the wrong root cause for a week because the symptom points outward and the actual cause is that the SNAT allocator returned nothing.
The failure mode punishes success.
As traffic grows, dependence on the shared pool grows with it. The day you hit product-market fit is the day you find out you were living on borrowed ports.
The argument I make to customers who insist their app is fine and has been fine goes like this.
Relying on shared-pool SNAT ports is not a strategy. It is an accident that has not caught up with you yet. The platform guarantees 128 ports per instance and offers a variable number of bonus ports whose availability depends on the behavior of tenants you will never meet. Any architecture that requires more than 128 ports per instance to function correctly is one noisy neighbor away from an incident. And the incident will look like something else. And it will happen at the worst possible time.
The fix is not to hope the pool stays generous. The fix is to stop needing it.
Singleton HttpClient. Kill the per-request new HttpClient() pattern and let a static instance live for the lifetime of the process. This one change typically drops SNAT consumption by an order of magnitude for HTTP-heavy apps. Works on every version of .NET going back to 4.5. No excuse.
NAT Gateway. VNet integrate the app and route outbound through a NAT Gateway. Your ceiling moves from 128 per instance to roughly 64,000 per public IP. The ports are yours. Not shared. Not weather.

Private Endpoints. For traffic to Azure PaaS destinations, Private Endpoints remove the traffic from the SNAT budget entirely. It never touches the public path. It never touches the pool. The cheapest port is the one you never had to allocate.
These moves take you from “works when the stamp is quiet” to “works because the design does not depend on the stamp being anything in particular.”
If your production incident story includes the phrase “and it just started working again,” you were not resolved. You were rescued.
By the weather.
Weather changes.