This week was primarily focused on ensuring the stability of various services ahead of the Chinese New Year.

Recently, a particular service frequently reported timeouts during peak traffic hours. I initially notified the service owner to investigate the issue. However, after several days, no clear root cause was identified.

Given the severity of the alarms—with timeout rates reaching up to 20% on some nodes—I decided to step in and troubleshoot the issue myself. During this period, overall traffic had increased significantly, roughly doubling compared to the end of December, likely due to the approaching holiday. My first assumption was insufficient service capacity, so I performed a capacity expansion.

However, the expansion did not improve the situation. The timeout rate and alarm frequency remained largely unchanged.

I then began analyzing the service’s source code. I found that each request first called a downstream service and then performed an asynchronous database write. Since the database operation was asynchronous and non-blocking, it was unlikely to be the bottleneck—although I had initially spent quite some time investigating that direction. This led me to suspect the downstream service call instead.

I logged into one of the nodes and examined the logs. It turned out that all downstream requests were consistently routed to the same IP and port—specifically, through a Scarecrow node. To eliminate potential issues caused by Scarecrow forwarding, I migrated the service to the cloud.

However, after switching traffic, I observed a large number of timeouts occurring in a different geographic region. Even after scaling resources, the issue persisted, which was quite puzzling. I temporarily rolled back the change.

Curious about why the problem worsened after moving to the cloud, I conducted further monitoring and analysis. I discovered that after traffic was redirected, one of the downstream service nodes in the cloud had extremely high CPU usage, while the others remained underutilized.

This strongly suggested a load balancing issue.

I revisited the service code and confirmed the suspicion: after retrieving the list of available downstream nodes from the service registry, the client always selected the first node in the list. As a result, nearly all traffic from a specific geographic region was routed to a single node.

Previously, when traffic passed through Scarecrow, load balancing was implicitly handled at that layer. This masked the issue, as Scarecrow nodes are proxy services capable of handling high concurrency without significant performance degradation. Once traffic bypassed Scarecrow and was routed directly to business nodes, the imbalance became critical, leading to severe overload and timeouts.

To validate this hypothesis, I checked the monitoring data for Scarecrow and found that one node had indeed been heavily overloaded, with CPU usage exceeding 95%.

After identifying the root cause, I proceeded with a two-step solution. First, to quickly mitigate the issue before the holiday, I scaled up the downstream service by doubling the CPU cores of each node. This significantly increased processing capacity and reduced timeout rates.

Then, I gradually shifted traffic back to the cloud. As expected, one node initially experienced higher CPU usage, but the system remained stable overall. Eventually, the service stabilized, alarms were cleared, and the timeout rate dropped to zero.

The final step was to fix the underlying issue. After the holiday, I modified the service code to introduce a proper load balancing mechanism using a random selection strategy. Once the new version was deployed, traffic was evenly distributed across all downstream nodes, and CPU usage became balanced.

At that point, the problem was fully resolved.