This week, my main focus was optimizing a Java service that had long been experiencing abnormal CPU behavior.

At first glance, the issue appeared to be that CPU usage could not scale up under load. A natural starting point was to check whether the service was constrained by an insufficient number of worker threads. However, further observation revealed a different pattern: CPU usage could increase, but once it did, the system quickly ran into severe timeout issues.

This suggested that the problem was not about CPU utilization itself, but rather about what happened when the system attempted to utilize more CPU.

Historically, this service had already been flagged for poor performance, and there were even suggestions to deprecate it. Based on this background, I suspected that the root cause might lie in the framework rather than the business logic.

I began by reviewing the framework implementation. The service uses Netty as its NIO server framework. Incoming requests are handled by I/O threads and then dispatched to worker threads for business processing.

In a previous issue, insufficient worker threads had caused performance bottlenecks under high I/O load. However, in this case, logs indicated that the number of worker threads was sufficient. This led me to consider whether the issue might originate from the client side.

Within the microservice architecture, this service also acts as a client when invoking other services. This became a key investigation direction.

By analyzing the code, I found that the framework uses a proxy class (ObjectProxy) based on Java’s InvocationHandler to intercept RPC calls. When a service method is invoked, ObjectProxy interacts with a ProtocolInvoker to retrieve a list of available downstream nodes from the service registry (refreshed every 30 seconds). The list is then passed to a LoadBalancer to select a target node.

The actual invocation is performed by a protocol-specific Invoker, which maintains long-lived connections to downstream services. Depending on whether the call is synchronous or asynchronous, the request is sent and handled differently.

Each downstream service consists of multiple nodes. For each node, the framework creates:

  • Two I/O threads (for NIO event handling)
  • Multiple TCP connections (by default equal to the number of CPU cores)

Each I/O thread manages a Selector, which polls for events on its associated connections. For each TCP connection, a TCPSession is maintained.

When a request is sent, a Ticket object is created to track the request-response lifecycle:

  • For synchronous calls, the thread blocks until a response is received
  • For asynchronous calls, the Ticket stores a callback, which is triggered upon response arrival

At the network level, the framework relies on Java NIO (SocketChannel). Incoming data is buffered, and packet boundaries are determined using a length-prefixed protocol. The framework checks whether the buffer contains enough bytes to form a complete packet before processing it.

Once a full packet is received, it is dispatched for further processing. At this stage, a worker thread from a thread pool is used to handle the Ticket. This thread pool is relatively small, with:

  • Core size = number of CPU cores
  • Maximum size = 2 × number of CPU cores

From a design perspective, this architecture is fairly standard and widely used. After thoroughly reviewing the client-side implementation, I did not find any obvious flaws.

At this point, it became increasingly likely that the issue was not on the client side, but rather on the server side or in the interaction between services. Further investigation would need to focus on the downstream services and their behavior under load.