As we enter 2023, the year ahead is expected to be challenging.
One major task is migrating all services from physical servers to the cloud. In parallel, several new team members need to be onboarded quickly and brought up to a level where they can independently handle production issues and optimize services. This will allow me to gradually transfer routine work and focus on more long-term and critical goals.
At a personal level, I also feel that I have reached a stage where my technical growth will significantly influence my direction over the next 7–8 years.
This week, my main focus was designing and standardizing logging and traceability across multiple services.
To enable cross-service log tracing, the first step is to unify the TraceId.
However, the existing services used inconsistent formats—some used Long,
others used String. In addition, the technology stacks varied across services,
making it difficult to directly adopt a standard distributed tracing framework
across the board.
Given these constraints, I chose a pragmatic approach: design a custom TraceId that is simple, compatible, and easy to roll out incrementally.
The TraceId is defined as a 16-digit Long, structured as follows:
- Prefix (2 digits):
99, indicating the unified TraceId format - Service identifier (2 digits): identifies the originating service
- Timestamp (4 digits): current microseconds (partial)
- Random component (8 digits): two concatenated 4-digit random numbers
Although this design does not guarantee absolute uniqueness, it is sufficient for practical tracing purposes in the current system.
For Java services, special care is required when generating TraceIds under
concurrency. To avoid contention, a per-thread random number generator is
recommended (e.g., using ThreadLocal).
When handling incoming requests:
- Extract the TraceId from the request if present
- If it already follows the unified format (prefix
99), reuse it - Otherwise, generate a new TraceId and store it in the MDC
For asynchronous execution, the MDC context must be explicitly propagated to child threads; otherwise, trace information will be lost.
When invoking downstream services, the TraceId must be included in the request to maintain trace continuity across services.
Finally, after the request is completed, the MDC should be cleared to avoid leaking context into subsequent requests.
This design provides a lightweight and compatible tracing mechanism that can be gradually adopted across heterogeneous services.