Repository navigation
Improve thread safety of shared mutable state #6932
Description
Activity
@xxo1shine For BackupManager.status, could you clarify the call path from a stale read to actual block-production behavior? Does the consensus thread directly rely on this field when deciding whether to participate in a production slot?
@jakamobiii Yes. BackupManager.status is read from the block-production path to determine whether the local node currently has the active role required for production. Since the status is updated by the backup keep-alive and UDP handling threads, a stale read may cause the production thread to temporarily act on an outdated role. This does not change consensus validation rules, because an invalid block would still be rejected by the protocol, but it may affect the availability of the individual witness by causing it to attempt or skip production based on stale local state.
- added a parent issue
on Sep 9, 2026 [DISCUSS] A related concurrency issue was observed in
BackupServerduring CI.org.tron.common.backup.BackupServerTest > test FAILED org.junit.runners.model.TestTimedOutException: test timed out after 60 seconds at sun.misc.Unsafe.park(Native Method) at java.util.concurrent.locks.LockSupport.parkNanos(LockSupport.java:215) at java.util.concurrent.locks.AbstractQueuedSynchronizer$ConditionObject.awaitNanos(AbstractQueuedSynchronizer.java:2078) at java.util.concurrent.ThreadPoolExecutor.awaitTermination(ThreadPoolExecutor.java:1475) at java.util.concurrent.Executors$DelegatedExecutorService.awaitTermination(Executors.java:675) at org.tron.common.es.ExecutorServiceManager.shutdownAndAwaitTermination(ExecutorServiceManager.java:90) at org.tron.common.backup.socket.BackupServer.close(BackupServer.java:106) at org.tron.common.backup.BackupServerTest.tearDown(BackupServerTest.java:42)If
close()runs beforebind()publisheschannel, it skips closing the channel. A subsequent successful bind leaves the server thread blocked incloseFuture().sync(), causingBackupServerTestto time out during teardown.This requires lifecycle coordination:
channelcan legitimately still be null when shutdown starts, so addingvolatilealone would not fix the race.This is another concurrency issue in the backup subsystem. It does not address the
BackupManager.statusvisibility issue described in item 8.One point to tighten for item 3:
volatileonly solves visibility, not the lifecycle race.fetchBlock()performs a null-check followed by assignment, whilefetchBlockProcess(fetchBlockInfo)operates on a previously captured object and later clearsthis.fetchBlockInfounconditionally. If the old request is completed and a new request is installed between those steps, the scheduled worker can erase the newer request; concurrent fetch callers can also both observe null and overwrite each other. AnAtomicReferencelifecycle looks safer: usecompareAndSet(null, newInfo)when installing, andcompareAndSet(observedInfo, null)on success or timeout. A deterministic test using an old scheduled snapshot followed by a newly installed request would help pin down this case.@halibobo1205, the
BackupServerlifecycle race you identified is addressed in PR #6961, which has not yet been merged. After bind completes, the code checksshutdownand immediately closes the newly bound channel if shutdown has already started. The PR also adds a deterministic regression test.This is separate from the cross-thread visibility issue involving
BackupManager.statusin item 8. Declaringstatusasvolatileaddresses the stale-read issue described there. The bind/close race involves operation ordering and requires additional lifecycle coordination;volatilealone is insufficient for that race.@lxcmyf,
volatileaddresses the cross-thread visibility issue originally described in item 3. Here,fetchBlockInfosupports an auxiliary block-fetch optimization: the normal block-fetch request has already been sent before this field is set, so losing this auxiliary state does not mean the original request is lost.The compound-operation races you identified are indeed not resolved by
volatile, but they have not yet been shown to cause persistent block-fetch failures. Given the scope of this fix, I would prefer to usevolatilefirst and avoid expanding the changes to request lifecycle management. Atomic installation and identity-checked cleanup can be evaluated separately as a follow-up improvement.Following the visibility/lifecycle distinction discussed above, I think item 5 could use a more specific invariant: each peer should be removed, accounted for, and cleaned up only once.
PeerManageralready uses a synchronized list and atomic counters, butcheck()does not acquire the class monitor used byremove(Channel). Both removal paths also decrement the counter without checking whetherpeers.remove(peer)succeeded.One possible interleaving is:
check()takes a snapshot containing a peer whosedisconnectTimeis older thanDISCONNECTION_TIME_OUT.- The regular disconnect callback (
P2pEventHandlerImpl.onDisconnect->PeerManager.remove(Channel)) removes that peer, decrements its counter, and callsonDisconnect(). check()resumes with its old snapshot. Removal now returnsfalse, but it still decrements the counter and callsonDisconnect()again.
The reverse order (callback locates the peer,
check()removes it first) has the same effect. A late callback is exactly the situationcheck()exists for, so the two paths are expected to be active at the same time by design. The visible effect today is limited to the active/passive counters in node info drifting, since the secondonDisconnect()runs on already-cleared state; the point is more that the invariant is not expressed anywhere than that the drift is harmful.Could both paths share a removal helper that updates membership and counters under the same lock and reports whether it actually removed the peer? Only the successful caller would perform cleanup, which could remain outside the membership lock.
That helper is also what makes this testable:
check()isprivate staticwith no seam to pause after the snapshot, and aPeerConnectionbuilt in a unit test has no wired services foronDisconnect(). A regression test at the helper level could remove the same peer twice (once throughremove(Channel), once directly) and verify, for both active and passive peers, that the second call reports no removal and the counter is decremented once.@waynercheung, the invariant you identified—each peer should be removed, accounted for, and cleaned up only once—is more precise.
check()and the disconnect callback can process the same peer. Currently, both paths decrement the counter even when removal fails, which can cause the counters to drift andonDisconnect()to run twice. I agree that both paths should share a synchronized helper that removes the peer, updates the counter, and reports whether removal succeeded. Only the caller that successfully removes the peer should perform cleanup. The regression test will cover repeated removal of both active and passive peers and verify that the second attempt neither changes the counter nor triggers cleanup.Thanks @xxo1shine, that matches what I had in mind for item 5; happy to review the PR.
For item 2, I traced a possible consequence that may help with prioritization. The proposed atomic check-and-set already addresses the underlying issue.
SyncService.processBlock()on the peer's message-handling thread andprocessSyncBlock()on the sync-handle-block thread can both reachsyncNext(peer). They do not share a lock protecting thesyncChainRequested == nullcheck, andforkLockis acquired only afterward, so both callers can pass the check and send aSyncBlockChainMessage.If both requests receive replies, one problematic interleaving is that the first
ChainInventoryMessageclears the pending request and the second arrives before another request is registered.ChainInventoryMsgHandler.check()then throwsBAD_MESSAGE, which maps toBAD_PROTOCOLandchannel.close(BAD_PEER_BAN_TIME), with a one-hour ban duration. This gives a code path from duplicate local requests to penalizing a healthy peer, although I have not reproduced it end to end. The window requires the message-handling caller to be stalled between the check and the assignment (for example, waiting onforkLockduring a fork switch), so I would expect this to be rare.A dedicated per-peer lock covering the check, summary generation, and assignment looks reasonable for preventing duplicate installation. It would also avoid sharing the peer monitor with
checkAndPutAdvInvRequest(). The implementation should account for the block-processing caller already holdingblockLockwhen checking lock order.For regression coverage, we could hold the first caller inside summary generation, ensure the second caller has reached lock contention, then release the first and assert that only one
SyncBlockChainMessageis sent. The response-side rejection is already covered byChainInventoryMsgHandlerTest, so the new test only needs the sender side.@waynercheung, thank you for the additional context. The
syncNext()call inprocessBlock()requests the next batch early after a block is received, while the call inprocessSyncBlock()is triggered only after block processing completes and the fetch queue is empty. They serve different purposes and occur at different stages. Normally, once the former setssyncChainRequested, the latter returns immediately.There is no need to introduce a dedicated per-peer lock here. We will reuse the existing lock to keep the null check, state update, and send operation in
syncNext()atomic. This closes the duplicate-request window under an extreme interleaving without introducing another lock or additional lock-order complexity.Reusing an existing lock works for me; the requirement is to serialize competing syncNext() calls across the null check, assignment, and send.
One lock-order detail is worth keeping in mind, since #6970 now calls syncNext() from ChainInventoryMsgHandler while holding blockLock. Extending the existing forkLock section looks like a reasonable option, preserving the existing blockLock -> forkLock order on that path. Using blockLock would also follow that order, although it would make message-handling callers contend with block processing.
The SyncService monitor is the one to avoid here: handleSyncBlock() holds it while acquiring blockLock, so making syncNext() synchronized would introduce the reverse order through ChainInventoryMsgHandler.
The atomic section also needs to cover the direct caller in SyncService.processBlock(). A same-peer concurrency regression against the real syncNext(), verifying that only one request is sent while it remains outstanding, would confirm that. Happy to review the update.
@waynercheung, noted on the lock ordering. We will account for it in the implementation and add the corresponding concurrency regression tests.
- linked a pull request that will close this issuefix(net): harden shared state concurrency #6970
on Sep 22, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNo status
Summary
There are several instances of shared mutable state in the P2P / net modules of java-tron that are accessed concurrently by multiple threads, including net workers, synchronization threads, block-fetch thread pools, and scheduled tasks, but are not protected by synchronization, volatile, or thread-safe containers.
Such unsynchronized concurrent reads and writes constitute data races under the Java Memory Model and may result in stale reads, lost updates, inconsistent collection state, and race windows in compound operations such as check-then-set. Under high connection pressure or intensive multi-threaded scheduling, these issues may affect the determinism and stability of node behavior.
We recommend adding appropriate synchronization, visibility, and atomicity guarantees to these shared states.
Root Cause
The following eight instances of shared state lack adequate concurrency protection:
SyncService.syncBlockInProcessuses a non-thread-safe HashSet and is accessed concurrently by multiple threads without synchronization. HashSet does not guarantee internal consistency under concurrent structural modifications.SyncService.syncNextperforms a check-then-set operation onpeer.syncChainRequested(checking whether it is unset before assigning a value). The operation is not atomic, leaving a race window when multiple threads enter the code concurrently. This may result in duplicate requests or inconsistent state.fetchBlockInfois not declared volatile and is concurrently read and written by three thread pools. Without a visibility guarantee, a write performed by one thread is not guaranteed to become visible to other threads in a timely manner.cheatWitnessInfoMapuses a non-thread-safe HashMap. Concurrent reads/writes or structural modifications are not safely supported and may result in lost updates or inconsistent observations. If iteration occurs concurrently with modification, it may also trigger ConcurrentModificationException.PeerManager.check()modifies the peers collection and related counters without synchronization. Reads, writes, and iteration may therefore occur concurrently and interfere with each other.BlockChainMetricManager.getDupWitness()has a write-read ordering race involvingdupWitnessBlockNum: the block-production thread first calls counterInc (making the counter key visible) and then performs put. Meanwhile, the metrics thread checks the counter key and callsdupWitnessBlockNum.get(witness). If the read occurs between these two operations, it may return null.MessageCountuses ordinary fields forszCount[],index, andtotalCount.add()performs non-atomic read-modify-write operations, whileupdate()updates the rolling window without synchronization. These methods can be invoked concurrently by multiple sender threads, potentially resulting in lost increments and inconsistent counters.BackupManager.statusis an ordinary field updated by the Backup keep-alive scheduled task and UDP event-handling thread, while being read by the consensus block-production thread, Relay scheduled task, and Metrics thread. The field is neither declared volatile nor protected by consistent synchronization. Under the Java Memory Model, reader threads may continue observing a stale role.The underlying cause is consistent across all cases: shared mutable state is accessed concurrently without adequate synchronization, visibility, or atomicity guarantees.
Impact
Suggested Fix
Apply appropriate concurrency protection to each shared state based on its access pattern:
syncBlockInProcess: Replace it with a thread-safe collection, such asCollections.synchronizedSetorConcurrentHashMap.newKeySet(), or consistently synchronize access to the set.syncChainRequestedinSyncService.syncNext: Replace the check-then-set operation with an atomic operation, such as performing both operations inside a synchronized block or using an atomic reference with CAS, eliminating the race window.fetchBlockInfo: Declare the field as volatile (or use an atomic reference) to guarantee cross-thread visibility. If compound updates are involved, synchronization should be added as well.cheatWitnessInfoMap: Replace HashMap with ConcurrentHashMap.PeerManager.check(): Synchronize reads and writes to peers and the related counters, or use thread-safe containers to ensure that modifications and iteration do not occur concurrently.getDupWitness / dupWitnessBlockNum: Ensure that the companion map becomes visible before the counter key by changing the write order todupWitnessBlockNum.put(...)followed bycounterInc(...). On the read side, usedupWitnessBlockNum.getOrDefault(witness, 0L)to eliminate the potential NullPointerException caused by automatic unboxing, providing an additional safeguard.MessageCount: Use a thread-safe implementation and synchronizeadd,add(int),getCount, andupdateconsistently to ensure cross-thread visibility and prevent lost updates.BackupManager.status: At minimum, declare the field volatile so that consensus, Relay, Metrics, and other reader threads observe the latest role.General Principle
For each shared state, identify the threads that access it and the corresponding access patterns, then choose the minimal and correct concurrency mechanism:
The goal is to avoid unprotected shared mutable state while keeping synchronization overhead and implementation complexity to a minimum.