Skip to content

feat(cambricon): add CNCL backend for local multi-device collectives - #72

Merged
Ziminli merged 7 commits into
masterfrom
feat/add-cncl
Sep 23, 2026
Merged

Ziminli merged 7 commits into
masterfrom
feat/add-cncl

Conversation

@baominghelly

@baominghelly baominghelly commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add full Cambricon CNCL backend support and integrate it with existing CCL/MPI architecture.

This PR adds CNCL device and backend integration, common CCL communicator initialization, CNCL rank-based initialization for native and MPI-hybrid multi-threaded programs, CNCL collectives and point-to-point communication, and the required build, dispatch, and example updates.

Changes

CNCL Backend Integration

  • Add CNCL CMake configuration and auto-detection for Cambricon environments.
  • Register CNCL in backend manifests, device/backend compatibility traits, backend priorities, and bridge generation.
  • Add CNCL API wrappers for:
    • AllReduce
    • AllGather
    • Send
    • Recv
    • GetUniqueId
    • CommInitRank
    • CommDestroy
  • Add CNCL error-checking macros and Cambricon-specific API bindings.
  • Increase INFINICCL_UNIQUE_ID_BYTES to accommodate cnclCliqueId.

Communicator Initialization

  • Implement CNCL rank initialization through cnclInitComms.
  • Coordinate multiple local thread requests into a single process-local CNCL initialization call.
  • Support native multi-threaded initialization and MPI-process plus multi-threaded MxN initialization.
  • Validate clique IDs, global ranks, local devices, duplicate ranks, and duplicate devices.
  • Queue separate initialization groups until the active CNCL group completes.
  • Refine common CCL CommInitRank handling and safely copy backend-specific unique IDs.
  • Update OpenMPI communicator initialization to preserve local communicator-size metadata.

Point-to-point Communication

  • Add compatibility handling for CNCL versions that do not accept a nullptr queue.
  • Preserve asynchronous behavior for caller-provided queues while synchronizing internally created fallback queues.

Documentation Update

  • Update README.md and .github/pull_request_template.md to include CNCL as a valid supported backend.

Platform and Backend Affected

Platform

  • CPU
  • NVIDIA GPU
  • Iluvatar GPU
  • MetaX GPU
  • Moore Threads GPU
  • Cambricon MLU
  • HYGON DCU

Backend

  • OpenMPI
  • MPICH
  • NCCL/RCCL
  • MCCL
  • CNCL

OpenMPI is marked because its CommInitAll implementation is adapted to the corrected multi-handle interface; its communication path remains unchanged.

NCCL/RCCL and MCCL are marked because common ccl implementation has changes.

Performance Impact

  • No performance impact
  • Performance improved
  • Performance regression possible

The branch enables native CNCL communication instead of relying on fallback paths. No formal benchmark comparison is included.

Known Issues & Future Work

  • CNCL allows each clique to be initialized only once per process, so local initialization requests must be coordinated and serialized.
  • CNCL currently does not support float64 reduction data or the Avg reduction operation.
  • CNCL versions up to 1.30.8 require a non-null queue for point-to-point operations.
  • CommInitAll is not supported yet.
  • A future improvement could generalize communicator-role-aware backend selection for APIs that support both intra-node CCL and inter-node MPI backends.

Test Results

Test Involved Platform

  • CPU
  • NVIDIA GPU
  • Iluvatar GPU
  • MetaX GPU
  • Moore Threads GPU
  • Cambricon MLU
  • HYGON DCU

Test Involved Backend

  • OpenMPI
  • MPICH
  • NCCL/RCCL
  • MCCL
  • CNCL

Cambricon CNCL + OMPI:
ccl_mpi_hybrid_all_gather.log
ccl_mpi_hybrid_all_reduce.log
ccl_mpi_hybrid_send_recv.log

Cambricon CNCL:
ccl_all_gather.log
ccl_all_reduce.log
ccl_send_recv.log

NVIDIA NCCL + OMPI:
ccl_mpi_hybrid_all_gather.log
ccl_mpi_hybrid_all_reduce.log
ccl_mpi_hybrid_send_recv.log


Checklist

Title, Branch, and Commits

  • PR title follows Conventional Commits.
  • Branch name follows the repository branch naming convention.
  • Each commit message follows Conventional Commits.
  • Relative to the stacked base, this PR is a single squashable commit.
  • No stray merge commits from master are present.
  • No fixup!, squash!, or wip commits remain.

Scope and Design

  • Changes are minimal and limited to CNCL support plus the required CommInitAll interface correction.
  • No dead code, debug output, or unowned TODOs were introduced.
  • No unrelated formatting churn was introduced.
  • The existing public C API signatures remain unchanged.

General Code Hygiene

  • Comments are limited to non-obvious intent.
  • Modified and added files end with trailing newlines.
  • git diff --check passes.
  • Comments and error messages are in English and follow repository conventions.

C++ Specific

  • Code follows the repository C++ style.
  • Applicable formatting was checked on the validated implementation.
  • No exceptions are introduced.
  • Error handling follows repository conventions.
  • N/A: No constructor initializer-list changes are included.

Python Specific

N/A: This PR does not modify Python files.

Testing

  • Applicable local multi-device CNCL operations were built and tested successfully on Cambricon hardware.

Build, CI, and Tooling

  • CNCL backend auto-detection and build registration are included.
  • git diff --check passes; hosted CI will run on this Draft PR.

Documentation

  • README backend documentation can be added following maintainer feedback.
  • N/A: This change is additive and has no user-visible breaking change.

Security and Safety

  • No secrets, internal URLs, customer data, or personal hardware identifiers are included.
  • No third-party source code is introduced.
  • Communicator inputs and CNCL type mappings are validated before dispatch.

@baominghelly
baominghelly marked this pull request as ready for review September 9, 2026 06:06
@baominghelly
baominghelly requested a review from Ziminli September 9, 2026 06:07
@Ziminli
Ziminli force-pushed the feat/nccl-all-gather branch from 0251d7f to 3816cd5 Compare September 15, 2026 02:34
Base automatically changed from feat/nccl-all-gather to master September 15, 2026 10:23
…elevant code

 - support `GetUniqueId` and `CommInitRank` for CNCL
 - remove the currently unsupported and irrelevant CCL backend of `CommInitAll`
 - set `INFINICCL_UNIQUE_ID_BYTES` to 136 to accommodate `cnclCliqueId`
 - organize code for `CommInitAll` and `CommInitRank`
- coordinate process-local `CommInitRank` requests into one `cnclInitComms`
  call per clique
- support native multi-threaded, MPI-hybrid, and MxN initialization
- validate local ranks, global ranks, devices, and deferred initialization groups
- add CNCL `Send` and `Recv` providers through the common CCL implementation
- provide a fallback queue for CNCL versions that reject `nullptr` queues
- add CNCL error checking and preserve local communicator metadata
@Ziminli
Ziminli merged commit b0bc9ba into master Sep 23, 2026
2 checks passed
@Ziminli
Ziminli deleted the feat/add-cncl branch September 23, 2026 12:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants