Skip to content

prov/efa: CUDA VMM buffers fail RMA when SYNC_MEMOPS is unsupported #12853

Description

@nvpohanh

[by Codex]

Describe the bug

The EFA RDM provider rejects CUDA Virtual Memory Management (VMM) buffers during RMA. Registration can succeed, but the first transfer fails when the provider tries to enable synchronous memory operations with the pointer-level CU_POINTER_ATTRIBUTE_SYNC_MEMOPS attribute.

CUDA VMM allocations created with cuMemCreate/cuMemMap do not support that pointer attribute and return CUDA_ERROR_NOT_SUPPORTED (801). cuda_set_sync_memops() converts the CUDA error to -FI_EINVAL, so fi_read/fi_write fails with Invalid argument.

This affects applications that use CUDA VMM-backed buffers, including PyTorch's expandable_segments:True allocator and custom torch.cuda.MemPool allocators built on CUDA VMM.

To Reproduce

Use two EFA-capable GPU hosts with CUDA HMEM support enabled in libfabric:

  1. On both hosts, confirm that the EFA RDM provider is available:

    fi_info -p efa -t FI_EP_RDM
  2. Establish an FI_EP_RDM connection between the hosts.

  3. On one host, reserve and map device memory with the CUDA Driver VMM API:

    • cuMemAddressReserve
    • cuMemCreate
    • cuMemMap
    • cuMemSetAccess
  4. Register that pointer as FI_HMEM_CUDA memory with fi_mr_regattr.

  5. Use the registered region as an RMA source or destination and call fi_write or fi_read.

As an application-level A/B reproduction, run the same two-node GPU transfer with an ordinary cudaMalloc buffer and then with a VMM-backed buffer. The cudaMalloc case succeeds. The VMM case reaches efa_rdm_attempt_to_sync_memops_iov() and fails before the RMA is submitted.

PyTorch can provide the VMM-backed variant by setting the allocator configuration before process startup:

export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

Then allocate the registered transfer buffer with PyTorch and repeat the same EFA RMA operation. Unsetting expandable_segments makes the transfer succeed.

Expected behavior

The EFA provider should support RMA with CUDA VMM buffers, or use a synchronization mechanism that is valid for them. For CUDA 12.1 and newer, the PSM3 provider already handles this distinction by enabling CU_CTX_SYNC_MEMOPS at the context level instead of requiring the unsupported pointer attribute for VMM allocations.

Output

The relevant output is equivalent to:

cuPointerSetAttribute(CU_POINTER_ATTRIBUTE_SYNC_MEMOPS): CUDA_ERROR_NOT_SUPPORTED
efa_rdm_attempt_to_sync_memops_iov: Unable to set memops for CUDA pointer
fi_write/fi_read: Invalid argument

The current call path is:

efa_rdm_attempt_to_sync_memops_iov()
  -> ofi_hmem_set_sync_memops(FI_HMEM_CUDA, ...)
    -> cuda_set_sync_memops()
      -> cuPointerSetAttribute(CU_POINTER_ATTRIBUTE_SYNC_MEMOPS, ...)

Environment:

  • Linux, x86_64
  • EFA provider, FI_EP_RDM
  • NVIDIA GPUs with CUDA 13
  • CUDA HMEM enabled
  • Reproduced through Mooncake Transfer Engine's EFA transport (mooncake-transfer-engine-efa-cuda13==0.3.13.post1)

Additional context

  • EFA's sync-memops path:
    static inline int efa_rdm_attempt_to_sync_memops_iov(struct efa_rdm_ep *ep, struct iovec *iov, void **desc, int num_desc)
    {
    int err = 0, i;
    struct efa_rdm_mr *efa_rdm_mr;
    if (!desc)
    return err;
    if (OFI_UNLIKELY(ep->cuda_api_permitted)) {
    for (i = 0; i < num_desc; i++) {
    efa_rdm_mr = (struct efa_rdm_mr *) desc[i];
    if (efa_rdm_mr && efa_rdm_mr->needs_sync) {
    err = cuda_set_sync_memops(iov[i].iov_base);
    if (err) {
    EFA_WARN(FI_LOG_MR,
    "Unable to set memops for "
    "cuda ptr %p\n",
    iov[i].iov_base);
    return err;
    }
    efa_rdm_mr->needs_sync = false;
    }
    }
    }
    return err;
  • Shared CUDA helper using the pointer-level attribute:

    libfabric/src/hmem_cuda.c

    Lines 290 to 318 in e2c27ad

    /**
    * @brief Set CU_POINTER_ATTRIBUTE_SYNC_MEMOPS for a cuda ptr
    * to ensure any synchronous copies are completed prior
    * to any other access of the memory region, which ensure
    * the data consistency for CUDA IPC.
    *
    * @param ptr the cuda ptr
    * @return int 0 on success, -FI_EINVAL on failure.
    */
    int cuda_set_sync_memops(void *ptr)
    {
    CUresult cu_result;
    const char *cu_error_name;
    const char *cu_error_str;
    int true_flag = 1;
    cu_result = cuda_ops.cuPointerSetAttribute(&true_flag,
    CU_POINTER_ATTRIBUTE_SYNC_MEMOPS,
    (CUdeviceptr) ptr);
    if (cu_result == CUDA_SUCCESS)
    return FI_SUCCESS;
    ofi_cuGetErrorName(cu_result, &cu_error_name);
    ofi_cuGetErrorString(cu_result, &cu_error_str);
    FI_WARN(&core_prov, FI_LOG_CORE,
    "Failed to set CU_POINTER_ATTRIBUTE_SYNC_MEMOPS: %s:%s\n",
    cu_error_name, cu_error_str);
    return -FI_EINVAL;
    }
  • PSM3's handling for VMM allocations and CUDA 12.1+ context-level sync:
    /*
    * CUDA documentation dictates the use of SYNC_MEMOPS attribute when the buffer
    * pointer received into PSM has been allocated by the application and is the
    * target of GPUDirect DMA operations.
    *
    * Normally, CUDA is permitted to implicitly execute synchronous memory
    * operations as asynchronous operations, relying on commands arriving via CUDA
    * for proper sequencing. GDR, however, bypasses CUDA, enabling races, e.g.
    * cuMemcpy sequenced before a GDR operation.
    *
    * SYNC_MEMOPS avoids this optimization.
    *
    * Note that allocations via the "VMM" API, i.e. cuMemCreate, do not support the
    * SYNC_MEMOPS pointer attribute, and will return 801 (not supported). If we're
    * using the newer context-level sync flag available in CUDA 12.1+ to avoid this
    * issue, we will not set the pointer-level sync flag here.
    */
    static void psm3_cuda_mark_buf_synchronous(const void *buf)
    {
    bool check_for_not_supported = false;
    switch (psm3_cuda_sync_mode) {
    case PSM3_CUDA_SYNC_CTX:
    #if PSM3_CUDA_HAVE_CTX_SYNC_MEMOPS
    // sync set at the context-level; nothing to do here
    return;
    #else
    // otherwise, intentional fall through to PTR behavior
    #endif
    case PSM3_CUDA_SYNC_PTR:
    // pointer level sync, handling all errors
    break;
    case PSM3_CUDA_SYNC_PTR_RELAXED:
    // pointer level sync, ignoring not supported
    check_for_not_supported = true;
    break;
    case PSM3_CUDA_SYNC_NONE:
    return;
    }
    CUresult cudaerr;
    int true_flag = 1;
    cudaerr = PSM3_CUDA_EXEC(cuPointerSetAttribute,
    &true_flag, CU_POINTER_ATTRIBUTE_SYNC_MEMOPS, (CUdeviceptr)buf);
    if_pf (check_for_not_supported && cudaerr == CUDA_ERROR_NOT_SUPPORTED) {
    #ifdef PSM_DEBUG
    // query the handle just to be sure it is in fact a VMM alloc
    CUmemGenericAllocationHandle h;
    PSM3_CUDA_CALL(cuMemRetainAllocationHandle, &h, (CUdeviceptr)buf);
    PSM3_CUDA_CALL(cuMemRelease, h);
    #endif
    return;
    }
    PSM3_CUDA_CHECK(cuPointerSetAttribute, cudaerr);

Would using CU_CTX_SYNC_MEMOPS for CUDA 12.1+ in the EFA provider, with an appropriate fallback for older CUDA versions, be the preferred fix?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions