【问题标题】:Cuda version issue while using Detectron2 in Google Colab在 Google Colab 中使用 Detectron2 时出现 Cuda 版本问题
【发布时间】:2020-10-02 02:30:39
【问题描述】:

我正在尝试使用 CUDA 10.0 版在 Colab 上运行 Detectron2 模块,但从今天开始,Cuda 编译器的版本出现了一些问题。

运行!nvidia-smi 后得到的输出是:

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 450.36.06    Driver Version: 418.67       CUDA Version: 10.1     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla P100-PCIE...  Off  | 00000000:00:04.0 Off |                    0 |
| N/A   36C    P0    26W / 250W |      0MiB / 16280MiB |      0%      Default |
|                               |                      |                 ERR! |
+-------------------------------+----------------------+----------------------+

+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|  No running processes found                                                 |
+-----------------------------------------------------------------------------+

我在运行!nvcc --version 后得到的是:

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2018 NVIDIA Corporation
Built on Sat_Aug_25_21:08:01_CDT_2018
Cuda compilation tools, release 10.0, V10.0.130

我无法理解不匹配的原因。运行!python -m detectron2.utils.collect_env 后检测器的输出也是:

----------------------  ----------------------------------------------------------------------------
sys.platform            linux
Python                  3.6.9 (default, Apr 18 2020, 01:56:04) [GCC 8.4.0]
numpy                   1.18.5
detectron2              0.1.3 @/content/gdrive/My Drive/Data/Table_Struct/detectron2_repo/detectron2
Compiler                GCC 7.5
CUDA compiler           CUDA 10.1
detectron2 arch flags   sm_60
DETECTRON2_ENV_MODULE   <not set>
PyTorch                 1.4.0+cu100 @/usr/local/lib/python3.6/dist-packages/torch
PyTorch debug build     False
GPU available           True
GPU 0                   Tesla K80
CUDA_HOME               /usr/local/cuda
Pillow                  7.0.0
torchvision             0.5.0+cu100 @/usr/local/lib/python3.6/dist-packages/torchvision
torchvision arch flags  sm_35, sm_50, sm_60, sm_70, sm_75
fvcore                  0.1.1
cv2                     4.1.2
----------------------  ----------------------------------------------------------------------------
PyTorch built with:
  - GCC 7.3
  - Intel(R) Math Kernel Library Version 2019.0.4 Product Build 20190411 for Intel(R) 64 architecture applications
  - Intel(R) MKL-DNN v0.21.1 (Git Hash 7d2fd500bc78936d1d648ca713b901012f470dbc)
  - OpenMP 201511 (a.k.a. OpenMP 4.5)
  - NNPACK is enabled
  - CUDA Runtime 10.0
  - NVCC architecture flags: -gencode;arch=compute_37,code=sm_37;-gencode;arch=compute_50,code=sm_50;-gencode;arch=compute_60,code=sm_60;-gencode;arch=compute_61,code=sm_61;-gencode;arch=compute_70,code=sm_70;-gencode;arch=compute_75,code=sm_75;-gencode;arch=compute_37,code=compute_37
  - CuDNN 7.6.3
  - Magma 2.5.1
  - Build settings: BLAS=MKL, BUILD_NAMEDTENSOR=OFF, BUILD_TYPE=Release, CXX_FLAGS= -Wno-deprecated -fvisibility-inlines-hidden -fopenmp -DUSE_FBGEMM -DUSE_QNNPACK -DUSE_PYTORCH_QNNPACK -O2 -fPIC -Wno-narrowing -Wall -Wextra -Wno-missing-field-initializers -Wno-type-limits -Wno-array-bounds -Wno-unknown-pragmas -Wno-sign-compare -Wno-unused-parameter -Wno-unused-variable -Wno-unused-function -Wno-unused-result -Wno-strict-overflow -Wno-strict-aliasing -Wno-error=deprecated-declarations -Wno-stringop-overflow -Wno-error=pedantic -Wno-error=redundant-decls -Wno-error=old-style-cast -fdiagnostics-color=always -faligned-new -Wno-unused-but-set-variable -Wno-maybe-uninitialized -fno-math-errno -fno-trapping-math -Wno-stringop-overflow, DISABLE_NUMA=1, PERF_WITH_AVX=1, PERF_WITH_AVX2=1, PERF_WITH_AVX512=1, USE_CUDA=ON, USE_EXCEPTION_PTR=1, USE_GFLAGS=OFF, USE_GLOG=OFF, USE_MKL=ON, USE_MKLDNN=ON, USE_MPI=OFF, USE_NCCL=ON, USE_NNPACK=ON, USE_OPENMP=ON, USE_STATIC_DISPATCH=OFF, 

我的猜测是 Colab 上的 CUDA 版本与我使用的 Detectron2 不匹配。如果是这样,我该如何更改才能使其在 Google Colab 上运行。

【问题讨论】:

  • 没有不匹配。驱动程序告诉您(通过 nvidia-smi)它支持的最大 cuda 版本。而已。您正在尝试解决一个不存在的问题
  • 那么我为 Cuda 10.0 构建的 pytorch 应该可以正常工作吗?但我收到错误“RuntimeError: CUDA error: no kernel image is available for execution on the device”
  • 我无法评论 PyTorch。如果这就是您的实际问题,那么将其设为 PyTorch 问题
  • 我认为这是由于 Pytorch 和 CUDA 版本不匹配造成的。 pytorch 是为 CUDA 10.0 版构建的,但系统不提供该环境。
  • 可能是 CUDA 工具包版本,而不是 CUDA 驱动程序。这不是一个与 CUDA 编程相关的问题——它是关于让 PyTorch 在 colab 上运行,我认为这不是 Stack Overflow 的真正主题

标签: pytorch google-colaboratory


【解决方案1】:

问题出在编译的 Detectron2 Cuda 运行时版本上,一旦我重新编译 Detectron2,错误就解决了。

这是!python -m detectron2.utils.collect_env 命令的结果:

----------------------  ----------------------------------------------------------------------------
sys.platform            linux
Python                  3.6.9 (default, Apr 18 2020, 01:56:04) [GCC 8.4.0]
numpy                   1.18.5
detectron2              0.1.3 @/content/gdrive/My Drive/Data/Table_Struct/detectron2_repo/detectron2
Compiler                GCC 7.5
CUDA compiler           CUDA 10.0
detectron2 arch flags   sm_75
DETECTRON2_ENV_MODULE   <not set>
PyTorch                 1.4.0+cu100 @/usr/local/lib/python3.6/dist-packages/torch
PyTorch debug build     False
GPU available           True
GPU 0                   Tesla T4
CUDA_HOME               /usr/local/cuda
Pillow                  7.0.0
torchvision             0.5.0+cu100 @/usr/local/lib/python3.6/dist-packages/torchvision
torchvision arch flags  sm_35, sm_50, sm_60, sm_70, sm_75
fvcore                  0.1.1
cv2                     4.1.2
----------------------  ----------------------------------------------------------------------------
PyTorch built with:
  - GCC 7.3
  - Intel(R) Math Kernel Library Version 2019.0.4 Product Build 20190411 for Intel(R) 64 architecture applications
  - Intel(R) MKL-DNN v0.21.1 (Git Hash 7d2fd500bc78936d1d648ca713b901012f470dbc)
  - OpenMP 201511 (a.k.a. OpenMP 4.5)
  - NNPACK is enabled
  - CUDA Runtime 10.0
  - NVCC architecture flags: -gencode;arch=compute_37,code=sm_37;-gencode;arch=compute_50,code=sm_50;-gencode;arch=compute_60,code=sm_60;-gencode;arch=compute_61,code=sm_61;-gencode;arch=compute_70,code=sm_70;-gencode;arch=compute_75,code=sm_75;-gencode;arch=compute_37,code=compute_37
  - CuDNN 7.6.3
  - Magma 2.5.1
  - Build settings: BLAS=MKL, BUILD_NAMEDTENSOR=OFF, BUILD_TYPE=Release, CXX_FLAGS= -Wno-deprecated -fvisibility-inlines-hidden -fopenmp -DUSE_FBGEMM -DUSE_QNNPACK -DUSE_PYTORCH_QNNPACK -O2 -fPIC -Wno-narrowing -Wall -Wextra -Wno-missing-field-initializers -Wno-type-limits -Wno-array-bounds -Wno-unknown-pragmas -Wno-sign-compare -Wno-unused-parameter -Wno-unused-variable -Wno-unused-function -Wno-unused-result -Wno-strict-overflow -Wno-strict-aliasing -Wno-error=deprecated-declarations -Wno-stringop-overflow -Wno-error=pedantic -Wno-error=redundant-decls -Wno-error=old-style-cast -fdiagnostics-color=always -faligned-new -Wno-unused-but-set-variable -Wno-maybe-uninitialized -fno-math-errno -fno-trapping-math -Wno-stringop-overflow, DISABLE_NUMA=1, PERF_WITH_AVX=1, PERF_WITH_AVX2=1, PERF_WITH_AVX512=1, USE_CUDA=ON, USE_EXCEPTION_PTR=1, USE_GFLAGS=OFF, USE_GLOG=OFF, USE_MKL=ON, USE_MKLDNN=ON, USE_MPI=OFF, USE_NCCL=ON, USE_NNPACK=ON, USE_OPENMP=ON, USE_STATIC_DISPATCH=OFF, 

【讨论】:

    猜你喜欢
    • 2021-04-21
    • 1970-01-01
    • 2019-07-25
    • 1970-01-01
    • 2020-01-07
    • 2021-10-08
    • 1970-01-01
    • 2020-09-06
    • 1970-01-01
    相关资源
    最近更新 更多