【发布时间】:2020-04-02 14:19:39
【问题描述】:
我有 4 个 GPU (0,1,2,3),我想在 GPU 2 上运行一个 Jupyter 笔记本,在 GPU 0 上运行另一个。因此,在执行后,
export CUDA_VISIBLE_DEVICES=0,1,2,3
对于我使用的 GPU 2 笔记本,
device = torch.device( f'cuda:{2}' if torch.cuda.is_available() else 'cpu')
device, torch.cuda.device_count(), torch.cuda.is_available(), torch.cuda.current_device(), torch.cuda.get_device_properties(1)
在创建新模型或加载模型后,
model = nn.DataParallel( model, device_ids = [ 0, 1, 2, 3])
model = model.to( device)
然后,当我开始训练模型时,我得到了,
RuntimeError Traceback (most recent call last)
<ipython-input-18-849ffcb53e16> in <module>
46 with torch.set_grad_enabled( phase == 'train'):
47 # [N, Nclass, H, W]
---> 48 prediction = model(X)
49 # print( prediction.shape, y.shape)
50 loss_matrix = criterion( prediction, y)
~/.local/lib/python3.6/site-packages/torch/nn/modules/module.py in __call__(self, *input, **kwargs)
491 result = self._slow_forward(*input, **kwargs)
492 else:
--> 493 result = self.forward(*input, **kwargs)
494 for hook in self._forward_hooks.values():
495 hook_result = hook(self, input, result)
~/.local/lib/python3.6/site-packages/torch/nn/parallel/data_parallel.py in forward(self, *inputs, **kwargs)
144 raise RuntimeError("module must have its parameters and buffers "
145 "on device {} (device_ids[0]) but found one of "
--> 146 "them on device: {}".format(self.src_device_obj, t.device))
147
148 inputs, kwargs = self.scatter(inputs, kwargs, self.device_ids)
RuntimeError: module must have its parameters and buffers on device cuda:0 (device_ids[0]) but found one of them on device: cuda:2
【问题讨论】:
-
DataParallel要求在其device_ids列表中的第一个设备上提供每个输入张量。它基本上将该设备用作分散到其他 gpus 之前的暂存区域,并且它是在从前向返回之前收集最终输出的设备。如果您希望设备 2 成为您的主要设备,我认为device_ids = [2, 0, 1, 3]会起作用,尽管我尚未对此进行测试。 -
我同意你的观点,因为设置 device_ids = [2] 可以。我希望 DataParallel 文档在这方面做得更好。我将在今天晚些时候将此评论作为答案。谢谢!
标签: parallel-processing pytorch