【问题标题】:start two scripts in parallel and stop one based on the other’s return并行启动两个脚本并根据另一个的返回停止一个脚本
【发布时间】:2019-04-22 08:33:48
【问题描述】:

我想在不同的 GPU 上并行启动两个不同的 python 脚本(tensorflow 对象检测 train.py 和 eval.py),当 train.py 完成后,杀死 eval.py。

我有以下代码来并行启动两个子进程 (How to terminate a python subprocess launched with shell=True)。但是子进程是在同一个设备上启动的(我可以猜到原因。我只是不知道如何在不同的设备上启动它们)。

start_train = “CUDA_DEVICE_ORDER= PCI_BUS_ID CUDA VISIBLE_DEVICES=0 train.py ...”

start_eval = “CUDA_DEVICE_ORDER= PCI_BUS_ID CUDA VISIBLE_DEVICES=1 eval.py ...”

commands = [start_train, start_eval]

procs = [subprocess.Popen(i, shell=True, stdout=subprocess.PIPE, preexec_fn=os.setsid) for i in commands]

在这一点之后,我不知道如何进行。我需要像下面这样的东西吗?我应该使用p.communicate() 来避免死锁吗?或者如果我只为 train.py 调用 wait() 或communicate() 就足够了,因为我只需要它的完成。

for p in procs:
    p.wait() # I assume this command won’t affect the parallel running

然后我需要以某种方式使用以下命令。我不需要来自 train.py 的返回值,而是来自 subprocess 的返回码。 Popen.returncode documentationwait() 和communicate() 看起来需要一个返回码设置。我不明白如何设置。我更喜欢像

if train is done without any error:
    os.killpg(os.getpgid(procs[1].pid), signal.SIGTERM) 
else:
    write the error to the console, or to a file (but how?)

或者?

train_return = proc[0].wait() 
if train_return == 0:
    os.killpg(os.getpgid(procs[1].pid), signal.SIGTERM) 

解决问题后的更新:

这是我的主要内容:

if __name__ == "__main__":
    exp = 1
    go = True
    while go:


        create_dir(os.path.join(MAIN_PATH,'kitti',str(exp),'train'))
        create_dir(os.path.join(MAIN_PATH,'kitti',str(exp),'eval'))


        copy_tree(os.path.join(MAIN_PATH,"kitti/eval_after_COCO"), os.path.join(MAIN_PATH,"kitti",str(exp),"eval"))
        copy_tree(os.path.join(MAIN_PATH,"kitti/train_after_COCO"), os.path.join(MAIN_PATH,"kitti",str(exp),"train"))

        err_log = open('./kitti/'+str(exp)+'/error_log' + str(exp) + '.txt', 'w')

        train_command = CUDA_COMMAND_PREFIX + "0 python3 " + str(MAIN_PATH) + "legacy/train.py \
                                            --logtostderr --train_dir " + str(MAIN_PATH) + "kitti/" \
                                            + str(exp) + "/train/ --pipeline_config_path " + str(MAIN_PATH) \
                                            + "kitti/faster_rcnn_resnet101_coco.config"
        eval_command = CUDA_COMMAND_PREFIX + "1 python3 " + str(MAIN_PATH) + "legacy/eval.py \
                                            --logtostderr --eval_dir " + str(MAIN_PATH) + "kitti/" \
                                            + str(exp) + "/eval/ --pipeline_config_path " + str(MAIN_PATH) \
                                            + "kitti/faster_rcnn_resnet101_coco.config --checkpoint_dir " + \
                                            str(MAIN_PATH) + "kitti/" + str(exp) + "/train/"

        os.system("python3 dataset_tools/random_sampler_with_replacement.py --random_set_id " + str(exp))
        time.sleep(20)
        update_train_set(exp)



        train_proc = subprocess.Popen(train_command,
                                  stdout=subprocess.PIPE,
                                  stderr=err_log, # write errors to a file
                                  shell=True)
        time.sleep(20)      
        eval_proc = subprocess.Popen(eval_command,
                                 stdout=subprocess.PIPE,
                                 shell=True)
        time.sleep(20)

        if train_proc.wait() == 0: # successfull termination
            os.killpg(os.getpgid(eval_proc.pid), subprocess.signal.SIGTERM)

        clean_train_set(exp)
        time.sleep(20)
        exp += 1
        if exp == 51:
            go = False

【问题讨论】:

    标签: python-3.x subprocess python-multiprocessing kill-process


    【解决方案1】:

    默认情况下,即使您有多个 GPU,TensorFlow 也会将操作分配给“/gpu:0”(或“/cpu:0”)。解决它的唯一方法是使用上下文管理器在您的一个脚本中手动将每个操作分配给第二个 GPU

    with tf.device("/gpu:1"):
        # your ops here
    

    更新

    如果我理解正确,您需要的是以下内容:

    import subprocess
    import os
    err_log = open('error_log.txt', 'w')
    train_proc = subprocess.Popen(start_train,
                                  stdout=subprocess.PIPE,
                                  stderr=err_log, # write errors to a file
                                  shell=True)
    eval_proc = subprocess.Popen(start_eval,
                                 stdout=subprocess.PIPE,
                                 shell=True)
    
    if train_proc.wait() == 0: # successfull termination
        os.killpg(os.getpgid(eval_proc.pid), subprocess.signal.SIGTERM)
    # else, errors will be written to the 'err_log.txt' file
    

    【讨论】:

    • 我可以使用该命令在首选 gpu 上启动脚本。但是我的问题不同..
    • 谢谢!这是工作。不过我有一个问题,我想循环使用这个设置(比如 100 次)。换句话说,在 train 完成并且 eval 被杀死之后,我想像那些一样启动另一对子进程。然而,在 eval 被杀死后,什么也没有重新开始,并且“终止”打印在我的控制台上。如何防止这种情况发生,并在 eval 被杀死后启动其他进程?
    • 日志文件没有错误,表示训练完成,模型保存到磁盘。在控制台上,再次没有错误,它只显示“终止”。我将我的主要方法添加到我的第一个问题,因为评论太长了
    • 我认为os.killpg(os.getpgid(eval_proc.pid), subprocess.signal.SIGTERM) 正在杀死主进程,因此它不能超出那条线。我尝试了 try/excepts,但是我还没有发现错误
    • eval_proc.kill()代替os.killpg(os.getpgid(eval_proc.pid), subprocess.signal.SIGTERM)
    猜你喜欢
    • 2018-02-04
    • 2020-08-28
    • 1970-01-01
    • 1970-01-01
    • 2022-08-22
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多