【问题标题】:How to run ray correctly?如何正确运行光线?
【发布时间】:2020-08-12 14:52:21
【问题描述】:

试图了解如何使用ray 正确编程。

下面的结果似乎与ray 的性能改进不一致,正如here 所解释的那样。

环境:

  • Python 版本:3.6.10
  • 射线版本:0.7.4

以下是机器规格:

>>> import psutil
>>> psutil.cpu_count(logical=False)
4
>>> psutil.cpu_count(logical=True)
8
>>> mem = psutil.virtual_memory()
>>> mem.total
33707012096 # 32 GB

一、传统python多处理用Queue(multiproc_function.py):

import time
from multiprocessing import Process, Queue

N_PARALLEL = 8
N_LIST_ITEMS = int(1e8)

def loop(n, nums, q):
    print(f"n = {n}")
    s = 0
    start = time.perf_counter()
    for e in nums:
        s += e
    t_taken = round(time.perf_counter() - start, 2)
    q.put((n, s, t_taken))

if __name__ == '__main__':
    results = []

    nums = list(range(N_LIST_ITEMS))

    q = Queue()

    procs = []
    for i in range(N_PARALLEL):
        procs.append(Process(target=loop, args=(i, nums, q)))

    for proc in procs:
        proc.start()

    for proc in procs:
        n, s, t_taken = q.get()
        results.append((n, s, t_taken))

    for proc in procs:
        proc.join()

    for r in results:
        print(r)

结果是:

$ time python multiproc_function.py
n = 0
n = 1
n = 2
n = 3
n = 4
n = 5
n = 6
n = 7
(0, 4999999950000000, 11.12)
(1, 4999999950000000, 11.14)
(2, 4999999950000000, 11.1)
(3, 4999999950000000, 11.23)
(4, 4999999950000000, 11.2)
(6, 4999999950000000, 11.22)
(7, 4999999950000000, 11.24)
(5, 4999999950000000, 11.54)

real    0m19.156s
user    1m13.614s
sys     0m24.496s

在运行期间检查 htop 时,内存从 2.6 GB 基本消耗变为 8 GB,并且所有 8 个处理器都已完全消耗。另外,从user+sys > real 可以清楚地看出并行处理正在发生。

这是光线测试代码(ray_test.py):

import time
import psutil
import ray

N_PARALLEL = 8
N_LIST_ITEMS = int(1e8)

use_logical_cores = False
num_cpus = psutil.cpu_count(logical=use_logical_cores)
if use_logical_cores:
    print(f"Setting num_cpus to # logical cores  = {num_cpus}")
else:
    print(f"Setting num_cpus to # physical cores = {num_cpus}")
ray.init(num_cpus=num_cpus)

@ray.remote
def loop(nums, n):
    print(f"n = {n}")
    s = 0
    start = time.perf_counter()
    for e in nums:
        s += e
    t_taken = round(time.perf_counter() - start, 2)
    return (n, s, t_taken)

if __name__ == '__main__':
    nums = list(range(N_LIST_ITEMS))
    list_id = ray.put(nums)
    results = ray.get([loop.remote(list_id, i) for i in range(N_PARALLEL)])
    for r in results:
        print(r)

结果是:

$ time python ray_test.py
Setting num_cpus to # physical cores = 4
2020-04-28 16:52:51,419 INFO resource_spec.py:205 -- Starting Ray with 18.16 GiB memory available for workers and up to 9.11 GiB for objects. You can adjust these settings with ray.remote(memory=<bytes>, object_store_memory=<bytes>).
(pid=78483) n = 2
(pid=78485) n = 1
(pid=78484) n = 3
(pid=78486) n = 0
(pid=78484) n = 4
(pid=78483) n = 5
(pid=78485) n = 6
(pid=78486) n = 7
(0, 4999999950000000, 5.12)
(1, 4999999950000000, 5.02)
(2, 4999999950000000, 4.8)
(3, 4999999950000000, 4.43)
(4, 4999999950000000, 4.64)
(5, 4999999950000000, 4.61)
(6, 4999999950000000, 4.84)
(7, 4999999950000000, 4.99)

real    0m45.082s
user    0m22.163s
sys     0m10.213s

real 的时间比 python 多处理的时间长得多。此外,real 大于 user+sys。在检查htop 时,内存达到了 30 GB,并且内核也没有完全饱和。所有这些似乎都与ray 应该做的事情相矛盾。

然后我将use_logical_cores 设置为True。由于内存不足,运行被终止:

$ time python ray_test.py
Setting num_cpus to # logical cores  = 8
2020-04-28 16:27:43,709 INFO resource_spec.py:205 -- Starting Ray with 17.29 GiB memory available for workers and up to 8.65 GiB for objects. You can adjust these settings with ray.remote(memory=<bytes>, object_store_memory=<bytes>).
Killed

real    0m25.205s
user    0m15.056s
sys     0m4.028s

我在这里做错了吗?

【问题讨论】:

  • 如果将列表更改为数组nums = np.arange(N_LIST_ITEMS),则运行时会减半。
  • 在这种情况下,每个 Ray worker 都会反序列化整数列表,这会占用大量内存和时间。 Numpy 数组避免了这个问题,因为 Ray 将它们存储在共享内存中。
  • 我认为在您的示例中,多处理在后台使用fork 创建程序的副本,因此工作人员自动拥有列表的副本,从而避免了反序列化步骤。这可以在像这样的特殊情况下工作,您有一个大型(且相对简单)的对象,您希望所有工作人员都拥有一个副本。
  • @RobertNishihara 感谢您的回复。我试过nums = np.arange(N_LIST_ITEMS) 而不是nums = list(range(N_LIST_ITEMS))。它运行得更慢,需要real 0m54.076s user 0m7.388s sys 0m2.585s

标签: python python-3.x parallel-processing multiprocessing ray


【解决方案1】:

首先,Ray 不保证 CPU 亲和力或资源隔离。这可能是它具有非饱和 CPU 使用率的原因。 (虽然我不是 100% 确定)。您可以尝试使用 psutil 设置 cpu 亲和性,看看核心是否仍未饱和。 (https://psutil.readthedocs.io/en/latest/#psutil.Process.cpu_affinity)。

关于结果,您介意尝试最新版本的 Ray 吗?从 0.7.4 版本开始,Ray 在性能和内存管理方面取得了相当不错的进步。

【讨论】:

  • 谢谢。我升级到 ray v. 0.8.4。它的运行速度快了大约 5%。我使用p = psutil.Process(); p.cpu_affinity([])将cpu affinity设置为使用所有内核,结果是一样的。
  • 内存使用情况如何?另外,你应该这样做stackoverflow.com/questions/61051911/…
猜你喜欢
  • 2011-03-14
  • 1970-01-01
  • 2017-07-02
  • 1970-01-01
  • 2011-09-27
  • 1970-01-01
  • 2017-03-30
  • 1970-01-01
  • 2010-10-26
相关资源
最近更新 更多