【问题标题】:How to use MPI_Gatherv for collecting strings of diiferent length from different processor including master node?如何使用 MPI_Gather 从包括主节点在内的不同处理器收集不同长度的字符串?
【发布时间】:2015-10-31 15:41:15
【问题描述】:

我正在尝试将来自所有处理器(包括主节点)的不同长度的不同字符串收集到主节点的单个字符串(字符数组)中。这是 MPI_Gatherv 的原型:

int MPI_Gatherv(const void *sendbuf, int sendcount, MPI_Datatype sendtype,
            void *recvbuf, const int *recvcounts, const int *displs,
            MPI_Datatype recvtype, int root, MPI_Comm comm)**.

我无法定义一些参数,例如recvbufrecvcountsdispls。任何人都可以为此提供 C 源代码示例吗?

【问题讨论】:

  • mpi-forum 上有一些使用MPI_Gatherv() 的例子。你可以调用MPI_Gather()来收集数组recvcounts中每个进程的每个字符串的长度,计算位移为displs[0]=0, displs[i]=displs[i-1]+recvcounts[i-1],为recvbuf分配足够的空间,调用MPI_Gatherv(),最后设置空终止recvbuf 的字符以确保有效的可打印字符串。

标签: c arrays parallel-processing mpi


【解决方案1】:

正如已经指出的,有很多使用MPI_Gatherv 的例子,包括这里的堆栈溢出;可以在here 找到一个答案,它首先描述了 scatter 和 gather 的工作原理,然后是 scatterv/gatherv 变体如何扩展它。

至关重要的是,对于更简单的 Gather 操作,每个块的大小相同,MPI 库可以轻松地预先计算每个块在最终编译数组中的位置;在更一般的 collectv 操作中,这不太清楚,您可以选择 - 实际上是要求 - 准确说明每个项目应该从哪里开始。

这里唯一额外的复杂性是你正在处理字符串,所以你可能不希望所有东西都被挤在一起;你需要额外的空格填充,当然最后还有一个空终止符。

假设你有五个进程想要发送字符串:

Rank 0: "Hello"    (len=5)
Rank 1: "world!"   (len=6)
Rank 2: "Bonjour"  (len=7)
Rank 3: "le"       (len=2)
Rank 4: "monde!"   (len=6)

您希望将其组装成一个全局字符串:

Hello world! Bonjour le monde!\0
          111111111122222222223
0123456789012345678901234567890

recvcounts={5,6,7,2,6};  /* just the lengths */
displs = {0,6,13,21,24}; /* cumulative sum of len+1 for padding */

可以看到位移0为0,位移i等于(recvcounts[j]+1)之和对于j=0..i-1:

   i    count[i]   count[i]+1   displ[i]   displ[i]-displ[i-1]
   ------------------------------------------------------------
   0       5          6           0    
   1       6          7           6                 6
   2       7          8          13                 7
   3       2          3          21                 8
   4       6          7          24                 3

这是直接实现的:

#include <stdio.h>
#include <string.h>
#include <stdlib.h>
#include "mpi.h"

#define nstrings 5
const char *const strings[nstrings] = {"Hello","world!","Bonjour","le","monde!"};

int main(int argc, char **argv) {

    MPI_Init(&argc, &argv); 

    int rank, size;
    MPI_Comm_rank(MPI_COMM_WORLD, &rank);
    MPI_Comm_size(MPI_COMM_WORLD, &size);

    /* Everyone gets a string */    

    int myStringNum = rank % nstrings;
    char *mystring = (char *)strings[myStringNum];
    int mylen = strlen(mystring);

    printf("Rank %d: %s\n", rank, mystring);

    /*
     * Now, we Gather the string lengths to the root process, 
     * so we can create the buffer into which we'll receive the strings
     */

    const int root = 0;
    int *recvcounts = NULL;

    /* Only root has the received data */
    if (rank == root)
        recvcounts = malloc( size * sizeof(int)) ;

    MPI_Gather(&mylen, 1, MPI_INT,
               recvcounts, 1, MPI_INT,
               root, MPI_COMM_WORLD);

    /*
     * Figure out the total length of string, 
     * and displacements for each rank 
     */

    int totlen = 0;
    int *displs = NULL;
    char *totalstring = NULL;

    if (rank == root) {
        displs = malloc( size * sizeof(int) );

        displs[0] = 0;
        totlen += recvcounts[0]+1;

        for (int i=1; i<size; i++) {
           totlen += recvcounts[i]+1;   /* plus one for space or \0 after words */
           displs[i] = displs[i-1] + recvcounts[i-1] + 1;
        }

        /* allocate string, pre-fill with spaces and null terminator */
        totalstring = malloc(totlen * sizeof(char));            
        for (int i=0; i<totlen-1; i++)
            totalstring[i] = ' ';
        totalstring[totlen-1] = '\0';
    }

    /* 
     * Now we have the receive buffer, counts, and displacements, and 
     * can gather the strings 
     */

    MPI_Gatherv(mystring, mylen, MPI_CHAR,
                totalstring, recvcounts, displs, MPI_CHAR,
                root, MPI_COMM_WORLD);


    if (rank == root) {
        printf("%d: <%s>\n", rank, totalstring);
        free(totalstring);
        free(displs);
        free(recvcounts);
    }

    MPI_Finalize();
    return 0;
}

跑步给出:

$ mpicc -o gatherstring gatherstring.c -Wall -std=c99
$ mpirun -np 5 ./gatherstring
Rank 0: Hello
Rank 3: le
Rank 4: monde!
Rank 1: world!
Rank 2: Bonjour
0: <Hello world! Bonjour le monde!>

【讨论】:

    【解决方案2】:

    MPI_Gather+MPI_Gatherv 需要计算所有等级的位移,当您的字符串长度几乎相似时,我发现这是不必要的。相反,您可以将MPI_Allreduce+MPI_Gather 与填充的字符串接收缓冲区一起使用。填充是基于使用MPI_Allreduce 计算的最长可用字符串完成的。代码如下:

    #include <stdio.h>
    #include <stdlib.h>
    #include <time.h>
    #include <string.h>
    
    #include <mpi.h>
    
    int main(int argc, char** argv) {
        MPI_Init(NULL, NULL);
        int rank;
        int nranks;
        MPI_Comm_rank(MPI_COMM_WORLD, &rank);
        MPI_Comm_size(MPI_COMM_WORLD, &nranks);
    
        srand(time(NULL) + rank);
        int my_len =  (rand() % 10) + 1; // str_len   \in [1, 9]
        int my_char = (rand() % 26) + 65; // str_char \in [65, 90] = [A, Z]
    
        char my_str[my_len + 1];
        memset(my_str, my_char, my_len);
        my_str[my_len] = '\0';
        printf("rank %d of %d has string=%s with size=%zu\n",
                rank, nranks, my_str, strlen(my_str));
    
        int max_len = 0;
        MPI_Allreduce(&my_len, &max_len, 1, 
                      MPI_INT, MPI_MAX, MPI_COMM_WORLD); 
    
        // + 1 for taking account of null pointer at the end ['\n']
        char *my_str_padded[max_len + 1]; 
        memset(my_str_padded, '\0', max_len + 1);
        memcpy(my_str_padded, my_str, my_len);
    
        char *all_str = NULL;
        if(!rank) {
            int all_len = (max_len + 1) * nranks;
            all_str = malloc(all_len * sizeof(char));   
            memset(all_str, '\0', all_len);
        }
    
        MPI_Gather(my_str_padded, max_len + 1, MPI_CHAR, 
                         all_str, max_len + 1, MPI_CHAR, 0, MPI_COMM_WORLD);
    
        if(!rank) {
            char *str_idx = all_str;
            int rank_idx = 0;
            while(*str_idx) {
                printf("rank %d sent string=%s with size=%zu\n", 
                        rank_idx, str_idx, strlen(str_idx));
                str_idx = str_idx + max_len + 1;
                rank_idx++;
            }
    
        }
    
        MPI_Finalize();
        return(0);  
    }
    

    请记住,在选择使用位移的MPI_Gather+MPI_Gatherv 和有时使用填充的MPI_AllReduce+MPI_Gather 之间存在权衡,因为前者需要更多时间来计算位移,而后者需要更多存储空间来对齐接收缓冲区。

    我还用大字符串缓冲区对这两种方法进行了基准测试,找不到任何显着的运行时差异。

    【讨论】:

    • 您正在比较 MPI_Gather()+MPI_Gatherv()(首先收集字符串长度,然后是字符串)与 MPI_Allreduce()+MPI_Gather()(获取最大字符串长度,然后收集它们),这需要额外的内存进行填充,并且可能会崩溃(尝试发送更多数据可能会导致缓冲区溢出)。我很难相信你的方法即使没有崩溃也会更快。
    • 我已经对这两种方法进行了基准测试,是的,使用MPI_Allreduce+MPI_Gather 找不到任何显着的性能提升。关于额外的内存,只要子字符串的大小几乎相同并且您 memset 它们,这种方法会产生相似的运行时间。在某些情况下,这可能效率不高,例如,当子字符串大小变化很大时。
    • 那么您是否能够找到(显着的)性能改进?
    • 从我的测试用例来看,MPI_Allreduce+GatherGather 平均快了 10 毫秒(~355.77 毫秒),子字符串大小为 1 MB,最后有一点变化以强制执行不同的大小。此外,我们跳过for-loop 来计算位移和!我相信,在某些情况下,这个答案的效果甚至比MPI_Gather + MPI_Gatherv..
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2012-12-23
    • 2017-07-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-09-23
    • 2020-07-08
    相关资源
    最近更新 更多