【问题标题】:1D FFT plus kernel calculation with managedCUDA使用 managedCUDA 进行 1D FFT 和内核计算
【发布时间】:2015-12-08 19:10:38
【问题描述】:

我正在尝试进行 FFT 加内核计算。 FFT:托管CUDA库 kernel calc : 自己的内核

C#代码

public void cuFFTreconstruct() {
                CudaContext ctx = new CudaContext(0);
                CudaKernel cuKernel = ctx.LoadKernel("kernel_Array.ptx", "cu_ArrayInversion");

                float[] fData = new float[Resolution * Resolution * 2];
                float[] result = new float[Resolution * Resolution * 2];
                CudaDeviceVariable<float> devData = new CudaDeviceVariable<float>(Resolution * Resolution * 2);
                CudaDeviceVariable<float> copy_devData = new CudaDeviceVariable<float>(Resolution * Resolution * 2);

                int i, j;
                Random rnd = new Random();
                double avrg = 0.0;

                for (i = 0; i < Resolution; i++)
                {
                    for (j = 0; j < Resolution; j++)
                    {
                        fData[(i * Resolution + j) * 2] = i + j * 2;
                        fData[(i * Resolution + j) * 2 + 1] = 0.0f;
                    }
                }

                devData.CopyToDevice(fData);

                CudaFFTPlan1D plan1D = new CudaFFTPlan1D(Resolution * 2, cufftType.C2C, Resolution * 2);
                plan1D.Exec(devData.DevicePointer, TransformDirection.Forward);

                cuKernel.GridDimensions = new ManagedCuda.VectorTypes.dim3(Resolution / 256, Resolution, 1);
                cuKernel.BlockDimensions = new ManagedCuda.VectorTypes.dim3(256, 1, 1);

                cuKernel.Run(devData.DevicePointer, copy_devData.DevicePointer, Resolution);

                devData.CopyToHost(result);

                for (i = 0; i < Resolution; i++)
                {
                    for (j = 0; j < Resolution; j++)
                    {
                        ResultData[i, j, 0] = result[(i * Resolution + j) * 2];
                        ResultData[i, j, 1] = result[(i * Resolution + j) * 2 + 1];
                    }
                }   
                ctx.FreeMemory(devData.DevicePointer);
                ctx.FreeMemory(copy_devData.DevicePointer);
            }

内核代码

    //Includes for IntelliSense 
    #define _SIZE_T_DEFINED
    #ifndef __CUDACC__
    #define __CUDACC__
    #endif
    #ifndef __cplusplus
    #define __cplusplus
    #endif


    #include <cuda.h>
    #include <device_launch_parameters.h>
    #include <texture_fetch_functions.h>
    #include "float.h"
    #include <builtin_types.h>
    #include <vector_functions.h>

    // Texture reference
    texture<float2, 2> texref;

    extern "C"
    {
        __global__ void cu_ArrayInversion(float* data_A, float* data_B, int Resolution)
        {
            int image_x = blockIdx.x * blockDim.x + threadIdx.x;
            int image_y = blockIdx.y;

            data_B[(Resolution * image_x + image_y) * 2] = data_A[(Resolution * image_y + image_x) * 2];
            data_B[(Resolution * image_x + image_y) * 2 + 1] = data_A[(Resolution * image_y + image_x) * 2 + 1];
        }
    }

但是这个程序不能很好地工作。 发生以下错误:

ErrorLaunchFailed:执行内核时设备发生异常。常见原因包括取消引用无效的设备指针和访问越界共享内存。 上下文不能被使用,所以它必须被销毁(并且应该创建一个新的)。 此上下文中的所有现有设备内存分配都是无效的,如果程序要继续使用 CUDA,则必须重新构建。

【问题讨论】:

    标签: c# cuda managed-cuda


    【解决方案1】:

    FFT 计划将元素的数量(即复数的数量)作为参数。所以删除计划构造函数的第一个参数中的* 2。而且批次数的乘以二也没有意义……

    此外,我将使用 float2cuFloatComplex 类型(在 ManagedCuda.VectorTypes 中)来表示复数而不是两个原始浮点数。要释放内存,请使用 CudaDeviceVariable 的 Dispose 方法。否则稍后会被 GC 内部调用。

    主机代码将如下所示:

    int Resolution = 512;
    CudaContext ctx = new CudaContext(0);
    CudaKernel cuKernel = ctx.LoadKernel("kernel.ptx", "cu_ArrayInversion");
    
    //float2 or cuFloatComplex
    float2[] fData = new float2[Resolution * Resolution];
    float2[] result = new float2[Resolution * Resolution];
    CudaDeviceVariable<float2> devData = new CudaDeviceVariable<float2>(Resolution * Resolution);
    CudaDeviceVariable<float2> copy_devData = new CudaDeviceVariable<float2>(Resolution * Resolution);
    
    int i, j;
    Random rnd = new Random();
    double avrg = 0.0;
    
    for (i = 0; i < Resolution; i++)
    {
    for (j = 0; j < Resolution; j++)
    {
        fData[(i * Resolution + j)].x = i + j * 2;
        fData[(i * Resolution + j)].y = 0.0f;
    }
    }
    
    devData.CopyToDevice(fData);
    
    //Only Resolution times in X and Resolution batches
    CudaFFTPlan1D plan1D = new CudaFFTPlan1D(Resolution, cufftType.C2C, Resolution);
    plan1D.Exec(devData.DevicePointer, TransformDirection.Forward);
    
    cuKernel.GridDimensions = new ManagedCuda.VectorTypes.dim3(Resolution / 256, Resolution, 1);
    cuKernel.BlockDimensions = new ManagedCuda.VectorTypes.dim3(256, 1, 1);
    
    cuKernel.Run(devData.DevicePointer, copy_devData.DevicePointer, Resolution);
    
    devData.CopyToHost(result);
    
    for (i = 0; i < Resolution; i++)
    {
        for (j = 0; j < Resolution; j++)
        {
            //ResultData[i, j, 0] = result[(i * Resolution + j)].x;
            //ResultData[i, j, 1] = result[(i * Resolution + j)].y;
        }
    }
    
    //And better free memory using Dispose()
    //ctx.FreeMemory is only meant for raw device pointers obtained from somewhere else...
    devData.Dispose();
    copy_devData.Dispose();
    plan1D.Dispose();
    //For Cuda Memory checker and profiler:
    CudaContext.ProfilerStop();
    ctx.Dispose();
    

    【讨论】:

    • 请同时发布您更新的主机代码或仔细检查它是否符合上述代码。如果我让您的两个内核都使用我在此处发布的主机代码运行,那么一切都运行良好。 Cuda 内存检查器没有找到任何东西,也没有错误消息。
    • 感谢您的参与。我检查了我的代码。但是我找不到错误。我的程序是通过参考网站制作的:(设置:managedcuda.codeplex.com/documentatio,dll:github.com/kunzmi/managedCuda,示例代码:github.com/kunzmi/managedCuda)。当我使用 managedCUDA 制作 2D cuda 和 1D cuda 时,程序运行良好,我可以获得良好的 FFT 结果。
    • 请按照您的使用方式发布您的主机代码,以便重现您的问题。再说一遍:我上面贴的主机代码和你的内核一起运行没有问题,所以肯定有区别。
    • 我现在设法重现了您的问题:如果我将主机代码编译为 64 位,将 cuda 设备代码编译为 32 位地址空间,我会得到同样的错误。您可以尝试检查这一点并将两个代码编译为相同的硬件架构吗?实际上,从 PTX 加载内核时应该已经抛出错误,但 Nvidia 似乎已经改变了这一点......
    • 当我将架构从 32 位更改为 64 位(内核和主机)时,我的程序运行良好。感谢您的支持!!
    【解决方案2】:

    感谢您的建议。

    我尝试了建议的代码。 但是,错误仍然存​​在。 (错误:ErrorLaunchFailed:执行内核时设备发生异常。常见原因包括取消引用无效的设备指针和访问越界共享内存。无法使用上下文,因此必须销毁它(并且应该是新的已创建)。来自此上下文的所有现有设备内存分配都是无效的,如果程序要继续使用 CUDA,则必须重新构建。)

    为了使用float2,我把cu代码改成如下

     extern "C"
    {
    
    __global__ void cu_ArrayInversion(float2* data_A, float2* data_B, int Resolution)
        {
        int image_x = blockIdx.x * blockDim.x + threadIdx.x;
        int image_y = blockIdx.y;
    
        data_B[(Resolution * image_x + image_y)].x = data_A[(Resolution * image_y + image_x)].x;
        data_B[(Resolution * image_x + image_y)].y = data_A[(Resolution * image_y + image_x)].y;
    }
    

    当程序执行“cuKernel.Run”时,进程停止。

    ptx 文件

    .version 4.3
    .target sm_20
    .address_size 32
    
        // .globl   cu_ArrayInversion
    .global .texref texref;
    
    .visible .entry cu_ArrayInversion(
        .param .u32 cu_ArrayInversion_param_0,
        .param .u32 cu_ArrayInversion_param_1,
        .param .u32 cu_ArrayInversion_param_2
    )
    {
        .reg .f32   %f<5>;
        .reg .b32   %r<17>;
    
    
        ld.param.u32    %r1, [cu_ArrayInversion_param_0];
        ld.param.u32    %r2, [cu_ArrayInversion_param_1];
        ld.param.u32    %r3, [cu_ArrayInversion_param_2];
        cvta.to.global.u32  %r4, %r2;
        cvta.to.global.u32  %r5, %r1;
        mov.u32     %r6, %ctaid.x;
        mov.u32     %r7, %ntid.x;
        mov.u32     %r8, %tid.x;
        mad.lo.s32  %r9, %r7, %r6, %r8;
        mov.u32     %r10, %ctaid.y;
        mad.lo.s32  %r11, %r10, %r3, %r9;
        shl.b32     %r12, %r11, 3;
        add.s32     %r13, %r5, %r12;
        mad.lo.s32  %r14, %r9, %r3, %r10;
        shl.b32     %r15, %r14, 3;
        add.s32     %r16, %r4, %r15;
        ld.global.v2.f32    {%f1, %f2}, [%r13];
        st.global.v2.f32    [%r16], {%f1, %f2};
        ret;
    }
    

    【讨论】:

      【解决方案3】:

      感谢您的留言。

      主机代码

      using System;
      using System.Collections.Generic;
      using System.ComponentModel;
      using System.Data;
      using System.Drawing;
      using System.Linq;
      using System.Text;
      using System.Threading.Tasks;
      using System.Windows.Forms;
      using System.Drawing.Imaging;
      using ManagedCuda;
      using ManagedCuda.CudaFFT;
      using ManagedCuda.VectorTypes;
      
      
      namespace WFA_CUDA_FFT
      {
          public partial class CuFFTMain : Form
          {
              float[, ,] FFTData2D;
              int Resolution;
      
              const int cuda_blockNum = 256;
      
              public CuFFTMain()
              {
                  InitializeComponent();
                  Resolution = 1024;
              }
      
              private void button1_Click(object sender, EventArgs e)
              {
                  cuFFTreconstruct();
              }
              public void cuFFTreconstruct()
              {
                  CudaContext ctx = new CudaContext(0);
                  ManagedCuda.BasicTypes.CUmodule cumodule = ctx.LoadModule("kernel.ptx");
                  CudaKernel cuKernel = new CudaKernel("cu_ArrayInversion", cumodule, ctx);
                  float2[] fData = new float2[Resolution * Resolution];
                  float2[] result = new float2[Resolution * Resolution];
                  FFTData2D = new float[Resolution, Resolution, 2];
                  CudaDeviceVariable<float2> devData = new CudaDeviceVariable<float2>(Resolution * Resolution);
                  CudaDeviceVariable<float2> copy_devData = new CudaDeviceVariable<float2>(Resolution * Resolution);
      
                  int i, j;
                  Random rnd = new Random();
                  double avrg = 0.0;
      
                  for (i = 0; i < Resolution; i++)
                  {
                      for (j = 0; j < Resolution; j++)
                      {
                          fData[i * Resolution + j].x = i + j * 2;
                          avrg += fData[i * Resolution + j].x;
                          fData[i * Resolution + j].y = 0.0f;
                      }
                  }
      
                  avrg = avrg / (double)(Resolution * Resolution);
      
                  for (i = 0; i < Resolution; i++)
                  {
                      for (j = 0; j < Resolution; j++)
                      {
                          fData[(i * Resolution + j)].x = fData[(i * Resolution + j)].x - (float)avrg;
                      }
                  }
      
                  devData.CopyToDevice(fData);
      
                  CudaFFTPlan1D plan1D = new CudaFFTPlan1D(Resolution, cufftType.C2C, Resolution);
                  plan1D.Exec(devData.DevicePointer, TransformDirection.Forward);
      
                  cuKernel.GridDimensions = new ManagedCuda.VectorTypes.dim3(Resolution / cuda_blockNum, Resolution, 1);
                  cuKernel.BlockDimensions = new ManagedCuda.VectorTypes.dim3(cuda_blockNum, 1, 1);
      
                  cuKernel.Run(devData.DevicePointer, copy_devData.DevicePointer, Resolution);
      
                  copy_devData.CopyToHost(result);
      
                  for (i = 0; i < Resolution; i++)
                  {
                      for (j = 0; j < Resolution; j++)
                      {
                          FFTData2D[i, j, 0] = result[i * Resolution + j].x;
                          FFTData2D[i, j, 1] = result[i * Resolution + j].y;
                      }
                  }
      
                  //Clean up
                  devData.Dispose();
                  copy_devData.Dispose();
                  plan1D.Dispose();
                  CudaContext.ProfilerStop();
                  ctx.Dispose();
              }
          }
      }
      

      内核代码

      //Includes for IntelliSense 
      #define _SIZE_T_DEFINED
      #ifndef __CUDACC__
      #define __CUDACC__
      #endif
      #ifndef __cplusplus
      #define __cplusplus
      #endif
      
      
      #include <cuda.h>
      #include <device_launch_parameters.h>
      #include <texture_fetch_functions.h>
      #include "float.h"
      #include <builtin_types.h>
      #include <vector_functions.h>
      #include <vector>
      
      // Texture reference
      texture<float2, 2> texref;
      
      extern "C"
      {
          // Device code
      
          __global__ void cu_ArrayInversion(float2* data_A, float2* data_B, int Resolution)
          {
              int image_x = blockIdx.x * blockDim.x + threadIdx.x;
              int image_y = blockIdx.y;
      
              data_B[(Resolution * image_x + image_y)].y = data_A[(Resolution * image_y + image_x)].x;
              data_B[(Resolution * image_x + image_y)].x = data_A[(Resolution * image_y + image_x)].y;
          }
      }
      

      首先我用 .Net4.5 编译。 该程序无法运行,并显示错误(System.BadImageFormatException)。 然而,当 FFT 函数被注释掉时,内核程序运行。

      第二次我从 .Net 4.5 更改为 .Net 4.0。 FFT 函数有效,但内核不运行并显示错误。

      我的电脑是 windows 8.1 pro,我使用的是 Visual Studio 2013。

      【讨论】:

        猜你喜欢
        • 2014-10-14
        • 2012-07-05
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多