【发布时间】:2013-08-08 21:52:05
【问题描述】:
我有两个版本的内核执行相同的任务 - 填充链接单元列表 - 两个内核之间的区别在于存储粒子位置的数据类型,第一个使用浮点数组来存储位置(4 个浮点每个粒子由于 128 位读取/写入),第二个使用 vec3f 结构数组来存储位置(一个包含 3 个浮点数的结构)。
使用 nvprof 进行一些测试,我发现第二个内核(使用 vec3f)比第一个内核运行得更快:
Time(%) Time Calls Avg Min Max Name
42.88 37.26s 2 18.63s 23.97us 37.26s adentu_grid_cuda_filling_kernel(int*, int*, int*, float*, int, _vec3f, _vec3f, _vec3i)
11.00 3.93s 2 1.97s 25.00us 3.93s adentu_grid_cuda_filling_kernel(int*, int*, int*, _vec3f*, int, _vec3f, _vec3f, _vec3i)
测试已完成,尝试使用 256 和 512000 个粒子填充链接的单元格列表。
我的问题是,这里发生了什么?我认为由于合并的内存,浮点数组应该做更好的内存访问,而不是使用具有未对齐内存的 vec3f 结构数组。我有什么误解?
这些是内核,第一个内核:
__global__ void adentu_grid_cuda_filling_kernel (int *head,
int *linked,
int *cellnAtoms,
float *pos,
int nAtoms,
vec3f origin,
vec3f h,
vec3i nCell)
{
int idx = threadIdx.x + blockIdx.x * blockDim.x;
if (idx >= nAtoms)
return;
vec3i cell;
vec3f _pos = (vec3f){(float)pos[idx*4+0], (float)pos[idx*4+1], (float)pos[idx*4+2]};
cell.x = floor ((_pos.x - origin.x)/h.x);
cell.y = floor ((_pos.y - origin.y)/h.y);
cell.z = floor ((_pos.z - origin.z)/h.z);
int c = nCell.x * nCell.y * cell.z + nCell.x * cell.y + cell.x;
int i;
if (atomicCAS (&head[c], -1, idx) != -1){
i = head[c];
while (atomicCAS (&linked[i], -1, idx) != -1)
i = linked[i];
}
atomicAdd (&cellnAtoms[c], 1);
}
这是第二个内核:
__global__ void adentu_grid_cuda_filling_kernel (int *head,
int *linked,
int *cellNAtoms,
vec3f *pos,
int nAtoms,
vec3f origin,
vec3f h,
vec3i nCell)
{
int idx = threadIdx.x + blockIdx.x * blockDim.x;
if (idx >= nAtoms)
return;
vec3i cell;
vec3f _pos = pos[idx];
cell.x = floor ((_pos.x - origin.x)/h.x);
cell.y = floor ((_pos.y - origin.y)/h.y);
cell.z = floor ((_pos.z - origin.z)/h.z);
int c = nCell.x * nCell.y * cell.z + nCell.x * cell.y + cell.x;
int i;
if (atomicCAS (&head[c], -1, idx) != -1){
i = head[c];
while (atomicCAS (&linked[i], -1, idx) != -1)
i = linked[i];
}
atomicAdd (&cellNAtoms[c], 1);
}
这是 vec3f 结构:
typedef struct _vec3f {float x, y, z} vec3f;
【问题讨论】:
-
您的内核调用是否正确cuda error checking?您可能没有为“快速”情况正确分配内存,并且内核试图越界访问,因此在启动后不久就失败了。 SO 希望您为此类问题提供SSCCE,而不仅仅是内核代码。很可能在您在此处显示的代码之外解释了差异。
-
是的,我检查了 cuda 错误,没有发生任何错误。
-
那么,为了解释时间差异,我建议发布 SSCCE,因为您所展示的两种情况实际上都是 AoS,正如您收到的答案所示。
-
感谢@RobertCrovella 提供制作 SSCCE 测试用例的建议。在新的测试用例中,两个内核的执行时间相同,所以我猜 problem 在内核之外。
标签: cuda