【发布时间】:2011-05-22 22:52:45
【问题描述】:
我正在 MATLAB 中使用 K-means 进行一些聚类。您可能知道用法如下:
[IDX,C] = kmeans(X,k)
其中 IDX 给出 X 中每个数据点的簇号,C 给出每个簇的质心。我需要获取离质心最近的数据点的索引(实际数据集 X 中的行号)。有谁知道我该怎么做? 谢谢
【问题讨论】:
标签: matlab cluster-analysis k-means
我正在 MATLAB 中使用 K-means 进行一些聚类。您可能知道用法如下:
[IDX,C] = kmeans(X,k)
其中 IDX 给出 X 中每个数据点的簇号,C 给出每个簇的质心。我需要获取离质心最近的数据点的索引(实际数据集 X 中的行号)。有谁知道我该怎么做? 谢谢
【问题讨论】:
标签: matlab cluster-analysis k-means
@Dima 提到的“蛮力方法”如下所示
%# loop through all clusters
for iCluster = 1:max(IDX)
%# find the points that are part of the current cluster
currentPointIdx = find(IDX==iCluster);
%# find the index (among points in the cluster)
%# of the point that has the smallest Euclidean distance from the centroid
%# bsxfun subtracts coordinates, then you sum the squares of
%# the distance vectors, then you take the minimum
[~,minIdx] = min(sum(bsxfun(@minus,X(currentPointIdx,:),C(iCluster,:)).^2,2));
%# store the index into X (among all the points)
closestIdx(iCluster) = currentPointIdx(minIdx);
end
要获取距离聚类中心k最近的点的坐标,请使用
X(closestIdx(k),:)
【讨论】:
蛮力方法是运行 k-means,然后将集群中的每个数据点与质心进行比较,并找到最接近它的那个。这在 matlab 中很容易做到。
另一方面,您可能想尝试k-medoids 聚类算法,它为您提供一个数据点作为每个聚类的“中心”。这是matlab implementation。
【讨论】:
其实kmeans已经给你答案了,如果我没听错的话:
[IDX,C, ~, D] = kmeans(X,k); % D is the distance of each datapoint to each of the clusters
[minD, indMinD] = min(D); % indMinD(i) is the index (in X) of closest point to the i-th centroid
【讨论】: