【问题标题】:efficiently calculating connected components in pyspark在 pyspark 中有效地计算连通分量
【发布时间】:2017-09-25 01:59:26
【问题描述】:

我正在尝试为城市中的朋友寻找连接的组件。我的数据是带有城市属性的边列表。

城市 | SRC |目的地

休斯顿凯尔 -> 本尼

休斯顿本尼 -> 查尔斯

休斯顿查尔斯 -> 丹尼

奥马哈卡罗尔 -> 布赖恩

等等。

我知道 pyspark 的 GraphX 库的 connectedComponents 函数将遍历图形的所有边以找到连接的组件,我想避免这种情况。我该怎么做?

编辑: 我以为我可以做类似的事情

从数据框中选择 connected_components(*) 按城市分组

connected_components 生成项目列表的位置。

【问题讨论】:

标签: graph spark-dataframe spark-graphx connected-components graphframes


【解决方案1】:

假设你的数据是这样的

import org.apache.spark._
import org.graphframes._

val l = List(("Houston","Kyle","Benny"),("Houston","Benny","charles"),
            ("Houston","Charles","Denny"),("Omaha","carol","Brian"),
            ("Omaha","Brian","Daniel"),("Omaha","Sara","Marry"))
var df = spark.createDataFrame(l).toDF("city","src","dst")

创建要为其运行连接组件的城市列表 cities = List("Houston","Omaha")

现在对城市列表中的每个城市的城市列运行过滤器,然后从生成的数据框创建边和顶点数据框。从这些边和顶点数据框创建一个图框并运行连接组件算法

val cities = List("Houston","Omaha")

for(city <- cities){
    val edges = df.filter(df("city") === city).drop("city")
    val vert = edges.select("src").union(edges.select("dst")).
                     distinct.select(col("src").alias("id"))
    val g = GraphFrame(vert,edges)
    val res = g.connectedComponents.run()
    res.select("id", "component").orderBy("component").show()
}

输出

|     id|   component|
+-------+------------+
|   Kyle|249108103168|
|charles|249108103168|
|  Benny|249108103168|
|Charles|721554505728|
|  Denny|721554505728|
+-------+------------+

+------+------------+                                                           
|    id|   component|
+------+------------+
| Marry|858993459200|
|  Sara|858993459200|
| Brian|944892805120|
| carol|944892805120|
|Daniel|944892805120|
+------+------------+

【讨论】:

  • 感谢您的工作!好吧,该死的。我认为可能有一些更接近金属的东西,而不仅仅是循环遍历我想要阻止的值,但我仍然感谢您的回答
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-08-23
  • 1970-01-01
  • 2019-05-10
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多