【问题标题】:Kolmogorov Smirnov Test in Spark (Python) not working?Spark(Python)中的Kolmogorov Smirnov测试不起作用?
【发布时间】:2017-05-19 11:06:56
【问题描述】:

我在 Python spark-ml 中进行了正态性测试,发现我认为是一个错误。

这是设置,我有一个标准化的数据集(范围 -1,到 1)。

当我做直方图时,我可以清楚地看到数据不正常:

>>> prices_norm.histogram(10)

([-1.0, -0.8, -0.6, -0.4, -0.2, 0.0, 0.2, 0.4, 0.6, 0.8, 1.0],
 [226, 269, 119, 95, 52, 26, 8, 2, 2, 5])

当我运行 Kolmgorov-Smirnov 测试时,我得到以下结果:

>>> testResults = Statistics.kolmogorovSmirnovTest(prices_norm, "norm")
>>> print testResults

Kolmogorov-Smirnov test summary:
degrees of freedom = 0 
statistic = 0.46231145770077375 
pValue = 1.742039845709087E-11 
Very strong presumption against null hypothesis: Sample follows theoretical distribution.

Kolmgorov-Smirnov 检验将 零假设 (H0) 定义为:数据遵循指定分布 (http://www.itl.nist.gov/div898/handbook/eda/section3/eda35g.htm)。

在这种情况下,p 值非常低,因此我们应该拒绝原假设。这是有道理的,因为这显然是不正常的。

那么为什么,它会说:

Sample follows theoretical distribution

这不是错的吗?它不应该说样本不遵循理论分布吗?我错过了什么吗?

【问题讨论】:

  • 我认为Sample follows theoretical distribution. 只是重申了零假设。
  • 如果数据确实服从正态分布,输出是什么?
  • @siwica,你是对的,它只是重申了零假设。

标签: python pyspark apache-spark-mllib kolmogorov-smirnov


【解决方案1】:

这把我逼疯了,所以我直接去看了源码:

git://git.apache.org/spark.git
spark/mllib/src/main/scala/org/apache/spark/mllib/stat/test/KolmogorovSmirnovTest.scala

代码正确,null Hypothesis设置为:

object NullHypothesis extends Enumeration {
  type NullHypothesis = Value
  val OneSampleTwoSided = Value("Sample follows theoretical distribution")
}

字符串消息的措辞只是重申原假设

Very strong presumption against null hypothesis: Sample follows theoretical distribution.
                                                 ________________________________________
                                                                    H0

可以说,这种措辞令人困惑,因为它可以用两种方式来解释。但这确实是正确的。

【讨论】:

    猜你喜欢
    • 2018-03-10
    • 2016-10-04
    • 2011-12-15
    • 2021-06-12
    • 2012-06-08
    • 1970-01-01
    • 1970-01-01
    • 2018-06-01
    • 1970-01-01
    相关资源
    最近更新 更多