【问题标题】:Fitting (a gaussian) with Scipy vs. ROOT et al用 Scipy 与 ROOT 等人拟合(高斯)
【发布时间】:2017-08-07 12:38:30
【问题描述】:

我现在多次偶然发现使用scipy.curve_fit 在 python 中进行拟合比使用其他工具(例如根(https://root.cern.ch/)

例如,在拟合高斯时,使用 scipy 我主要得到一条直线:

对应代码:

def fit_gauss(y, x = None):
    n = len(y)  # the number of data
    if x is None:
       x = np.arange(0,n,1)
    mean = y.mean()
    sigma = y.std()

    def gauss(x, a, x0, sigma):
        return a * np.exp(-(x - x0) ** 2 / (2 * sigma ** 2))

    popt, pcov = curve_fit(gauss, x, y, p0=[max(y), mean, sigma])

    plt.plot(x, y, 'b+:', label='data')
    plt.plot(x, gauss(x, *popt), 'ro:', label='fit')
    plt.legend()
    plt.title('Gauss fit for spot')
    plt.xlabel('Pixel (px)')
    plt.ylabel('Intensity (a.u.)')
    plt.show()

使用 ROOT,我得到了完美的匹配,甚至没有给出启动参数:

同样,对应的代码:

import ROOT
import numpy as np

y = np.array([2., 2., 11., 0., 5., 7., 18., 12., 19., 20., 36., 11., 21., 8., 13., 14., 8., 3., 21., 0., 24., 0., 12., 0., 8., 11., 18., 0., 9., 21., 17., 21., 28., 36., 51., 36., 47., 69., 78., 73., 52., 81., 96., 71., 92., 70., 84.,72., 88., 82., 106., 101., 88., 74., 94., 80., 83., 70., 78., 85., 85., 56., 59., 56., 73., 33., 49., 50., 40., 22., 37., 26., 6., 11., 7., 26., 0., 3., 0., 0., 0., 0., 0., 3., 9., 0., 31., 0., 11., 0., 8., 0., 9., 18.,9., 14., 0., 0., 6., 0.])
x = np.arange(0,len(y),1)
#yerr= np.array([0.1,0.2,0.1,0.2,0.2])
graph = ROOT.TGraphErrors()
for i in range(len(y)):
    graph.SetPoint(i, x[i], y[i])
    #graph.SetPointError(i, yerr[i], yerr[i])
func = ROOT.TF1("Name", "gaus")
graph.Fit(func)

canvas = ROOT.TCanvas("name", "title", 1024, 768)
graph.GetXaxis().SetTitle("x") # set x-axis title
graph.GetYaxis().SetTitle("y") # set y-axis title
graph.Draw("AP")

有人可以向我解释一下,为什么结果差异如此之大? scipy 中的实现是否糟糕/依赖于良好的启动参数? 有什么办法吗?我需要自动处理很多拟合,但在目标计算机上无权访问 ROOT,因此它只能与 python 一起使用。

当从 ROOT 拟合中获取结果并将它们作为起始参数提供给 scipy 时,拟合也适用于 scipy...

【问题讨论】:

  • 您在第二个代码示例中提供的数据得到了很好的输出(请参阅下面的答案)。

标签: python scipy curve-fitting gaussian


【解决方案1】:

您可能并不真的想使用ydata.mean() 作为高斯质心的初始值,或使用ydata.std() 作为方差的初始值——这些可能从xdata 中猜测得更好。我不知道这是否是造成最初麻烦的原因。

您可能会发现lmfit 库很有用。这提供了一种将您的模型函数 gauss 转换为具有 fit() 方法的模型类的方法,该方法使用从模型函数确定的命名参数。使用它,你的身材可能看起来像:

import numpy as np
import matplotlib.pyplot as plt

from lmfit import Model

def gauss(x, a, x0, sigma):
    return a * np.exp(-(x - x0) ** 2 / (2 * sigma ** 2))

ydata = np.array([2., 2., 11., 0., 5., 7., 18., 12., 19., 20., 36., 11., 21., 8., 13., 14., 8., 3., 21., 0., 24., 0., 12., 0., 8., 11., 18., 0., 9., 21., 17., 21., 28., 36., 51., 36., 47., 69., 78., 73., 52., 81., 96., 71., 92., 70., 84.,72., 88., 82., 106., 101., 88., 74., 94., 80., 83., 70., 78., 85., 85., 56., 59., 56., 73., 33., 49., 50., 40., 22., 37., 26., 6., 11., 7., 26., 0., 3., 0., 0., 0., 0., 0., 3., 9., 0., 31., 0., 11., 0., 8., 0., 9., 18.,9., 14., 0., 0., 6., 0.])

xdata = np.arange(0, len(ydata), 1)

# wrap your gauss function into a Model
gmodel = Model(gauss)
result = gmodel.fit(ydata, x=xdata, 
                    a=ydata.max(), x0=xdata.mean(), sigma=xdata.std())

print(result.fit_report())

plt.plot(xdata, ydata, 'bo', label='data')
plt.plot(xdata, result.best_fit, 'r-', label='fit')
plt.show()

还有几个附加功能。例如,您可能希望看到最合适的置信度,这将是(在主版本中,即将发布):

# add estimated band of uncertainty:
dely = result.eval_uncertainty(sigma=3)
plt.fill_between(xdata, result.best_fit-dely, result.best_fit+dely, color="#ABABAB")
plt.show()

给予:

【讨论】:

  • 我的答案的不错选择(赞成)。我喜欢不确定的部分(就像lmfit 一样;期待新版本!)。 sigma=3 部分究竟是做什么的?
【解决方案2】:

如果没有实际数据,重现您的结果并不容易,但人工创建的嘈杂数据对我来说看起来不错:

这是我正在使用的代码:

import numpy as np
import matplotlib.pyplot as plt
from scipy.optimize import curve_fit

# your gauss function
def gauss(x, a, x0, sigma):
    return a * np.exp(-(x - x0) ** 2 / (2 * sigma ** 2))

# create some noisy data
xdata = np.linspace(0, 4, 50)
y = gauss(xdata, 2.5, 1.3, 0.5)
y_noise = 0.4 * np.random.normal(size=xdata.size)
ydata = y + y_noise
# plot the noisy data
plt.plot(xdata, ydata, 'bo', label='data')

# do the curve fit using your idea for the initial guess
popt, pcov = curve_fit(gauss, xdata, ydata, p0=[ydata.max(), ydata.mean(), ydata.std()])

# plot the fit as well
plt.plot(xdata, gauss(xdata, *popt), 'r-', label='fit')

plt.show()

和你一样,我也使用 p0=[ydata.max(), ydata.mean(), ydata.std()] 作为初步猜测,对于不同的噪声强度,这似乎工作正常且稳健。

编辑

我刚刚意识到您实际上提供了数据;那么结果如下:

代码:

import numpy as np
import matplotlib.pyplot as plt
from scipy.optimize import curve_fit


def gauss(x, a, x0, sigma):
    return a * np.exp(-(x - x0) ** 2 / (2 * sigma ** 2))

ydata = np.array([2., 2., 11., 0., 5., 7., 18., 12., 19., 20., 36., 11., 21., 8., 13., 14., 8., 3., 21., 0., 24., 0., 12.,
0., 8., 11., 18., 0., 9., 21., 17., 21., 28., 36., 51., 36., 47., 69., 78., 73., 52., 81., 96., 71., 92., 70., 84.,72.,
88., 82., 106., 101., 88., 74., 94., 80., 83., 70., 78., 85., 85., 56., 59., 56., 73., 33., 49., 50., 40., 22., 37., 26.,
6., 11., 7., 26., 0., 3., 0., 0., 0., 0., 0., 3., 9., 0., 31., 0., 11., 0., 8., 0., 9., 18.,9., 14., 0., 0., 6., 0.])

xdata = np.arange(0, len(ydata), 1)

plt.plot(xdata, ydata, 'bo', label='data')

popt, pcov = curve_fit(gauss, xdata, ydata, p0=[ydata.max(), ydata.mean(), ydata.std()])
plt.plot(xdata, gauss(xdata, *popt), 'r-', label='fit')

plt.show()

【讨论】:

  • 按广告宣传。实际上,我真的不知道为什么我的版本不起作用。我最终只是简单地创建了一个新文件,复制粘贴除了配件之外的所有内容,然后添加了您对配件的答案,现在它可以工作了......不知道那里有什么疯狂的怪癖蟒蛇......
  • @user3696412:很高兴它现在已修复。 :) 我没有尝试“修复”您的代码,所以我也不知道;没有发现任何明显的问题。
猜你喜欢
  • 1970-01-01
  • 2018-02-04
  • 1970-01-01
  • 2011-02-10
  • 1970-01-01
  • 2018-08-28
  • 2017-06-14
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多