【发布时间】:2016-12-14 16:44:58
【问题描述】:
我一直在尝试使用 fmin_cg 来最小化 Logistic 回归的成本函数。
xopt = fmin_cg(costFn, fprime=grad, x0= initial_theta,
args = (X, y, m), maxiter = 400, disp = True, full_output = True )
这就是我调用 fmin_cg 的方式
这是我的 CostFn:
def costFn(theta, X, y, m):
h = sigmoid(X.dot(theta))
J = 0
J = 1 / m * np.sum((-(y * np.log(h))) - ((1-y) * np.log(1-h)))
return J.flatten()
这是我的毕业生:
def grad(theta, X, y, m):
h = sigmoid(X.dot(theta))
J = 1 / m * np.sum((-(y * np.log(h))) - ((1-y) * np.log(1-h)))
gg = 1 / m * (X.T.dot(h-y))
return gg.flatten()
好像是在抛出这个错误:
/Users/sugethakch/miniconda2/lib/python2.7/site-packages/scipy/optimize/linesearch.pyc in phi(s)
85 def phi(s):
86 fc[0] += 1
---> 87 return f(xk + s*pk, *args)
88
89 def derphi(s):
ValueError: operands could not be broadcast together with shapes (3,) (300,)
我知道这与我的尺寸有关。但我似乎无法弄清楚。 我是菜鸟,所以我可能会犯一个明显的错误。
我已阅读此链接:
fmin_cg: Desired error not necessarily achieved due to precision loss
但是,它似乎对我不起作用。
有什么帮助吗?
更新了 X,y,m,theta 的大小
(100, 3) ----> X
(100, 1) -----> 是的
100 ----> 米
(3, 1) ----> θ
这就是我初始化 X,y,m 的方式:
data = pd.read_csv('ex2data1.txt', sep=",", header=None)
data.columns = ['x1', 'x2', 'y']
x1 = data.iloc[:, 0].values[:, None]
x2 = data.iloc[:, 1].values[:, None]
y = data.iloc[:, 2].values[:, None]
# join x1 and x2 to make one array of X
X = np.concatenate((x1, x2), axis=1)
m, n = X.shape
ex2data1.txt:
34.62365962451697,78.0246928153624,0
30.28671076822607,43.89499752400101,0
35.84740876993872,72.90219802708364,0
.....
如果有帮助,我正在尝试用 Python 重新编写 Andrew Ng 为 Coursera 的 ML 课程编写的一项家庭作业
【问题讨论】:
-
X,Y,m,theta 的尺寸是多少?也许包括初始化这些变量的代码行。
-
更新了问题以反映您的 cmets @user2241910
-
如果您能以完整的可重现示例的形式发布您的代码,将会很有帮助。一个问题是你的目标函数似乎返回一个向量而不是一个标量,因为
m是一个(100,)数组。 -
我确实添加了初始化 X,y,m 的代码
-
目标函数是指costFn()? 'm' 似乎是一个整数,它确实从 costFn() 中返回一个“
”
标签: python optimization machine-learning scipy gradient