【问题标题】:How do you gather the elements of y_pred that do not correspond to the true label in a Keras/tf2.0 custom loss function?如何在 Keras/tf2.0 自定义损失函数中收集与真实标签不对应的 y_pred 元素?
【发布时间】:2020-04-14 04:26:05
【问题描述】:

下面是一个简单的 numpy 示例,说明我想做的事情:

import numpy as np

y_true = np.array([0,0,1])
y_pred = np.array([0.1,0.2,0.7])

yc = (1-y_true).astype('bool')

desired = y_pred[yc]

>>> desired
>>> array([0.1, 0.2])

所以ground truth对应的预测是0.7,我想对一个包含y_pred的所有元素的数组进行操作,除了ground truth元素。

我不确定如何在 Keras 中进行这项工作。这是损失函数中问题的一个工作示例。现在'期望'没有完成任何事情,但这是我需要处理的:

# using tensorflow 2.0.0 and keras 2.3.1

import tensorflow.keras.backend as K
import tensorflow as tf
from tensorflow.keras.layers import Input,Dense,Flatten
from tensorflow.keras.models import Model
from keras.datasets import mnist

(x_train, y_train), (x_test, y_test) = mnist.load_data()

# Normalize data.
x_train = x_train.astype('float32') / 255
x_test = x_test.astype('float32') / 255

# Convert class vectors to binary class matrices.
y_train = tf.keras.utils.to_categorical(y_train, 10)
y_test = tf.keras.utils.to_categorical(y_test, 10)

input_shape = x_train.shape[1:]


x_in = Input((input_shape))

x = Flatten()(x_in)
x = Dense(256,'relu')(x)
x = Dense(256,'relu')(x)
x = Dense(256,'relu')(x)

out = Dense(10,'softmax')(x)




def loss(y_true,y_pred):


    yc = tf.math.logical_not(kb.cast(y_true, 'bool'))
    desired = tf.boolean_mask(y_pred,yc,axis = 1)    #Remove and it runs


    CE = tf.keras.losses.categorical_crossentropy(
        y_true,
        y_pred)

    L = CE

    return L



model = Model(x_in,out)

model.compile('adam',loss = loss,metrics = ['accuracy'])


model.fit(x_train,y_train)

我最终得到一个错误

ValueError: Shapes (10,) and (None, None) are incompatible

其中 10 是类别数。最终目的是实现这一点:ComplementEntropy 在 Keras,我的问题似乎是第 26-28 行。

【问题讨论】:

  • 请同时提供rest that works fine 后面的代码。所以我们有你试图使用的全部损失。
  • 我添加了一个示例,您可以运行该示例来重现相同的错误。

标签: python tensorflow keras tensorflow2.0 loss-function


【解决方案1】:

您可以从Boolean_mask 中删除axis=1,它会运行。坦率地说,我不明白你为什么在这里需要axis=1。

def loss(y_true,y_pred):


    yc = tf.math.logical_not(K.cast(y_true, 'bool'))
    print(yc.shape)
    desired = tf.boolean_mask(y_pred, yc)    #Remove axis=1 and it runs


    CE = tf.keras.losses.categorical_crossentropy(
        y_true,
        y_pred)

    L = CE

    return L

这可能就是发生的事情。你有y_pred,它是一个二维张量(N=2)。然后你有一个 2D 蒙版 (K=2)。但是有这个条件K + axis <= N。如果您通过axis=1,则会失败。

【讨论】:

  • 是的,这是有道理的。我将尝试完成其余的代码。当我运行它时,我得到了这个,WARNING:tensorflow:Entity <function Function._initialize_uninitialized_variables.<locals>.initialize_variables at 0x000002CBB3D7E168> could not be transformed and will be executed as-is. Please report this to the AutoGraph team. When filing the bug, set the verbosity to 10 (on Linux, `export AUTOGRAPH_VERBOSITY=10`) and attach the full output. Cause: 它仍然运行,但知道为什么吗?
  • @NickMerrill,不太清楚为什么会出现这种情况。我去看看
【解决方案2】:

使用因此hv89 的答案,这是我如何在LeNet 上实现COT 的完整代码,来自参考论文。一个技巧是我实际上并没有在两个目标之间来回翻转,而是只有一个随机权重翻转s

# using tensorflow 2.0.0 and keras 2.3.1

import tensorflow.keras.backend as kb
import tensorflow as tf
from tensorflow.keras.layers import Conv2D, Input, Dense,Flatten,AveragePooling2D,GlobalAveragePooling2D
from tensorflow.keras.models import Model
from keras.datasets import mnist

(x_train, y_train), (x_test, y_test) = mnist.load_data()

# Normalize data.
x_train = x_train.astype('float32') / 255
x_test = x_test.astype('float32') / 255

#exapnd dims to fit chn format
x_train = np.expand_dims(x_train,axis=3)
x_test = np.expand_dims(x_test,axis=3)


# Convert class vectors to binary class matrices.
y_train = tf.keras.utils.to_categorical(y_train, 10)
y_test = tf.keras.utils.to_categorical(y_test, 10)

input_shape = x_train.shape[1:]

x_in = Input((input_shape))

act = 'tanh'
x = Conv2D(32, (5, 5), activation=act, padding='same',strides = 1)(x_in)
x = AveragePooling2D((2, 2),strides = (2,2))(x)
x = Conv2D(16, (5, 5), activation=act)(x)
x = AveragePooling2D((2, 2),strides = (2,2))(x)

conv_out = Flatten()(x)
z = Dense(120,activation = act)(conv_out)#120
z = Dense(84,activation = act)(z)#84
last = Dense(10,activation = 'softmax')(z)

model = Model(x_in,last)



def loss(y_true,y_pred, axis=-1):

    s = kb.round(tf.random.uniform( (1,), minval=0, maxval=1, dtype=tf.dtypes.float32))
    s_ = 1 - s

    y_pred = y_pred + 1e-8

    yg = kb.max(y_pred,axis=1)
    yc = tf.math.logical_not(kb.cast(y_true, 'bool'))
    yp_c = tf.boolean_mask(y_pred, yc)  

    ygc_ = 1/(1-yg+1e-8)
    ygc_ = kb.expand_dims(ygc_,axis=1)

    Px = yp_c*ygc_ +1e-8

    COT = kb.mean(Px*kb.log(Px),axis=1)

    CE = -kb.mean(y_true*kb.log(y_pred),axis=1)

    L = s*CE +s_*(1/(10-1))*COT

    return L


model.compile(loss=loss, 
              optimizer='adam', metrics=['accuracy'])


model.fit(x_train,y_train,epochs=20,batch_size = 128,validation_data= (x_test,y_test))

pred = model.predict(x_test)

pred_label = np.argmax(pred,axis=1)
label = np.argmax(y_test,axis=1)

cor = (pred_label == label).sum()
acc = print('acc:',cor/label.shape[0])

【讨论】:

    猜你喜欢
    • 2020-02-28
    • 2017-08-20
    • 2019-10-14
    • 2020-10-12
    • 2018-12-30
    • 2018-06-10
    • 2016-12-05
    • 1970-01-01
    • 2020-03-07
    相关资源
    最近更新 更多