【问题标题】:"Tap" a specific layer in existing Keras Model and make a branch to a new output?“点击”现有 Keras 模型中的特定层并分支到新输出?
【发布时间】:2020-03-29 19:28:16
【问题描述】:

环境:

我在 Google Colab 上使用 TF.Keras (Tensorflow 1.14),我的模型架构是 MobileNet V2 1.00 224。

问题:

我正在尝试(但失败)附加一个新层并向现有层进行新输出,该层不是我的模型的正常输出。即在 MobileNet V2 中更早地创建一个分支

我希望这个新分支用于回归输出 - 但我不希望该输出从 MobileNet 的最终嵌入层串行连接,而是更早的阶段(我不确定,我正在试验)。基本上是一个有自己输出的分支,然后是正常的、预训练的图像网络嵌入出来。

获取 MobileNet V2 作为 base_model:

  base_model = tf.keras.applications.MobileNetV2(input_shape=(IMG_SIZE, IMG_SIZE, 3),
                                                include_top=False,
                                                weights='imagenet')

  base_model.trainable = False

从 base_model 制作我的图层并制作我的新输出。

  # get layers from mobilenet base layer
  mobilenet_input = base_model.get_layer('input_1')
  mobilenet_output = base_model.get_layer('out_relu')

  # add our average pooling layer to our MobileNetV2 output like all of our other classifiers so we split our graph on the same nodes
  out_global_pooling = tf.keras.layers.GlobalAveragePooling2D(name='embedding_pooling')(mobilenet_output.output)
  out_global_pooling.trainable = False

  # Our new branch and outputs for the branch
  expanded_conv_depthwise_BN = base_model.get_layer('expanded_conv_depthwise_BN')
  regression_dropout = tf.keras.layers.Dropout(0.5) (expanded_conv_depthwise_BN.output)
  regression_global_pooling = tf.keras.layers.GlobalAveragePooling2D(name="regression_pooling")(regression_dropout)
  new_regression_output = tf.keras.layers.Dense(num_labels, activation = 'sigmoid', name = "cinemanet_output") (regression_global_pooling)

这看起来不错,我什至可以通过函数式 API 制作模型:

  model = tf.keras.Model(inputs=mobilenet_input.input, outputs=[out_global_pooling, new_regression_output])

我的培训代码

我的数据集是一组 30 个浮点数(10 个 RGB duplets),我想从输入图像中进行预测。我的数据集在训练“序列”模型时起作用,但在我尝试训练此模型时失败。

 ops.reset_default_graph()
  tf.keras.backend.set_learning_phase(1) # 0 testing, 1 training mode

# preview contents of CSV to verify things are sane
  import csv
  import math

  def lenopenreadlines(filename):
      with open(filename) as f:
          return len(f.readlines())

  def csvheaderrow(filename):
    with open(filename) as f:
      reader = csv.reader(f)
      return next(reader, None)

  # !head {label_file}

  NUM_IMAGES = ( lenopenreadlines(label_file) - 1) # remove header

  COLUMN_NAMES = csvheaderrow(label_file)
  
  LABEL_NAMES = COLUMN_NAMES[:]
  LABEL_NAMES.remove("filepath")

  ALL_LABELS.extend(LABEL_NAMES)

  # make our data set
  BATCH_SIZE = 256
  NUM_EPOCHS = 50
  FILE_PATH = ["filepath"]
  
  LABELS_TO_PRINT = ' '.join(LABEL_NAMES)
  print("Label contains: " + str(NUM_IMAGES) + " images")
  print("Label Are: " + LABELS_TO_PRINT)
  print("Creating Data Set From " + label_file)

  csv_dataset = get_dataset(label_file, BATCH_SIZE, NUM_EPOCHS, COLUMN_NAMES)

  #make a new data set from our csv by mapping every value to the above function
  split_dataset = csv_dataset.map(split_csv_to_path_and_labels)  

  # make a new datas set that loads our images from the first path 
  image_and_labels_ds = split_dataset.map(load_and_preprocess_image_batch, num_parallel_calls=AUTOTUNE)

  # update our image floating point range to match -1, 1
  ds = image_and_labels_ds.map(change_range)
  
  print(image_and_labels_ds)

  model = build_model(LABEL_NAMES, use_masked_loss)

  #split the final data set into train / validation splits to use for our model.
  DATASET_SIZE = NUM_IMAGES

  ds = ds.repeat()


  steps_per_epoch =  int(math.floor(DATASET_SIZE/BATCH_SIZE))
  history = model.fit(ds, epochs=NUM_EPOCHS, steps_per_epoch=steps_per_epoch, callbacks=[TensorBoardColabCallback(tbc)])


  print(history)

  # results = model.evaluate(test_dataset)
  # print('test loss, test acc:', results)
  export_model(model, model_name, LABEL_NAMES, date)

ValueError: Error when checking model target: 
the list of Numpy arrays that you are passing to your model is not the size the model expected.

Expected to see 2 array(s), but instead got the following list of 1 arrays:
[<tf.Tensor 'IteratorGetNext:1' shape=(?, 30) dtype=float32>]

如果我改为使用序列并天真地尝试针对移动网络的最终输出(而不是分支)训练我的回归任务 - 训练效果很好(尽管我得到的结果很差)。

我的模型摘要似乎告诉我,事情如我所料。我的辍学连接到expanded_conv_depthwise_BN。我的回归池连接到我的 dropout,并且我的输出层出现在连接到我的回归池的摘要中


Model: "model"
__________________________________________________________________________________________________
Layer (type)                    Output Shape         Param #     Connected to                     
==================================================================================================
input_1 (InputLayer)            [(None, 224, 224, 3) 0                                            
__________________________________________________________________________________________________
Conv1_pad (ZeroPadding2D)       (None, 225, 225, 3)  0           input_1[0][0]                    
__________________________________________________________________________________________________
Conv1 (Conv2D)                  (None, 112, 112, 32) 864         Conv1_pad[0][0]                  
__________________________________________________________________________________________________
bn_Conv1 (BatchNormalization)   (None, 112, 112, 32) 128         Conv1[0][0]                      
__________________________________________________________________________________________________
Conv1_relu (ReLU)               (None, 112, 112, 32) 0           bn_Conv1[0][0]                   
__________________________________________________________________________________________________
expanded_conv_depthwise (Depthw (None, 112, 112, 32) 288         Conv1_relu[0][0]                 
__________________________________________________________________________________________________
expanded_conv_depthwise_BN (Bat (None, 112, 112, 32) 128         expanded_conv_depthwise[0][0]    
__________________________________________________________________________________________________
expanded_conv_depthwise_relu (R (None, 112, 112, 32) 0           expanded_conv_depthwise_BN[0][0] 
__________________________________________________________________________________________________
expanded_conv_project (Conv2D)  (None, 112, 112, 16) 512         expanded_conv_depthwise_relu[0][0
__________________________________________________________________________________________


< snip for brevity >

________
block_16_project (Conv2D)       (None, 7, 7, 320)    307200      block_16_depthwise_relu[0][0]    
__________________________________________________________________________________________________
block_16_project_BN (BatchNorma (None, 7, 7, 320)    1280        block_16_project[0][0]           
__________________________________________________________________________________________________
Conv_1 (Conv2D)                 (None, 7, 7, 1280)   409600      block_16_project_BN[0][0]        
__________________________________________________________________________________________________
Conv_1_bn (BatchNormalization)  (None, 7, 7, 1280)   5120        Conv_1[0][0]                     
__________________________________________________________________________________________________
dropout (Dropout)               (None, 112, 112, 32) 0           expanded_conv_depthwise_BN[0][0] 
__________________________________________________________________________________________________
out_relu (ReLU)                 (None, 7, 7, 1280)   0           Conv_1_bn[0][0]                  
__________________________________________________________________________________________________
regression_pooling (GlobalAvera (None, 32)           0           dropout[0][0]                    
__________________________________________________________________________________________________
embedding_pooling (GlobalAverag (None, 1280)         0           out_relu[0][0]                   
__________________________________________________________________________________________________
cinemanet_output (Dense)        (None, 30)           990         regression_pooling[0][0]         
==================================================================================================
Total params: 2,258,974
Trainable params: 990
Non-trainable params: 2,257,984

【问题讨论】:

  • 能贴出训练代码吗?
  • 当然!生病编辑主要问题。
  • 所以这很有趣:如果我删除我的第一个输出(我不想训练的输出),一切正常。看来我的数据集需要为每个输出输出 2 个张量,即使我只想训练一个。
  • stackoverflow.com/questions/42785433/… 似乎是要做的事情,因为我想要整个网络(如果我只指定线性回归网络,则后面的层不包括在我想要的移动网络中)。

标签: tensorflow keras neural-network google-colaboratory


【解决方案1】:

看起来您的设置正确,但您的训练数据集不包含两个输出的张量。如果您只想训练新的输出,您可以为另一个提供虚拟张量(甚至是真实的训练数据),同时使用 0 的损失权重来防止参数更新。这也应该防止任何不是新输出层直接“上游”的参数在训练期间更新。

编译模型时,使用参数loss_weights 将权重作为列表(例如loss_weights=[0, 1])或字典(例如、@ 987654324@).

【讨论】:

  • 哦,这是一种很棒的处理方式。我最终做了一种稍微不同的方法,感觉更hacky,这在这个SO问题stackoverflow.com/questions/42785433/…中进行了概述,在那里我制作了两个模型,一个带有我想要训练的输出,另一个带有两个模型,最后保存了后者,因为它最终具有相同的层。不过,您的上述方法看起来要好得多。我会标记为答案并试一试,非常感谢!
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2017-05-25
  • 2019-03-21
  • 2018-05-16
  • 2021-01-05
  • 1970-01-01
  • 1970-01-01
  • 2021-03-21
相关资源
最近更新 更多