【问题标题】:Using Tensorflow Estimators with Dataset API results in strange steps behavior将 Tensorflow Estimators 与 Dataset API 一起使用会导致奇怪的步骤行为
【发布时间】:2021-04-10 08:43:09
【问题描述】:

我在 Tensorflow 的 Estimator 和 Dataset API 的训练循环行为方面遇到了一些问题。 代码如下(tf2.3):


NUM_EXAMPLES = X_train.shape[0] # dataset has 8000 elements
BATCH_SIZE = NUM_EXAMPLES
STEPS = NONE
N_EPOCHS = 100

def make_input_fn(X, y, n_epochs=N_EPOCHS, shuffle=True):
    dataset = tf.data.Dataset.from_tensor_slices((X.to_dict(orient='list'), y))
    if shuffle:
      dataset = dataset.shuffle(NUM_EXAMPLES)
    return dataset.repeat(n_epochs).batch(BATCH_SIZE)


estimator = tf.estimator.BoostedTreesClassifier(feature_cols, {
    'config': tf.estimator.RunConfig(
        model_dir=model_dir,
        save_checkpoints_steps=100
    ),
    'n_trees': 50,
    'max_depth': 6,
    'n_batches_per_layer': 1,
    'l2_regularization': 0.1
})

estimator.train(input_fn=lambda: make_input_fn(X_train, y_train), steps=STEPS)

我只是不明白我看到的行为。 TF 估计器训练的步数似乎被限制在 300 步,而不管我为 batch_size、训练步数或 epoch 数设置了什么。

我的数据集有 8K 个训练元素,当我选择 n_epochs=100batch_size=1000steps=None 时,我预计 tensorflow 将运行 100 (n_epochs) * 8 (steps required for 1 epoch) 步,但不,它运行 300 步。

下面实际上是对不同N_EPOCHSBATCH_SIZESTEPS的多个实验的总结,前3个对我来说很好,但其余的不行。

- steps N_EPOCHS BATCH_SIZE TF training steps (est.train) My expected # steps
1 None 100 8000 100 100
2 None 200 8000 200 200
3 None 300 8000 300 300
4 None 400 8000 300 400
5 None 100 1000 300 800
6 None 600 10 300 600 * 800
7 400 400 8000 300 400

可以看出,从第 4 行开始,我的期望不等于 tensorflow 训练运行的实际步数。这基本上意味着当我将batch_size 降低到说10 时,它只在 10 批错误的数据上运行 300 个 epoch,但我无法理解我的实现有什么不正确,查看文档任何帮助都非常感谢!

另外,我使用 train_and_evaluate 和 Specs 还是直接使用 train 都没有关系,为了简单起见,这里是 train。 train函数的日志如下(第7次实验):

INFO:tensorflow:Calling model_fn.
INFO:tensorflow:Done calling model_fn.
INFO:tensorflow:Create CheckpointSaverHook.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Graph was finalized.
INFO:tensorflow:Running local_init_op.
INFO:tensorflow:Done running local_init_op.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Calling checkpoint listeners before saving checkpoint 0...
INFO:tensorflow:Saving checkpoints for 0 into /tmp/estimator-run-1609770778/model.ckpt.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Calling checkpoint listeners after saving checkpoint 0...
INFO:tensorflow:loss = 0.69314593, step = 0
INFO:tensorflow:Calling checkpoint listeners before saving checkpoint 100...
INFO:tensorflow:Saving checkpoints for 100 into /tmp/estimator-run-1609770778/model.ckpt.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Calling checkpoint listeners after saving checkpoint 100...
INFO:tensorflow:global_step/sec: 2.88621
INFO:tensorflow:loss = 0.64561623, step = 99 (34.648 sec)
INFO:tensorflow:Calling checkpoint listeners before saving checkpoint 200...
INFO:tensorflow:Saving checkpoints for 200 into /tmp/estimator-run-1609770778/model.ckpt.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Calling checkpoint listeners after saving checkpoint 200...
INFO:tensorflow:global_step/sec: 2.9199
INFO:tensorflow:loss = 0.6292723, step = 199 (34.248 sec)
INFO:tensorflow:Calling checkpoint listeners before saving checkpoint 300...
INFO:tensorflow:Saving checkpoints for 300 into /tmp/estimator-run-1609770778/model.ckpt.
WARNING:tensorflow:Issue encountered when serializing resources.
Type is unsupported, or the types of the items don't match field type in CollectionDef. Note this is a warning and probably safe to ignore.
'_Resource' object has no attribute 'name'
INFO:tensorflow:Calling checkpoint listeners after saving checkpoint 300...
INFO:tensorflow:global_step/sec: 2.83013
INFO:tensorflow:loss = 0.6164282, step = 299 (35.334 sec)
INFO:tensorflow:Loss for final step: 0.6164282.

【问题讨论】:

    标签: tensorflow machine-learning tensorflow2.0 tensorflow-datasets tensorflow-estimator


    【解决方案1】:

    我认为这里的线索是n_trees * max_depth = 300:

        'n_trees': 50,
        'max_depth': 6,
    

    另外,看看this test中的那一行:

          # It will stop after 5 steps because of the max depth and num trees.
          num_steps = 100
    

    我不知道确切的逻辑,我猜它只是在每批一棵树上添加一层。你设置了树的数量和最大深度,一旦所有的树都建好了,它就不能继续训练了。

    【讨论】:

    • 有趣,这意味着应该将 batch_size 设置为完整数据集以实际传达 300 个 epoch 的训练?
    猜你喜欢
    • 2019-11-15
    • 2018-03-06
    • 1970-01-01
    • 2018-09-16
    • 2020-12-11
    • 2015-11-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多