【问题标题】:Bazel error parsing tf.estimator modelBazel 错误解析 tf.estimator 模型
【发布时间】:2018-08-14 05:39:52
【问题描述】:

我正在尝试使用tf.estimatorexport_savedmodel() 制作一个*.pb 模型,它是一个简单的分类器来分类虹膜数据集(4 个特征,3 个类):

import tensorflow as tf


num_epoch = 500
num_train = 120
num_test = 30

# 1 Define input function
def input_function(x, y, is_train):
    dict_x = {
        "thisisinput" : x,
    }

    dataset = tf.data.Dataset.from_tensor_slices((
        dict_x, y
    ))

    if is_train:
        dataset = dataset.shuffle(num_train).repeat(num_epoch).batch(num_train)
    else:   
        dataset = dataset.batch(num_test)

    return dataset


def my_serving_input_fn():
    input_data = tf.placeholder(tf.string, [None], name='input_tensors')
    receiver_tensors = {"inputs" : input_data}

    # 2 Define feature columns
    feature_columns = [
        tf.feature_column.numeric_column(key="thisisinput", shape=4),]
    features = tf.parse_example(
        input_data, 
        tf.feature_column.make_parse_example_spec(feature_columns))

    return tf.estimator.export.ServingInputReceiver(features, receiver_tensors)


def main(argv):
    tf.set_random_seed(1103) # avoiding different result of random

    # 2 Define feature columns
    feature_columns = [
        tf.feature_column.numeric_column(key="thisisinput", shape=4),]

    # 3 Define an estimator
    classifier = tf.estimator.DNNClassifier(
        feature_columns=feature_columns,
        hidden_units=[10],
        n_classes=3,
        optimizer=tf.train.GradientDescentOptimizer(0.001),
        activation_fn=tf.nn.relu,
        model_dir = 'modeliris2/'
    )

    # Train the model
    classifier.train(
        input_fn=lambda:input_function(xtrain, ytrain, True)
    )

    # Evaluate the model
    eval_result = classifier.evaluate(
        input_fn=lambda:input_function(xtest, ytest, False)
    )

    print('\nTest set accuracy: {accuracy:0.3f}\n'.format(**eval_result))
    print('\nSaving models...')
    classifier.export_savedmodel("modeliris2pb", my_serving_input_fn)


if __name__ == "__main__":
    tf.logging.set_verbosity(tf.logging.INFO)
    tf.app.run(main)

这将产生一个saved_model.pb 文件。我已经确认该模型有效。我还可以制作另一个程序来加载和运行它。现在,我想用 Bazel 来总结和冻结模型。如果我构建 Bazel 然后运行以下命令:

bazel-bin/tensorflow/tools/graph_transforms/summarize_graph \
--in_graph=saved_model.pb

我收到以下错误:

[libprotobuf ERROR external/protobuf_archive/src/google/protobuf/text_format.cc:307] 解析文本格式 tensorflow 时出错。GraphDef:1:1:文本中遇到无效的控制字符。
[libprotobuf ERROR external/protobuf_archive/src/google/protobuf/text_format.cc:307] 解析文本格式 tensorflow.GraphDef 时出错:1:4:解释非 ascii 代码点 218。
[libprotobuf ERROR external/protobuf_archive/src/google/protobuf/text_format.cc:307] 解析文本格式 tensorflow.GraphDef 时出错:1:4:预期标识符,得到:�
2018-08-14 11:50:17.759617:E tensorflow/tools/graph_transforms/summarize_graph_main.cc:320] 加载图“saved_model.pb”失败,无法将 saved_model.pb 解析为二进制原型
(文件 saved_model.pb 的文本和二进制解析均失败)
2018-08-14 11:50:17.759670:E tensorflow/tools/graph_transforms/summarize_graph_main.cc:322] 用法:bazel-bin/tensorflow/tools/graph_transforms/summarize_graph
标志:
--in_graph="" 字符串输入图形文件名
--print_structure=false bool 是否打印图的网络连接

我不明白这个错误。我尝试使用inception pb file,它运行良好,所以我认为问题在于tf.estimator 如何构建.pb 文件。

使用export_savedmodel()tf.estimator 创建保存的模型时是否遗漏了什么?

更新

Tensorflow 版本:v1.9.0-0-g25c197e023 1.9.0

tf_env_collect.sh 的结果:

== cat /etc/issue ===============================================
Linux rianadam 4.15.0-32-generic #35-Ubuntu SMP Fri Aug 10 17:58:07 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux
VERSION="18.04.1 LTS (Bionic Beaver)"
VERSION_ID="18.04"
VERSION_CODENAME=bionic

== are we in docker =============================================
No

== compiler =====================================================
c++ (Ubuntu 7.3.0-16ubuntu3) 7.3.0
Copyright (C) 2017 Free Software Foundation, Inc.
This is free software; see the source for copying conditions.  There is NO
warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.


== uname -a =====================================================
Linux rianadam 4.15.0-32-generic #35-Ubuntu SMP Fri Aug 10 17:58:07 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux

== check pips ===================================================
numpy               1.15.0 
protobuf            3.6.0  
tensorflow-gpu      1.9.0  

== check for virtualenv =========================================
True

== tensorflow import ============================================
tf.VERSION = 1.9.0
tf.GIT_VERSION = v1.9.0-0-g25c197e023
tf.COMPILER_VERSION = v1.9.0-0-g25c197e023
Sanity check: array([1], dtype=int32)
/home/rian/NgodingYuk/tf_env/env/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary incompatibility. Expected 96, got 88
return f(*args, **kwds)
/home/rian/NgodingYuk/tf_env/env/lib/python3.6/importlib/_bootstrap.py:219: RuntimeWarning: numpy.dtype size changed, may indicate binary incompatibility. Expected 96, got 88
return f(*args, **kwds)

== env ==========================================================
LD_LIBRARY_PATH /usr/local/cuda/lib64:/usr/local/cuda-9.0/lib64:/usr/local/cuda/lib64:/usr/local/cuda-9.0/lib64:
DYLD_LIBRARY_PATH is unset

== nvidia-smi ===================================================
Tue Aug 21 11:13:55 2018       
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 390.77                 Driver Version: 390.77                    |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|===============================+======================+======================|
|   0  GeForce 920M        Off  | 00000000:04:00.0 N/A |                  N/A |
| N/A   51C    P0    N/A /  N/A |    367MiB /  2004MiB |     N/A      Default |
+-------------------------------+----------------------+----------------------+

+-----------------------------------------------------------------------------+
| Processes:                                                       GPU Memory |
|  GPU       PID   Type   Process name                             Usage      |
|=============================================================================|
|    0                    Not Supported                                       |
+-----------------------------------------------------------------------------+

== cuda libs  ===================================================
/usr/local/cuda-9.0/lib64/libcudart_static.a
/usr/local/cuda-9.0/lib64/libcudart.so.9.0.176
/usr/local/cuda-9.0/doc/man/man7/libcudart.7
/usr/local/cuda-9.0/doc/man/man7/libcudart.so.7

【问题讨论】:

  • 您是否针对此问题向 tensorflow 提出过问题?
  • 我不太确定这是否是一个错误,在 Tensorflow 问题的说明中据说先在这里询问:/

标签: python tensorflow bazel tensorflow-estimator


【解决方案1】:

当我尝试从使用custom tf.Estimator 训练的模型中查找输入/输出节点时遇到了同样的问题。错误是因为您在使用 export_savedmodel 时得到的输出是 servable (据我所知,这是一个 GraphDef和其他元数据),而不仅仅是GraphDef

要找到输入和输出节点,你可以这样做。

# -*- coding: utf-8 -*-

import tensorflow as tf
from tensorflow.saved_model import tag_constants

with tf.Session(graph=tf.Graph()) as sess:
    gf = tf.saved_model.loader.load(
        sess,
        [tf.saved_model.tag_constants.SERVING],
        "/path/to/saved/model/")

    nodes = gf.graph_def.node
    print([n.name + " -> " + n.op for n in nodes
           if n.op in ('Softmax', 'Placeholder')])

    # ... ['Placeholder -> Placeholder',
    #      'dnn/head/predictions/probabilities -> Softmax']

我也使用了罐装的 DNNEstimator,所以 OP 的节点应该和我的一样,其他用户,您的操作名称可能与 PlaceholderSoftmax 不同,具体取决于您的分类器。

现在您有了输入/输出节点的名称,您可以冻结图形,地址为here

如果您想使用已训练参数的值,例如量化权重,您需要运行 tensorflow/python/tools/freeze_graph.py 将检查点值转换为图形文件本身中的嵌入常量.

#!/bin/bash

python ./freeze_graph.py \
  --in_graph="/path/to/model/saved_model.pb" \
  --input_checkpoint="/MyModel/model.ckpt-xxxx" \
  --output_graph="/home/user/pruned_saved_model_or_whatever.pb" \
  --input_saved_model_dir="/path/to/model" \
  --output_node_names="dnn/head/predictions/probabilities" \

那么假设你已经建立了graph_transforms

#!/bin/bash

tensorflow/bazel-bin/tensorflow/tools/graph_transforms/summarize_graph \
  --in_graph=pruned_saved_model_or_whatever.pb

输出

Found 1 possible inputs: (name=Placeholder, type=string(7), shape=[?])
No variables spotted.
Found 1 possible outputs: (name=dnn/head/predictions/probabilities, op=Softmax)
Found 256974297 (256.97M) const parameters, 0 (0) variable parameters, and 0 
control_edges
Op types used: 155 Const, 41 Identity, 32 RegexReplace, 18 Gather, 9 
StridedSlice, 9 MatMul, 6 Shape, 6 Reshape, 6 Relu, 5 ConcatV2, 4 BiasAdd, 4 
Add, 3 ExpandDims, 3 Pack, 2 NotEqual, 2 Where, 2 Select, 2 StringJoin, 2 Cast, 
2 DynamicPartition, 2 Fill, 2 Maximum, 1 Size, 1 Unique, 1 Tanh, 1 Sum, 1 
StringToHashBucketFast, 1 StringSplit, 1 Equal, 1 Squeeze, 1 Square, 1 
SparseToDense, 1 SparseSegmentSqrtN, 1 SparseFillEmptyRows, 1 Softmax, 1 
FloorDiv, 1 Rsqrt, 1 FloorMod, 1 HashTableV2, 1 LookupTableFindV2, 1 Range, 1 
Prod, 1 Placeholder, 1 ParallelDynamicStitch, 1 LookupTableSizeV2, 1 Max, 1 Mul
To use with tensorflow/tools/benchmark:benchmark_model try these arguments:
bazel run tensorflow/tools/benchmark:benchmark_model -- -- 
graph=pruned_saved_model.pb --show_flops --input_layer=Placeholder -- 
input_layer_type=string --input_layer_shape=-1 -- 
output_layer=dnn/head/predictions/probabilities

希望这会有所帮助。

更新(2018-12-03)

我打开的一个相关github issue 似乎已在工单末尾列出的详细博客文章中解决。

【讨论】:

  • 哇,谢谢你,你的代码帮助我找到输入和输出节点,并且现在使用你的一步一步总结图形工作:)
  • 我刚刚意识到我的 pruned_saved_model_or_whatever.pb 无法使用tensorflow.contrib.from_saved_model 打开。错误说MetaGraphDef associated with tags 'serve' could not be found in SavedModel,你知道为什么会这样吗?
  • 冻结图表后,它就不再是saved model
猜你喜欢
  • 2018-06-05
  • 2019-01-05
  • 2018-01-27
  • 1970-01-01
  • 1970-01-01
  • 2017-05-12
  • 2016-04-21
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多