【问题标题】:Profiling TensorFlow using tfprof使用 tfprof 分析 TensorFlow
【发布时间】:2017-02-17 23:33:09
【问题描述】:

我正在尝试分析 TensorFlow 的计算/内存使用情况,并发现 tfprof 是适合我目的的正确工具。但是,我无法获得所有运营商的 FLOPS。

这是我在 TensorFlow 存储库 (tensorflow/models/image/cifar10/cifar10_train.py) 中使用 cifar10 教程的 tfprof 教程之后所做的:

run_metadata = tf.RunMetadata()

_, loss_value = sess.run([train_op, loss],
        options=tf.RunOptions(trace_level=tf.RunOptions.FULL_TRACE),
        run_metadata=run_metadata)

op_log = tfprof_log_pb2.OpLog()

// TODO: add op information

tf.contrib.tfprof.tfprof_logger.write_op_log(
        tf.get_default_graph(),
        log_dir="/tmp/log_dir",
        op_log=op_log,
        run_meta=run_metadata)

tf.contrib.tfprof.model_analyzer.print_model_analysis(
        tf.get_default_graph(),
        run_metadata=run_metadata,
        op_log=op_log,
        tfprof_options=tf.contrib.tfprof.model_analyzer.FLOAT_OPS_OPTIONS)

结果是

Parsing GraphDef...
Parsing RunMetadata...
Parsing OpLog...
Preparing Views...

=========================Options=============================
-max_depth                  10000
-min_bytes                  0
-min_micros                 0
-min_params                 0
-min_float_ops              1
-device_regexes             .*
-order_by                   float_ops
-account_type_regexes       .*
-start_name_regexes         .*
-trim_name_regexes
-show_name_regexes          .*
-hide_name_regexes
-account_displayed_op_only  true
-select                     float_ops
-viz                        false
-dump_to_file

==================Model Analysis Report======================
_TFProfRoot (0/5.23b flops)
  conv2/Conv2D (3.77b/3.77b flops)
  conv1/Conv2D (707.79m/707.79m flops)
  gradients/local3/MatMul_grad/MatMul (226.49m/226.49m flops)
  gradients/local3/MatMul_grad/MatMul_1 (226.49m/226.49m flops)
  local3/MatMul (226.49m/226.49m flops)
  gradients/local4/MatMul_grad/MatMul (18.87m/18.87m flops)
  gradients/local4/MatMul_grad/MatMul_1 (18.87m/18.87m flops)
  local4/MatMul (18.87m/18.87m flops)
  conv1/BiasAdd (4.72m/4.72m flops)
  conv2/BiasAdd (1.18m/1.18m flops)
  gradients/softmax_linear/MatMul_grad/MatMul (491.52k/491.52k flops)
  gradients/softmax_linear/MatMul_grad/MatMul_1 (491.52k/491.52k flops)
  softmax_linear/MatMul (491.52k/491.52k flops)

======================End of Report==========================

但是,结果并不包含所有的操作,例如最大池化、relu、conv 层的梯度。也许这些操作的失败统计数据没有定义(RegisterStatistics('flops'))。因此,为了提供运行时信息,如 tfprof 教程 11),我尝试创建 OpLog(参见上面的代码)。

但是,我不确定如何添加操作信息(如何获取操作的条目名称?)。有没有办法添加它包含的 ALL 操作?

或者任何其他工具而不是 tfprof?也许来自 NVIDIA 的分析工具?

【问题讨论】:

  • tfprof 链接已损坏。由于编辑队列已满,这里是working link。此外,tfprof 已被弃用。

标签: tensorflow profiling gpu


【解决方案1】:

您是对的,其他操作在没有 RegisterStatistics('flops') 之前没有 flops。欢迎您投稿。

我不确定 NVIDA 是否有相应的工具。

【讨论】:

  • 对于那些操作,即使是 OpLog 和 runtime_meta 也不提供 flops stat?
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-11-07
  • 1970-01-01
  • 1970-01-01
  • 2021-10-06
相关资源
最近更新 更多