【问题标题】:How to use gradient_override_map in Tensorflow 2.0?如何在 Tensorflow 2.0 中使用 gradient_override_map?
【发布时间】:2019-09-09 21:25:47
【问题描述】:

我正在尝试将 gradient_override_map 与 Tensorflow 2.0 一起使用。有一个example in the documentation,我也会在这里作为例子。

在 2.0 中,GradientTape 可用于计算梯度,如下所示:

import tensorflow as tf
print(tf.version.VERSION)  # 2.0.0-alpha0

x = tf.Variable(5.0)
with tf.GradientTape() as tape:
    s_1 = tf.square(x)
print(tape.gradient(s_1, x))

还有tf.custom_gradient 装饰器,可用于为new 函数定义渐变(同样,使用example from the docs):

import tensorflow as tf
print(tf.version.VERSION)  # 2.0.0-alpha

@tf.custom_gradient
def log1pexp(x):
    e = tf.exp(x)

    def grad(dy):
        return dy * (1 - 1 / (1 + e))

    return tf.math.log(1 + e), grad

x = tf.Variable(100.)

with tf.GradientTape() as tape:
    y = log1pexp(x)

print(tape.gradient(y, x))

但是,我想替换标准函数的渐变,例如tf.square。我尝试使用以下代码:

@tf.RegisterGradient("CustomSquare")
def _custom_square_grad(op, grad):
  return tf.constant(0)

with tf.Graph().as_default() as g:
    x = tf.Variable(5.0)
    with g.gradient_override_map({"Square": "CustomSquare"}):
        with tf.GradientTape() as tape:
            s_2 = tf.square(x, name="Square")

    with tf.compat.v1.Session() as sess:
        sess.run(tf.compat.v1.global_variables_initializer())            
        print(sess.run(tape.gradient(s_2, x)))

但是,有两个问题:梯度替换似乎不起作用(它被评估为10.0 而不是0.0),我需要求助于session.run() 来执行图表。有没有办法在“原生”TensorFlow 2.0 中实现这一点?

在 TensorFlow 1.12.0 中,以下内容会产生所需的输出:

import tensorflow as tf
print(tf.__version__)  # 1.12.0

@tf.RegisterGradient("CustomSquare")
def _custom_square_grad(op, grad):
  return tf.constant(0)

x = tf.Variable(5.0)

g = tf.get_default_graph()
with g.gradient_override_map({"Square": "CustomSquare"}):
    s_2 = tf.square(x, name="Square")
grad = tf.gradients(s_2, x)

with tf.Session() as sess:
  sess.run(tf.global_variables_initializer())
  print(sess.run(grad))

【问题讨论】:

    标签: python tensorflow tensorflow2.0


    【解决方案1】:

    TensorFlow 2.0 中没有内置机制来覆盖范围内内置运算符的所有梯度。但是,如果您能够修改对内置运算符的每次调用的调用站点,则可以使用 tf.custom_gradient 装饰器,如下所示:

    @tf.custom_gradient
    def custom_square(x):
      def grad(dy):
        return tf.constant(0.0)
      return tf.square(x), grad
    
    with tf.Graph().as_default() as g:
      x = tf.Variable(5.0)
      with tf.GradientTape() as tape:
        s_2 = custom_square(x)
    
      with tf.compat.v1.Session() as sess:
        sess.run(tf.compat.v1.global_variables_initializer())            
        print(sess.run(tape.gradient(s_2, x)))
    

    【讨论】:

    • 您是否知道tf.compat.v1.Session()/sess.run() 在可预见的未来是否仍将是 TensorFlow 的一部分?
    • tf.compat.v1 兼容性模块包含最新版本 TF 1.x 中 tf 模块中的所有内容(除了tf.contrib)。在可预见的将来,没有计划将其从 TensorFlow 中移除,因为许多库仍然依赖于它,尽管新功能开发将集中在主模块上,并且新旧 API 之间的兼容性可能存在差距(尽管,幸运的是,这种情况有效!)。
    • 不使用compat,它会消失的
    【解决方案2】:

    除了mrry的回答,还有两点想补充:

    (1) 在 TF 2 中,我们可以使用 tf.GradientTape 而不用构建图,像这样:

    @tf.custom_gradient
    def custom_square(x):
      def grad(dy):
        return tf.constant(0.0)
      return tf.square(x), grad
    
    with tf.GradientTape() as tape:
      x = tf.Variable(5.0)
      s_2 = custom_square(x)
    
    print(tape.gradient(s_2,x).numpy())
    

    (2) 将您的custom grad 与上一个毕业生相乘

    小心,梯度计算是一个链式计算,我们应该将自定义梯度乘以dy(之前计算的梯度)。 如果不这样做,我们的自定义函数将在链式计算中被破坏。这是一个例子:

    @tf.custom_gradient
    def custom_square(x):
      def grad(dy):
        return tf.constant(4.0)
      return tf.square(x), grad
    
    with tf.GradientTape(persistent=True) as tape:
      x = tf.Variable(5.0)
      s_2 = custom_square(x)
      s_4 = custom_square(s_2)
    
    print("Grad from s_4 to x: ",tape.gradient(s_4,x).numpy())
    print("Grad from s_4 to s_2: ",tape.gradient(s_4,s_2).numpy())
    print("Grad from s_2 to x: ",tape.gradient(s_2,x).numpy())
    

    结果:

    Grad from s_4 to x:  4.0
    Grad from s_4 to s_2:  4.0
    Grad from s_2 to x:  4.0
    

    s_4x 的毕业生应该是16(从s_4s_2 和从s_2x 的累计毕业生)。

    但结果是 4。这意味着它没有从上一步累积梯度。

    将自定义 grad 乘以dy将解决问题:

    @tf.custom_gradient
    def custom_square(x):
      def grad(dy):
        return tf.constant(4.0)*dy
      return tf.square(x), grad
    
    with tf.GradientTape(persistent=True) as tape:
      x = tf.Variable(5.0)
      s_2 = custom_square(x)
      s_4 = custom_square(s_2)
    
    print("Grad from s_4 to x: ",tape.gradient(s_4,x).numpy())
    print("Grad from s_4 to s_2: ",tape.gradient(s_4,s_2).numpy())
    print("Grad from s_2 to x: ",tape.gradient(s_2,x).numpy())
    

    结果如下:

    Grad from s_4 to x:  16.0
    Grad from s_4 to s_2:  4.0
    Grad from s_2 to x:  4.0
    

    您可以在这里通过 Colab 尝试实现:https://colab.research.google.com/drive/1gbLopOLJiyznDA-Cr473bZEeWkWh_KGG?usp=sharing

    【讨论】:

      猜你喜欢
      • 2017-05-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-02-02
      相关资源
      最近更新 更多