【问题标题】:What is attention penalty in speech transformer paper? (updated)语音转换器论文中的注意力惩罚是什么? (更新)
【发布时间】:2020-01-08 13:29:18
【问题描述】:

github:https://github.com/sephiroce/tfsr/tree/exprimental

我正在尝试重现语音转换器论文 [1] 中描述的识别准确度。 注意力惩罚是一种我无法完全理解的技术。 这是论文中对注意力惩罚的描述。

“此外,我们通过添加 对更远位置对的注意力权重的惩罚更大。”

我的理解是,除了解码器中的第一个多头注意力之外,这意味着在缩放的注意力逻辑上(在掩蔽之前)上添加更远离对角线的更小的负值。

这是用于计算注意力权重的代码 sn-p。

  # Q * trans(K): (..., seq_len_q, seq_len_k)
  matmul_qk = tf.matmul(query, key, transpose_b=True)

  # scaled matmul_qk: ( Q * trans(K) ) / sqrt(d_k)
  dimension_of_key = tf.cast(tf.shape(key)[-1], tf.float32)
  scaled_attention_logits = matmul_qk / tf.math.sqrt(dimension_of_key)

  # add the mask to the scaled tensor
  if mask is not None:
    scaled_attention_logits += (mask * -1e9)

  # softmax is normalized on the last axis (seq_len_k) so that the scores
  # add up to 1.
  attention_weights = tf.nn.softmax(scaled_attention_logits, axis=-1)

  # Adding penalty to attention weights and linearly re-normalize it.
  if attention_penalty is not None and att_penalty_scale > 0:
    attention_weights += (attention_penalty * att_penalty_scale)
    attention_weights += tf.math.abs(tf.math.reduce_min(attention_weights))
    inv_sum = 1 / tf.math.reduce_sum(attention_weights, axis=-1)
    attention_weights = tf.einsum('ijlm,ijl->ijlm', attention_weights, inv_sum)

下面的源代码 sn-p 用于创建注意力惩罚矩阵。 我找不到任何有效的方法来为解码器中的第二个多头注意力权重创建注意力惩罚矩阵,因为注意力图不是对角线。因此,首先我试图将注意力惩罚应用于编码器。 源代码为距离对角线更远的元素分配线性更大的惩罚。
有两个超参数,例如 attention_penalty_scale(这类似于 Jindřich 建议的penalty_values)和对角线的宽度。
我也许可以添加一个选项,例如stripe_step_size。目前stripe_step_size可以解释为1。

def create_attention_penalty(inp_len, tar_len, num_heads, attention_penalty_width):
  max_inp_len = tf.cast(tf.math.reduce_max(inp_len), tf.int32)
  n_batch = tf.shape(inp_len)[0]

  enc_att_penalty = tf.ones([n_batch, num_heads, max_inp_len, max_inp_len])

  accum = tf.zeros(([n_batch, num_heads, max_inp_len, max_inp_len]))
  for i in range(attention_penalty_width - 1, max_inp_len - 1):
    accum += tf.linalg.band_part(enc_att_penalty, i, i, name=None) - 1

  enc_att_penalty = accum

  return enc_att_penalty, None

即使我按照我的理解实现了,我也无法获得任何准确性改进。这个实现还有另一个缺点。训练速度越来越慢。

Q) 如何有效地将这种注意力惩罚方法应用于方形和非方形注意力权重?

参考
[1] Linhao Dong, Shuang Xu, Bo Xu, Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition, ICASSP 2018, https://ieeexplore.ieee.org/document/8462506

【问题讨论】:

    标签: tensorflow deep-learning speech-recognition tf.keras transformer


    【解决方案1】:

    我想你理解得很好。他们可能在对角线上做了一条条纹,比如:

    attention_penalty = (1 - tf.linalg.band_part(scaled_attention_logits, stripe_size, stripe_size)) * penalty
    

    但是,您可能需要更多地尝试 strip_sizepenalty_values 应该是什么,因为论文并没有说太多。或者您可以尝试写信给作者。

    【讨论】:

    • 非常感谢@Jindřich。我修改了源代码以在注意力权重上添加注意力惩罚,而不是注意力逻辑(也称为注意力能量)。我实现了“create_attention_penalty”方法。 (请检查更新的问题。)使用注意力惩罚掩蔽,即使我使用相同的注意力惩罚进行训练和解码,我也没有得到任何提高的准确性。我认为正如你所建议的,我需要找到一些好的超参数。
    猜你喜欢
    • 2019-10-18
    • 2011-07-30
    • 2017-10-08
    • 2016-12-29
    • 2014-12-25
    • 1970-01-01
    • 1970-01-01
    • 2017-10-29
    • 2014-05-31
    相关资源
    最近更新 更多