【发布时间】:2020-07-16 05:26:04
【问题描述】:
通常,转换器标记器将输入编码为字典。
{"input_ids": tf.int32, "attention_mask": tf.int32, "token_type_ids": tf.int32}
为了对大型数据集进行更好的性能处理,最好实现一个管道,其中包括使用Dataset.map 将标记器函数应用于输入数据集的每个元素。与 Tensorflow 教程中所做的完全相同:Load text。
但是,tf.py_function(用于包装 map python 函数)不支持返回如上所示的张量字典。
例如,如果Load text 中的分词器(编码器)返回以下字典:
{
"input_ids": [ 101, 13366, 2131, 1035, 6819, 2094, 1035, 102 ],
"attention_mask": [ 1, 1, 1, 1, 1, 1, 1, 1 ]
}
如何设置tf.py_function 的Tout 参数以获得所需的张量字典:
{
'input_ids': <tf.Tensor: shape=(16,), dtype=int32, numpy = array(
[ 101, 13366, 2131, 1035, 6819, 2094, 1035, 102 ], dtype=int32)>
'attention_mask': <tf.Tensor: shape=(16,), dtype=int32, numpy=array(
[ 1, 1, 1, 1, 1, 1, 1, 1 ], dtype=int32)>
}
?
【问题讨论】:
标签: python-3.x tensorflow2.0 huggingface-transformers