【问题标题】:How to build embedding layer with tensorflow for the following model?如何使用 tensorflow 为以下模型构建嵌入层?
【发布时间】:2019-03-07 23:16:21
【问题描述】:

这是一个用于情感分析的 keras 模型我需要将其转换为 tensorflow 我无法使用 tensorflow 构建嵌入层并使用混淆矩阵来评估此模型?我问过 tf-learn 是否和 tensorflow 一样

import os     
import numpy as np  
import pandas as pd  
import tensorflow as tf  
from tensorflow import set_random_seed
set_random_seed(2)
from nltk.tokenize import word_tokenize
from sklearn.utils import shuffle
from sklearn.model_selection import train_test_split
from sklearn.metrics import confusion_matrix
from keras.preprocessing.sequence import pad_sequences
from keras.preprocessing.text import Tokenizer
from sklearn.preprocessing import  LabelEncoder
from keras.utils import to_categorical
from keras.models import Sequential
from keras.layers.embeddings import Embedding
from keras.layers import Flatten
from keras.layers import Conv1D, MaxPooling1D
from keras.layers import Dense,Activation
from keras.layers import Dropout
from keras.callbacks import TensorBoard, ModelCheckpoint
import re
import string
import collections
import time
seed = 10

读取 CSV 文件

df=pd.read_csv('tweets-pos-neg.csv', usecols = ['text','airline_sentiment'])
df = df.reindex(['text','airline_sentiment'], axis=1) #reorder columns
df=df.apply(lambda x: x.astype(str).str.lower()) 

规范化文本

def normalize(text):
    text= re.sub(r"http\S+", r'', text) 
    text= re.sub(r"@\S+", r'', text)
    punctuation = re.compile(r'[!"#$%&()*+,-./:;<=>?@[\]^_`{|}~|0-9]')
    text = re.sub(punctuation, ' ', text)
    text= re.sub(r'(.)\1\1+', r'\1', text) 
    return text

清除文本

def prepareDataSets(df):
     sentences=[]
     for index, r in df.iterrows():
         text= normalize(r['text'])
         sentences.append([text,r['airline_sentiment']])
          df_sentences=pd.DataFrame(sentences,columns= 
          ['text','airline_sentiment'])
     return df_sentences
edit_df=prepareDataSets(df)
edit_df=shuffle(edit_df)
X=edit_df.iloc[:,0]
Y=edit_df.iloc[:,1]

将评论拆分为标记

 max_features = 50000
 tokenizer = Tokenizer(num_words=max_features, split=' ')
 tokenizer.fit_on_texts(X.values)
 #convert review tokens to integers
 X_seq = tokenizer.texts_to_sequences(X)

根据评论的最大长度使所有向量具有相同大小的填充序列

seq_len=35
X_pad = pad_sequences(X_seq,maxlen=seq_len)   

将目标值从字符串转换为整数

le=LabelEncoder()
Y_le=le.fit_transform(Y)
Y_le_oh=to_categorical(Y_le)

训练-测试-拆分

X_train, X_test, Y_train, Y_test = train_test_split(X_pad,Y_le_oh, test_size 
= 0.33, random_state = 42)
X_train, X_Val, Y_train, Y_Val = train_test_split(X_train,Y_train, test_size 
= 0.1, random_state = 42)
print(X_train.shape,Y_train.shape)
print(X_test.shape,Y_test.shape)
print(X_Val.shape,Y_Val.shape) 

创建模型

embedding_vecor_length = 32    #no of vector columns
model_cnn = Sequential()
model_cnn.add(Embedding(max_features, embedding_vecor_length, 
input_length=seq_len))
model_cnn.add(Conv1D(filters=100, kernel_size=2, padding='valid', 
activation='relu', strides=1))
model_cnn.add(MaxPooling1D(2))
model_cnn.add(Flatten())
model_cnn.add(Dense(256, activation='relu'))
model_cnn.add(Dense(2, activation='softmax'))
opt=tf.keras.optimizers.Adam(lr=0.001, decay=1e-6)
model_cnn.compile(loss='binary_crossentropy', optimizer=opt, metrics= 
['accuracy'])
print(model_cnn.summary())  

评估模型

history=model_cnn.fit(X_train, Y_train, epochs=3, batch_size=32, callbacks=[tensorboard], validation_data=(X_Val, Y_Val))
scores = model_cnn.evaluate(X_test, Y_test, verbose=0)
print("Accuracy: %.2f%%" % (scores[-1]*100))

【问题讨论】:

  • 除了将所有内容都转换为有点宽泛的 Tensorflow 并且您可能必须自己做之外,到目前为止您尝试了哪些方法,您遇到了什么问题?
  • @de1 在张量流中嵌入层我无法创建它我搜索了更多但我无法达到它的形状,我问这些步骤在张量流中是否也正确
  • 有嵌入的 word2vec 示例:tensorflow.org/tutorials/representation/word2vec 我认为 Stackoverflow 不是为一般代码审查而设计的。 codereview.stackexchange.com 可能会更好。
  • 澄清一下,您是否需要将model_cnn 从 keras 转换为 tensorflow(例如在生产中使用),还是需要将所有这些模型构建代码移植到 tensorflow? (顺便说一句,model_cnn.fit() 是构建模型的一部分,而不是评估它;compile() 只是在配置,实际上并没有在这一点上进行任何模型训练。)

标签: tensorflow keras deep-learning tensorboard sentiment-analysis


【解决方案1】:

如果您只需要使用 Tensorflow API 来训练/评估,您可以使用 model_to_estimator 函数构建一个 Estimator。

这是documentation 的示例。

【讨论】:

  • 不,我需要将此模型转换为 tensorflow 而不是 keras,我知道差异将从嵌入层开始,我无法使用 tensorflow 构建它..
猜你喜欢
  • 2019-02-09
  • 2020-01-23
  • 1970-01-01
  • 2020-10-14
  • 2017-04-18
  • 2018-09-29
  • 1970-01-01
  • 2021-09-26
  • 1970-01-01
相关资源
最近更新 更多