【问题标题】:My text classifier model doens't improve with multiple classes我的文本分类器模型没有改善多个类
【发布时间】:2020-03-13 10:31:03
【问题描述】:

我正在尝试训练一个文本分类模型,该模型采用从文章中嵌入的最多 300 个整数的列表。模型训练没有问题,除了准确率外,其他一切都不会上升。

目标由 41 个类别组成,从 0 到 41 编码为 int,然后进行归一化。

表格看起来像这样

另外,我不知道我的模型应该是什么样子,因为我参考了以下两个不同的示例

  • 一个输入列一输出列的二元分类器Example 1
  • 以多列作为输入的多类分类器Example 2

我已经尝试根据这两个模型修改我的模型,但模型精度不会改变,甚至每个 epoch 的精度都会降低

我应该为我的模型添加更多层还是我做了一些我没有意识到的愚蠢的事情?

注意:如果“df.pickle”下载链接损坏,请使用this link

from sklearn.model_selection import train_test_split
from urllib.request import urlopen
from os.path import exists
from os import mkdir
import tensorflow as tf
import pandas as pd
import pickle

# Define dataframe path
df_path = 'df.pickle'

# Check if local dataframe exists
if not exists(df_path):
  # Download binary from dropbox
  content = urlopen('https://ucd92a22d5e0d4d29b8edb608305.dl.dropboxusercontent.com/cd/0/get/Askx_25n3JI-jmnZsWXmMmRgd4O2EH1w9l0U6zCMq7xdSXs_IN_i2zuUviseqa9N7-WrReFbGhQi8CeseV5cNsFTO8dzRmSdxjr-MWEDQNpPaZ8Ik29E_58YAjY57qTc4CA/file#').read()

  # Write to file
  with open(df_path, 'wb') as file: file.write(content)

  # Load the dataframe from bytes
  df = pickle.loads(content)
# If the file exists (aka. downloaded)
else:
  # Load the dataframe from file
  df = pickle.load(open(df_path, 'rb'))

# Normalize the category
df['Category_Code'] = df['Category_Code'].apply(lambda x: x / 41)

train_df, test_df = [pd.DataFrame() for _ in range(2)]

x_train, x_test, y_train, y_test = train_test_split(df['Content_Parsed'], df['Category_Code'], test_size=0.15, random_state=8)
train_df['Content_Parsed'], train_df['Category_Code'] = x_train, y_train
test_df['Content_Parsed'], test_df['Category_Code'] = x_test, y_test

# Variable containing the number of words we want to keep in our vocabulary
NUM_WORDS = 10000
# Input/Token length
SEQ_LEN = 300

# Create tokenizer for our data
tokenizer = tf.keras.preprocessing.text.Tokenizer(num_words=NUM_WORDS, oov_token='<UNK>')
tokenizer.fit_on_texts(train_df['Content_Parsed'])

# Convert text data to numerical indexes
train_seqs=tokenizer.texts_to_sequences(train_df['Content_Parsed'])
test_seqs=tokenizer.texts_to_sequences(test_df['Content_Parsed'])

# Pad data up to SEQ_LEN (note that we truncate if there are more than SEQ_LEN tokens)
train_seqs=tf.keras.preprocessing.sequence.pad_sequences(train_seqs, maxlen=SEQ_LEN, padding="post")
test_seqs=tf.keras.preprocessing.sequence.pad_sequences(test_seqs, maxlen=SEQ_LEN, padding="post")

# Create Models folder if not exists
if not exists('Models'): mkdir('Models')

# Define local model path
model_path = 'Models/model.pickle'

# Check if model exists/pre-trained
if not exists(model_path):
  # Define word embedding size
  EMBEDDING_SIZE = 16

  # Create new model
  '''
  model = tf.keras.Sequential([
    tf.keras.layers.Embedding(NUM_WORDS, EMBEDDING_SIZE),
    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(EMBEDDING_SIZE)),
    # tf.keras.layers.Dense(EMBEDDING_SIZE, activation='relu'),
    tf.keras.layers.Dense(1, activation='sigmoid')
  ])
  '''
  model = tf.keras.Sequential([
      tf.keras.layers.Embedding(NUM_WORDS, EMBEDDING_SIZE),
      # tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(EMBEDDING_SIZE)),
      tf.keras.layers.GlobalAveragePooling1D(),
      tf.keras.layers.Dense(EMBEDDING_SIZE, activation='relu'),
      tf.keras.layers.Dense(1, activation='sigmoid')
  ])

  # Compile the model
  model.compile(
    optimizer='adam',
    loss='binary_crossentropy',
    metrics=['accuracy']
  )

  # Stop training when a monitored quantity has stopped improving.
  es = tf.keras.callbacks.EarlyStopping(monitor='val_acc', mode='max', patience=1)

  # Define batch size (Can be tuned to improve model accuracy)
  BATCH_SIZE = 16
  # Define number or cycle to train
  EPOCHS = 20

  # Using GPU (If error means you don't have GPU. Use CPU instead)
  with tf.device('/GPU:0'):
    # Train/Fit the model
    history = model.fit(
      train_seqs, 
      train_df['Category_Code'].values, 
      batch_size=BATCH_SIZE, 
      epochs=EPOCHS, 
      validation_split=0.2,
      validation_steps=30,
      callbacks=[es]
    )

  # Evaluate the model
  model.evaluate(test_seqs, test_df['Category_Code'].values)

  # Save the model into a file
  with open(model_path, 'wb') as file: file.write(pickle.dumps(model))
else:
  # Load the model
  model = pickle.load(open(model_path, 'rb'))

# Check the model
model.summary()

【问题讨论】:

    标签: python pandas tensorflow machine-learning text-classification


    【解决方案1】:

    经过 2 天的调整和理解更多示例后,我发现 this 网站很好地解释了多类分类。

    我所做的更改的详细信息包括:

    1. 由于我要为多类构建模型,在模型编译期间,模型应该使用categorical_crossentropy,因为它是损失函数而不是binary_crossentropy

    2. 模型应该产生与您的总类相似长度的输出数量,在我的例子中41

    3. 另外,在模型的最后一层你只需要一个激活函数,损失函数应该是"softmax"

    4. 您需要根据要分类的类数相应地调整图层。请参阅here,了解如何改进您的模型。

    最终代码如下所示

    from sklearn.model_selection import train_test_split
    from urllib.request import urlopen
    from functools import reduce
    from os.path import exists
    from os import listdir
    from sys import exit
    import tensorflow as tf
    import pandas as pd
    import pickle
    import re
    
    # Specify dataframe path
    df_path = 'df.pickle'
    # Check if the file exists
    if not exists(df_path):
      # Specify url of the dataframe binary
      url = 'https://www.dropbox.com/s/76hibe24hmpz3bk/df.pickle?dl=1'
      # Read the byte content from url
      content = urlopen(url).read()
      # Write to a file to save up time
      with open(df_path, 'wb') as file: file.write(pickle.dumps(content))
      # Unpickle the dataframe
      df = pickle.loads(content)
    else:
      # Load the pickle dataframe
      df = pickle.load(open(df_path, 'rb'))
    
    # Useful variables
    MAX_NUM_WORDS = 50000                        # Vocabulary size for our tokenizer
    MAX_SEQ_LENGTH = 600                         # Maximum length of tokens (for padding later)
    EMBEDDING_SIZE = 256                         # Embedding size (Tweak to improve accuracy)
    OUTPUT_LENGTH = len(df['Category'].unique()) # Number of class to be classified
    
    # Create our tokenizer
    tokenizer = tf.keras.preprocessing.text.Tokenizer(num_words=MAX_NUM_WORDS, lower=True)
    # Fit our tokenizer with words/tokens
    tokenizer.fit_on_texts(df['Content_Parsed'].values)
    # Get our token vocabulary
    word_index = tokenizer.word_index
    print('Found {} unique tokens'.format(len(word_index)))
    
    # Parse our text into sequence of numbers using our tokenizer
    X = tokenizer.texts_to_sequences(df['Content_Parsed'].values)
    # Pad the sequence up to the MAX_SEQ_LENGTH
    X = tf.keras.preprocessing.sequence.pad_sequences(X, maxlen=MAX_SEQ_LENGTH)
    print('Shape of feature tensor: {}'.format(X.shape))
    
    # Convert our labels into dummy variable (More info on the link provided above)
    Y = pd.get_dummies(df['Category']).values
    print('Shape of label tensor: {}'.format(Y.shape))
    
    # Split our features and labels into test and train dataset
    x_train, x_test, y_train, y_test = train_test_split(X, Y, test_size=0.1, random_state=42)
    print(x_train.shape, y_train.shape)
    print(x_test.shape, y_test.shape)
    
    # Creating our model
    model = tf.keras.Sequential()
    model.add(tf.keras.layers.Embedding(MAX_NUM_WORDS, EMBEDDING_SIZE, input_length=MAX_SEQ_LENGTH))
    model.add(tf.keras.layers.SpatialDropout1D(0.2))
    # The number 64 could be changed based on your model performance
    model.add(tf.keras.layers.LSTM(64, dropout=0.2, recurrent_dropout=0.2))
    # Our output layer with length similar to the OUTPUT_LENGTH
    model.add(tf.keras.layers.Dense(OUTPUT_LENGTH, activation='softmax'))
    # Compile our model with "categorical_crossentropy" loss function
    model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
    
    # Model variables
    EPOCHS = 100                          # Number of cycle to run (The early stopping may stop the training process accordingly)
    BATCH_SIZE = 64                       # Batch size (Tweaking this may improve model performance a bit)
    checkpoint_path = 'model_checkpoints' # Checkpoint path of our model
    
    # Use GPU if available
    with tf.device('/GPU:0'):
      # Fit/Train our model
      history = model.fit(
        x_train, y_train,
        epochs=EPOCHS,
        batch_size=BATCH_SIZE,
        validation_split=0.1,
        callbacks=[
          tf.keras.callbacks.EarlyStopping(monitor='val_loss', min_delta=0.0001),
          tf.keras.callbacks.ModelCheckpoint(
            checkpoint_path, 
            monitor='val_acc', 
            save_best_only=True, 
            save_weights_only=False
          )
        ],
        verbose=1
      )
    

    现在,我的模型准确度表现良好,并且每个 epoch 都有所提高,但由于验证准确度(76%~77% 的 val_acc)表现不佳,我可能需要稍微调整模型。

    下面提供了输出快照

    【讨论】:

      猜你喜欢
      • 2022-07-10
      • 2021-03-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-09-04
      • 2016-06-14
      • 1970-01-01
      • 2021-01-03
      相关资源
      最近更新 更多