【问题标题】:How to print feature values after transformation inside tensorflow model如何在张量流模型内转换后打印特征值
【发布时间】:2019-07-01 09:50:00
【问题描述】:

如何查看在 TensorFlow 模型中训练的最终特征的价值。就像在下面的情况下,我试图对我的列“x”进行多重加热,我想看看这些功能是如何进入我的模型的。

这在 sklearn 中很容易做到,但对于 Tensorflow 来说是新手,我不明白这怎么可能。

import tensorflow as tf
import pandas as pd

data = {'x':['a c', 'a b', 'b c'], 'y': [1, 1, 0]}

df = pd.DataFrame(data)
Y = df['y']
X = df.drop('y', axis=1)
indicator_features = [tf.feature_column.indicator_column(categorical_column=
      tf.feature_column.categorical_column_with_vocabulary_list(key = 'x', 
                                                 vocabulary_list = ['a','b','c']))]
model = tf.estimator.LinearClassifier(feature_columns=indicator_features,
                                                              model_dir = "/tmp/samplemodel")
training_input_fn = tf.estimator.inputs.pandas_input_fn(x = X,
                                                    y=Y,
                                                    batch_size=64,
                                                    shuffle= True,
                                                    num_epochs = None)

model.train(input_fn=training_input_fn,steps=1000)

【问题讨论】:

  • 你能在 sklearn 中发布你将如何做,以便我可以尝试复制它吗?
  • 你试过 tf.Print() 吗?
  • @Sharky...我不明白如何在这里使用 tf.Print().. 有什么想法吗?
  • @gorjan.. 我可以简单地做类似这个例子的事情......chrisalbon.com/machine_learning/preprocessing_structured_data/…
  • @KundanKumar,您可以将任何张量传递给它。我不确定它是否可以与 tf.estimator.inputs.pandas_input_fn 一起使用,但是您可以尝试像 tf.Print(x, [x]) 一样将“x”传递给它

标签: tensorflow machine-learning data-science tensorboard


【解决方案1】:

我已经能够通过在 tensorflow 中启用即时执行来打印这些值。 在下面发布我的解决方案。也欢迎您提出任何其他想法。

import tensorflow as tf
import tensorflow.feature_column as fc 
import pandas as pd

PATH = "/tmp/sample.csv"

tf.enable_eager_execution()

COLUMNS = ['education','label']
train_df = pd.read_csv(PATH, header=None, names = COLUMNS)

#train_df['education'] = train_df['education'].str.split(" ")
def easy_input_function(df, label_key, num_epochs, shuffle, batch_size):
  label = df[label_key]
  ed = tf.string_split(df['education']," ")
  df['education'] = ed
  ds = tf.data.Dataset.from_tensor_slices((dict(df),label))
  if shuffle:
    ds = ds.shuffle(10000)
  ds = ds.batch(batch_size).repeat(num_epochs)
  return ds

ds = easy_input_function(train_df, label_key='label', num_epochs=5, shuffle=False, batch_size=5)


for feature_batch, label_batch in ds.take(1):
  print('Some feature keys:', list(feature_batch.keys())[:5])
  print()
  print('A batch of education  :', feature_batch['education'])
  print()
  print('A batch of Labels:', label_batch )
  print(feature_batch)

education_vocabulary_list = [
    'Bachelors', 'HS-grad', '11th', 'Masters', '9th', 'Some-college',
    'Assoc-acdm', 'Assoc-voc', '7th-8th', 'Doctorate', 'Prof-school',
    '5th-6th', '10th', '1st-4th', 'Preschool', '12th']  
education = tf.feature_column.categorical_column_with_vocabulary_list('education', vocabulary_list=education_vocabulary_list)

fc.input_layer(feature_batch, [fc.indicator_column(education)])

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-04-11
    • 1970-01-01
    • 1970-01-01
    • 2019-10-07
    • 1970-01-01
    • 2020-02-29
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多