從具有眾多標簽的Pandas資料框創建Tensorflow資料集？-有解無憂

我正在嘗試將 Pandas 資料幀加載到張量資料集中。列是 text[string] 和 labels[a list in string format]

一行看起來像： text: "Hi, this is me in here, ...." 標簽: [0, 1, 1, 0, 1, 0, 0, 0, ...]

每個文本有 17 個標簽的概率。

我找不到將資料集加載為陣列的方法，并呼叫 model.fit() 我閱讀了許多答案，嘗試在 df_to_dataset() 中使用以下代碼。

我無法弄清楚我在這方面缺少什么..

labels = labels.apply(lambda x: np.asarray(literal_eval(x)))  # Cast to a list
labels = labels.apply(lambda x: [0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0])  # Straight out list ..

#  ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type list).

列印一行（來自回傳的資料集）顯示：

({'text': <tf.Tensor: shape=(), dtype=string, numpy=b'Text in here'>}, <tf.Tensor: shape=(), dtype=string, numpy=b'[0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1.0, 0, 0, 0, 0, 0, 0]'>)

當我不使用任何轉換時，model.fit 會發送一個例外，因為它不能使用字串。

UnimplementedError:  Cast string to float is not supported
     [[node sparse_categorical_crossentropy/Cast (defined at <ipython-input-102-71a9fbf2d907>:4) ]] [Op:__inference_train_function_1193273]

def df_to_dataset(dataframe, shuffle=True, batch_size=32):
  dataframe = dataframe.copy()
  labels = dataframe.pop('labels')

  ds = tf.data.Dataset.from_tensor_slices((dict(dataframe), labels))
  return ds

train_ds = df_to_dataset(df_train, batch_size=batch_size)
val_ds = df_to_dataset(df_val, batch_size=batch_size)
test_ds = df_to_dataset(df_test, batch_size=batch_size)

def build_classifier_model():
  text_input = tf.keras.layers.Input(shape=(), dtype=tf.string, name='text')

  preprocessing_layer = hub.KerasLayer(tfhub_handle_preprocess, name='preprocessing')
  encoder_inputs = preprocessing_layer(text_input)

  encoder = hub.KerasLayer(tfhub_handle_encoder, trainable=True, name='BERT_encoder')
  outputs = encoder(encoder_inputs)
  net = outputs['pooled_output']
  net = tf.keras.layers.Dropout(0.2)(net)
  net = tf.keras.layers.Dense(17, activation='softmax', name='classifier')(net)

  return tf.keras.Model(text_input, net)


classifier_model = build_classifier_model()

loss = 'sparse_categorical_crossentropy'
metrics = ["accuracy"]
classifier_model.compile(optimizer=optimizer,
                         loss=loss,
                         metrics=metrics)

history = classifier_model.fit(x=train_ds,
                               validation_data=val_ds,
                               epochs=epochs)

uj5u.com熱心網友回復：

您可以使用方法中的tf.strings函式map。

import tensorflow as tf

x = ['[0, 1, 0]', '[1, 1, 0]']


def splitter(string):
    string = tf.strings.substr(string, 1, tf.strings.length(string) - 2) # no brackets
    string = tf.strings.split(string, ', ')                              # isolate int
    string = tf.strings.to_number(string, out_type=tf.int32)             # as integer
    return string


ds = tf.data.Dataset.from_tensor_slices(x).map(splitter)

next(iter(ds))

<tf.Tensor: shape=(3,), dtype=int32, numpy=array([0, 1, 0])>

話雖如此，您也可以更改您的 DataFrame，以便目標是單熱編碼的。

uj5u.com熱心網友回復：

在使用tf.data.Dataset.from_tensor_slices. 這是一個簡單的作業示例：

import tensorflow as tf
import tensorflow_text as tf_text
import tensorflow_hub as hub
import pandas as pd

def build_classifier_model():
  text_input = tf.keras.layers.Input(shape=(), dtype=tf.string, name='text')

  preprocessing_layer = hub.KerasLayer('https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/1', name='preprocessing')
  encoder_inputs = preprocessing_layer(text_input)

  encoder = hub.KerasLayer('https://tfhub.dev/tensorflow/small_bert/bert_en_uncased_L-2_H-128_A-2/2', trainable=True, name='BERT_encoder')
  outputs = encoder(encoder_inputs)
  net = outputs['pooled_output']
  net = tf.keras.layers.Dropout(0.2)(net)
  net = tf.keras.layers.Dense(5, activation='softmax', name='classifier')(net)
  return tf.keras.Model(text_input, net)

def remove_and_split(s):
  s = s.replace('[', '') 
  s = s.replace(']', '')  
  return s.split(',')
 
def df_to_dataset(dataframe, shuffle=True, batch_size=2):
  dataframe = dataframe.copy()
  labels = tf.squeeze(tf.constant([dataframe.pop('labels')]), axis=0)
  ds = tf.data.Dataset.from_tensor_slices((dict(dataframe), labels)).batch(
        batch_size)
  return ds

dummy_data = {'text': [
"Improve the physical fitness of your goldfish by getting him a bicycle",
"You are unsure whether or not to trust him but very thankful that you wore a turtle neck",
"Not all people who wander are lost", 
"There is a reason that roses have thorns",
"Charles ate the french fries knowing they would be his last meal",
"He hated that he loved what she hated about hate",
], 'labels': ['[0, 1, 1, 1, 1]', '[1, 1, 1, 0, 0]', '[1, 0, 1, 0, 0]', '[1, 0, 1, 0, 0]', '[1, 1, 1, 0, 0]', '[1, 1, 1, 0, 0]']}  

df = pd.DataFrame(dummy_data)  
df["labels"] = df["labels"].apply(lambda x: [int(i) for i in remove_and_split(x)])
batch_size = 2

train_ds = df_to_dataset(df, batch_size=batch_size)
val_ds = df_to_dataset(df, batch_size=batch_size)
test_ds = df_to_dataset(df, batch_size=batch_size)

loss = 'categorical_crossentropy'
metrics = ["accuracy"]

classifier_model = build_classifier_model()
classifier_model.compile(optimizer='adam',
                         loss=loss,
                         metrics=metrics)

history = classifier_model.fit(x=train_ds,
                             validation_data=val_ds,
                              epochs=5)

并且不要忘記在tf.data.Dataset.from_tensor_slices使用 Bert 預處理層時包括批量大小。我還將您的損失函式更改為categorical_crossentropy，因為您正在使用單熱編碼標簽（至少可以從您的問題中推斷出）。該sparse_categorical_crossentropy損失函式需要整數標簽不是一個熱編碼。

轉載請註明出處，本文鏈接：https://www.uj5u.com/yidong/338089.html

標籤：熊猫张量流凯拉斯张量流数据集

上一篇：有條件的VAE-無法建立模型

下一篇：如何使用tf.function從TensorFlow中的一組函式中隨機選擇