嘗試在Tensorflow中結合數字和文本特征：ValueError：層模型需要2個輸入，但它收到1個輸入張量-有解無憂

我正在嘗試使用 wine 評論資料集進行沙盒專案，并希望將文本資料和一些工程數字特征組合到神經網路中，但我收到了一個值錯誤。

我擁有的三組功能是描述（實際評論）、按比例縮放的價格和按比例縮放的字數（描述的長度）。我將 y 目標變數轉換為代表好壞評論的二分變數，將其轉化為分類問題。

這些是否是最好的特性并不是重點，但我希望嘗試將 NLP 與元資料或數字資料結合起來。當我只使用描述運行代碼時，它作業正常，但添加其他變數會導致值錯誤。

y = df['y']
X = df.drop('y', axis=1)

# split up the data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33)

X_train.head();
description_train = X_train['description']
description_test = X_test['description']

#subsetting the numeric variables
numeric_train = X_train[['scaled_price','scaled_num_words']].to_numpy()
numeric_test = X_test[['scaled_price','scaled_num_words']].to_numpy()

MAX_VOCAB_SIZE = 60000
tokenizer = Tokenizer(num_words=MAX_VOCAB_SIZE)
tokenizer.fit_on_texts(description_train)
sequences_train = tokenizer.texts_to_sequences(description_train)
sequences_test = tokenizer.texts_to_sequences(description_test)

word2idx = tokenizer.word_index
V = len(word2idx)
print('Found %s unique tokens.' % V)
Found 31598 unique tokens.

nlp_train = pad_sequences(sequences_train)
print('Shape of data train tensor:', nlp_train.shape)
Shape of data train tensor: (91944, 136)

# get sequence length
T = nlp_train.shape[1]

nlp_test = pad_sequences(sequences_test, maxlen=T)
print('Shape of data test tensor:', nlp_test.shape)
Shape of data test tensor: (45286, 136)

data_train = np.concatenate((nlp_train,numeric_train), axis=1)
data_test = np.concatenate((nlp_test,numeric_test), axis=1)



# Choosing embedding dimensionality
D = 20

# Hidden state dimensionality
M = 40

nlp_input = Input(shape=(T,),name= 'nlp_input')
meta_input = Input(shape=(2,), name='meta_input')
emb = Embedding(V   1, D)(nlp_input)
emb = Bidirectional(LSTM(64, return_sequences=True))(emb)
emb = Dropout(0.40)(emb)
emb = Bidirectional(LSTM(128))(emb)
nlp_out = Dropout(0.40)(emb)
x = tf.concat([nlp_out, meta_input], 1)
x = Dense(64, activation='swish')(x)
x = Dropout(0.40)(x)
x = Dense(1, activation='sigmoid')(x)

model = Model(inputs=[nlp_input, meta_input], outputs=[x])

#next, create a custom optimizer
optimizer1 = RMSprop(learning_rate=0.0001)

# Compile and fit
model.compile(
  loss='binary_crossentropy',
  optimizer='adam',
  metrics=['accuracy']
)


print('Training model...')
r = model.fit(
  data_train,
  y_train,
  epochs=5, 
  validation_data=(data_test, y_test))

如果這太過分了，我深表歉意，但我想確保我沒有遺漏任何可能有用的相關線索或資訊。我從運行代碼中得到的錯誤是

ValueError: Layer model expects 2 input(s), but it received 1 input tensors. Inputs received: [<tf.Tensor 'IteratorGetNext:0' shape=(None, 138) dtype=float32>]

我該如何解決該錯誤？

uj5u.com熱心網友回復：

感謝您發布所有代碼。這兩行是問題所在：

data_train = np.concatenate((nlp_train,numeric_train), axis=1)
data_test = np.concatenate((nlp_test,numeric_test), axis=1)

無論其形狀如何，numpy 陣列都被解釋為一個輸入。使用tf.data.Dataset您的資料集并將其直接提供給您的模型：

train_dataset = tf.data.Dataset.from_tensor_slices((nlp_train, numeric_train))
labels = tf.data.Dataset.from_tensor_slices(y_train)
dataset = tf.data.Dataset.zip((train_dataset, train_dataset))
r = model.fit(dataset, epochs=5)

或者直接將您的資料model.fit()作為輸入串列提供給：

r = model.fit(
  [nlp_train, numeric_train],
  y_train,
  epochs=5, 
  validation_data=([nlp_test, numeric_test], y_test))

轉載請註明出處，本文鏈接：https://www.uj5u.com/yidong/350066.html

標籤：Python 张量流凯拉斯深度学习无印良品

上一篇：Dense()和Conv2d()可以充當同一層函式嗎？

下一篇：A100tensorflowgpu錯誤：“呼叫cuInit失敗：CUDA_ERROR_NOT_INITIALIZED：初始化錯誤”