Showing posts with label tensorflow. Show all posts
Showing posts with label tensorflow. Show all posts

Tuesday, October 16, 2018

Constant Validation Accuracy with a high loss in machine learning

Leave a Comment

I'm currently trying to do create an image classification model using Inception V3 with 2 classes. I have 1428 images which are balanced about 70/30. When I run my model I get a pretty high loss of as well as a constant validation accuracy. What might be causing this constant value?

data = np.array(data, dtype="float")/255.0 labels = np.array(labels,dtype ="uint8")  (trainX, testX, trainY, testY) = train_test_split(                             data,labels,                              test_size=0.2,                              random_state=42)   img_width, img_height = 320, 320 #InceptionV3 size  train_samples =  1145  validation_samples = 287 epochs = 20  batch_size = 32  base_model = keras.applications.InceptionV3(         weights ='imagenet',         include_top=False,          input_shape = (img_width,img_height,3))  model_top = keras.models.Sequential() model_top.add(keras.layers.GlobalAveragePooling2D(input_shape=base_model.output_shape[1:], data_format=None)), model_top.add(keras.layers.Dense(350,activation='relu')) model_top.add(keras.layers.Dropout(0.2)) model_top.add(keras.layers.Dense(1,activation = 'sigmoid')) model = keras.models.Model(inputs = base_model.input, outputs = model_top(base_model.output))   for layer in model.layers[:30]:   layer.trainable = False  model.compile(optimizer = keras.optimizers.Adam(                     lr=0.00001,                     beta_1=0.9,                     beta_2=0.999,                     epsilon=1e-08),                     loss='binary_crossentropy',                     metrics=['accuracy'])  #Image Processing and Augmentation  train_datagen = keras.preprocessing.image.ImageDataGenerator(           zoom_range = 0.05,           #width_shift_range = 0.05,            height_shift_range = 0.05,           horizontal_flip = True,           vertical_flip = True,           fill_mode ='nearest')   val_datagen = keras.preprocessing.image.ImageDataGenerator()   train_generator = train_datagen.flow(         trainX,          trainY,         batch_size=batch_size,         shuffle=True)  validation_generator = val_datagen.flow(                 testX,                 testY,                 batch_size=batch_size)  history = model.fit_generator(     train_generator,      steps_per_epoch = train_samples//batch_size,     epochs = epochs,      validation_data = validation_generator,      validation_steps = validation_samples//batch_size,     callbacks = [ModelCheckpoint]) 

This is my log when I run my model:

Epoch 1/20 35/35 [==============================]35/35[==============================] - 52s 1s/step - loss: 0.6347 - acc: 0.6830 - val_loss: 0.6237 - val_acc: 0.6875  Epoch 2/20 35/35 [==============================]35/35 [==============================] - 14s 411ms/step - loss: 0.6364 - acc: 0.6756 - val_loss: 0.6265 - val_acc: 0.6875  Epoch 3/20 35/35 [==============================]35/35 [==============================] - 14s 411ms/step - loss: 0.6420 - acc: 0.6743 - val_loss: 0.6254 - val_acc: 0.6875  Epoch 4/20 35/35 [==============================]35/35 [==============================] - 14s 414ms/step - loss: 0.6365 - acc: 0.6851 - val_loss: 0.6289 - val_acc: 0.6875  Epoch 5/20 35/35 [==============================]35/35 [==============================] - 14s 411ms/step - loss: 0.6359 - acc: 0.6727 - val_loss: 0.6244 - val_acc: 0.6875  Epoch 6/20 35/35 [==============================]35/35 [==============================] - 15s 415ms/step - loss: 0.6342 - acc: 0.6862 - val_loss: 0.6243 - val_acc: 0.6875 

2 Answers

Answers 1

I think you have too low learning rate and too few epochs. try with lr = 0.001 and epochs = 100.

Answers 2

Your accuracy is 68.25%. Given that your classes are split roughly 70/30 it is likely that your model is just predicting the same thing every time, ignoring the input. That would give the accuracy you are seeing. Your model has not yet learned from your data.

As Novak said, your learning rate seems very low, so maybe try increasing that first to see if that helps.

Read More

Wednesday, September 26, 2018

Keras: Accuracy Drops While Finetuning Inception

Leave a Comment

I am having trouble fine tuning an Inception model with Keras.

I have managed to use tutorials and documentation to generate a model of fully connected top layers that classifies my dataset into their proper categories with an accuracy over 99% using bottleneck features from Inception.

import numpy as np from keras.preprocessing.image import ImageDataGenerator from keras.models import Sequential from keras.layers import Dropout, Flatten, Dense from keras import applications   # dimensions of our images. img_width, img_height = 150, 150  #paths for saving weights and finding datasets top_model_weights_path = 'Inception_fc_model_v0.h5' train_data_dir = '../data/train2' validation_data_dir = '../data/train2'   #training related parameters? inclusive_images = 1424 nb_train_samples = 1424 nb_validation_samples = 1424 epochs = 50 batch_size = 16   def save_bottlebeck_features():     datagen = ImageDataGenerator(rescale=1. / 255)      # build bottleneck features     model = applications.inception_v3.InceptionV3(include_top=False, weights='imagenet', input_shape=(img_width,img_height,3))      generator = datagen.flow_from_directory(         train_data_dir,         target_size=(img_width, img_height),         batch_size=batch_size,         class_mode='categorical',         shuffle=False)      bottleneck_features_train = model.predict_generator(         generator, nb_train_samples // batch_size)      np.save('bottleneck_features_train', bottleneck_features_train)      generator = datagen.flow_from_directory(         validation_data_dir,         target_size=(img_width, img_height),         batch_size=batch_size,         class_mode='categorical',         shuffle=False)      bottleneck_features_validation = model.predict_generator(         generator, nb_validation_samples // batch_size)      np.save('bottleneck_features_validation', bottleneck_features_validation)  def train_top_model():     train_data = np.load('bottleneck_features_train.npy')     train_labels = np.array(range(inclusive_images))      validation_data = np.load('bottleneck_features_validation.npy')     validation_labels = np.array(range(inclusive_images))      print('base size ', train_data.shape[1:])      model = Sequential()     model.add(Flatten(input_shape=train_data.shape[1:]))     model.add(Dense(1000, activation='relu'))     model.add(Dense(inclusive_images, activation='softmax'))     model.compile(loss='sparse_categorical_crossentropy',              optimizer='Adam',              metrics=['accuracy'])      proceed = True      #model.load_weights(top_model_weights_path)      while proceed:         history = model.fit(train_data, train_labels,               epochs=epochs,               batch_size=batch_size)#,               #validation_data=(validation_data, validation_labels), verbose=1)         if history.history['acc'][-1] > .99:             proceed = False      model.save_weights(top_model_weights_path)   save_bottlebeck_features() train_top_model() 

Epoch 50/50 1424/1424 [==============================] - 17s 12ms/step - loss: 0.0398 - acc: 0.9909

I have also been able to stack this model on top of inception to create my full model and use that full model to successfully classify my training set.

from keras import Model from keras import optimizers from keras.callbacks import EarlyStopping  img_width, img_height = 150, 150  top_model_weights_path = 'Inception_fc_model_v0.h5' train_data_dir = '../data/train2' validation_data_dir = '../data/train2'   #how many inclusive examples do we have? inclusive_images = 1424 nb_train_samples = 1424 nb_validation_samples = 1424 epochs = 50 batch_size = 16  # build the complete network for evaluation base_model = applications.inception_v3.InceptionV3(weights='imagenet', include_top=False, input_shape=(img_width,img_height,3))  top_model = Sequential() top_model.add(Flatten(input_shape=base_model.output_shape[1:])) top_model.add(Dense(1000, activation='relu')) top_model.add(Dense(inclusive_images, activation='softmax'))  top_model.load_weights(top_model_weights_path)  #combine base and top model fullModel = Model(input= base_model.input, output= top_model(base_model.output))  #predict with the full training dataset results = fullModel.predict_generator(ImageDataGenerator(rescale=1. / 255).flow_from_directory(         train_data_dir,         target_size=(img_width, img_height),         batch_size=batch_size,         class_mode='categorical',         shuffle=False)) 

inspection of the results from processing on this full model match the accuracy of the bottleneck generated fully connected model.

import matplotlib.pyplot as plt import operator  #retrieve what the softmax based class assignments would be from results resultMaxClassIDs = [ max(enumerate(result), key=operator.itemgetter(1))[0] for result in results]  #resultMaxClassIDs should be equal to range(inclusive_images) so we subtract the two and plot the log of the absolute value  #looking for spikes that indicate the values aren't equal  plt.plot([np.log(np.abs(x)+10) for x in (np.array(resultMaxClassIDs) - np.array(range(inclusive_images)))]) 

results: spikes are misclassifications

Here is the problem: When I take this full model and attempt to train it, Accuracy drops to 0 even though validation remains above 99%.

model2 = fullModel  for layer in model2.layers[:-2]:     layer.trainable = False  # compile the model with a SGD/momentum optimizer # and a very slow learning rate. #model.compile(loss='binary_crossentropy', optimizer=optimizers.SGD(lr=1e-4, momentum=0.9),  metrics=['accuracy'])  model2.compile(loss='categorical_crossentropy',              optimizer=optimizers.SGD(lr=1e-4, momentum=0.9),               metrics=['accuracy'])  train_datagen = ImageDataGenerator(rescale=1. / 255)  test_datagen = ImageDataGenerator(rescale=1. / 255)  train_generator = train_datagen.flow_from_directory(     train_data_dir,     target_size=(img_height, img_width),     batch_size=batch_size,     class_mode='categorical')  validation_generator = test_datagen.flow_from_directory(     validation_data_dir,     target_size=(img_height, img_width),     batch_size=batch_size,     class_mode='categorical')  callback = [EarlyStopping(monitor='acc', min_delta=0, patience=3, verbose=0, mode='auto', baseline=None)] # fine-tune the model model2.fit_generator(     #train_generator,     validation_generator,     steps_per_epoch=nb_train_samples//batch_size,     validation_steps = nb_validation_samples//batch_size,     epochs=epochs,     validation_data=validation_generator) 

Epoch 1/50 89/89 [==============================] - 388s 4s/step - loss: 13.5787 - acc: 0.0000e+00 - val_loss: 0.0353 - val_acc: 0.9937

and it gets worse as things progress

Epoch 21/50 89/89 [==============================] - 372s 4s/step - loss: 7.3850 - acc: 0.0035 - val_loss: 0.5813 - val_acc: 0.8272

The only thing I could think of is that somehow the training labels are getting improperly assigned on this last train, but I've successfully done this with similar code using VGG16 before.

I have searched over the code trying to find a discrepancy to explain why a model making accurate predictions over 99% of the time drops its training accuracy while maintaining validation accuracy during fine tuning, but I can't figure it out. Any help would be appreciated.

Information about the code and environment:

Things that are going to stand out as weird, but are meant to be that way:

  • There is only 1 image per class. This NN is intended to classify objects whose environmental and orientation conditions are controlled. Their is only one acceptable image for each class corresponding to the correct environmental and rotational situation.
  • The test and validation set are the same. This NN is only ever designed to be used on the classes it is being trained on. The images it will process will be carbon copies of the class examples. It is my intent to overfit the model to these classes

I am using:

  • Windows 10
  • Python 3.5.6 under Anaconda client 1.6.14
  • Keras 2.2.2
  • Tensorflow 1.10.0 as the backend
  • CUDA 9.0
  • CuDNN 8.0

I have checked out:

  1. Keras accuracy discrepancy in fine-tuned model
  2. VGG16 Keras fine tuning: low accuracy
  3. Keras: model accuracy drops after reaching 99 percent accuracy and loss 0.01
  4. Keras inception v3 retraining and finetuning error
  5. How to find which version of TensorFlow is installed in my system?

but they appear unrelated.

2 Answers

Answers 1

Note: Since your problem is a bit strange and difficult to debug without having your trained model and dataset, this answer is just a (best) guess after considering many things that may have could go wrong. Please provide your feedback and I will delete this answer if it does not work.

Since the inception_V3 contains BatchNormalization layers, maybe the problem is due to (somehow ambiguous or unexpected) behavior of this layer when you set trainable parameter to False (1, 2, 3, 4).

Now, let's see if this is the root of the problem: as suggested by @fchollet, set the learning phase when defining the model for fine-tuning:

from keras import backend as K  K.set_learning_phase(0)  base_model = applications.inception_v3.InceptionV3(weights='imagenet', include_top=False, input_shape=(img_width,img_height,3))  for layer in base_model.layers:     layer.trainable = False  K.set_learning_phase(1)  top_model = Sequential() top_model.add(Flatten(input_shape=base_model.output_shape[1:])) top_model.add(Dense(1000, activation='relu')) top_model.add(Dense(inclusive_images, activation='softmax'))  top_model.load_weights(top_model_weights_path)  #combine base and top model fullModel = Model(input= base_model.input, output= top_model(base_model.output))  fullModel.compile(loss='categorical_crossentropy',              optimizer=optimizers.SGD(lr=1e-4, momentum=0.9),               metrics=['accuracy'])   ##################################################################### # Here, define the generators and then fit the model same as before # ##################################################################### 

Side Note: This is not causing any problem in your case, but keep in mind that when you use top_model(base_model.output) the whole Sequential model (i.e. top_model) is stored as one layer of fullModel. You can verify this by either using fullModel.summary() or print(fullModel.layers[-1]). Hence when you used:

for layer in model2.layers[:-2]:     layer.trainable = False  

you are actually not freezing the last layer of base_model as well. However, since it is a Concatenate layer, and therefore does not have trainable parameters, no problem occurs and it would behave as you intended.

Answers 2

Like the previous reply, I'll try to share some thoughts to see whether it helps.

There are a couple of things that called my attention (and maybe are worth reviewing). Note: some of them should have given you issues with the separate models as well.

  • Correct if I'm wrong, but it seems you used sparse_categorical_crossentropy for the first training while you used categorical_crossentropy for the second one. Is it correct? Because I believe they assume labels differently (sparse assumes integers and the other assumes one-hot).
  • Have you tried to set the layers you added in the end as trainable = True? I know that you have already set the others to trainable = False, but maybe that's something worth checking too.
  • It seems the data generator is not making use of the default preprocessing function used in Inception v3, which uses a per-mean channel.
  • Have you tried any experiment using Functional instead of Sequential API?

I hope that helps.

Read More

Sunday, September 23, 2018

Tensorflow eager choose checkpoint max to keep

Leave a Comment

I'm writing a process-based implementation of a3c with tensorflow in eager mode. After every gradient update, my general model writes its parameters as checkpoints to a folder. The workers then update their parameters by loading the last checkpoints from this folder. However, there is a problem.

Often times, while the worker is reading the last available checkpoint from the folder, the master network will write new checkpoints to the folder and sometimes will erase the checkpoint that the worker is reading. A simple solution would be raising the maximum of checkpoints to keep. However, tfe.Checkpoint and tfe.Saver don't have a parameter to choose the max to keep.

Is there a way to achieve this?

2 Answers

Answers 1

For the tf.train.Saver you can specify max_to_keep:

tf.train.Saver(max_to_keep = 10) 

and max_to_keep seems to be present in the both fte.Saver and it's tf.training.Saver.

I haven't tried if it works though.

Answers 2

It seems the suggested way of doing checkpoint deletion is to use the CheckpointManager.

import tensorflow as tf checkpoint = tf.train.Checkpoint(optimizer=optimizer, model=model) manager = tf.contrib.checkpoint.CheckpointManager(      checkpoint, directory="/tmp/model", max_to_keep=5) status = checkpoint.restore(manager.latest_checkpoint) while True: # train   manager.save() 
Read More

Tuesday, September 18, 2018

Trying to retrain a tensorflow model, input and output nodes disappear

Leave a Comment

I am trying to retrain the tensorflow deeplab model using MobileNet_V2. I have downloaded the checkpoint from the deeplab model zoo, about halfway down this page: https://github.com/tensorflow/models/blob/master/research/deeplab/g3doc/model_zoo.md Specifically, the mobilenetv2_coco_voc_trainaug one. I would like my retrained output to have the same graph, but different parameters as this one. (Well, almost the same graph, the final tensor should probably have a different shape because I am trying to work with a different number of classes.)

I assembled my own images into a tfrecord, labelled with just one class for now. This is practice for a dataset with 4 classes.

I then ran the following to retrain the network, producing .pbtxt, .meta, .index and .data-00000-of-00001 files:

PATH_TO_INITIAL_CHECKPOINT=/path/to/unzipped/files/model.ckpt-30000.index PATH_TO_TRAIN_DIR=/path/to/checkpoints/ PATH_TO_DATASET=/path/to/tfrecord python /path/to/tensorflow/models/research/deeplab/train.py \     --logtostderr \     --training_number_of_steps=900 \ # 90000 \     --train_split="train" \     --model_variant="mobilenet_v2" \     --output_stride=16 \     --decoder_output_stride=4 \     --train_crop_size=128 \     --train_crop_size=128 \     --train_batch_size=1 \     --dataset="cityscapes" \     --tf_initial_checkpoint=${PATH_TO_INITIAL_CHECKPOINT} \     --train_logdir=${PATH_TO_TRAIN_DIR} \     --dataset_dir=${PATH_TO_DATASET} \     --initialize_last_layer=False \     --last_layers_contain_logits_only=True \     --fine_tune_batch_norm=False 

Running bazel's summarize_graph on the downloaded file gives:

Found 1 possible inputs: (name=ImageTensor, type=uint8(4), shape=[1,?,?,3])  No variables spotted. Found 1 possible outputs: (name=SemanticPredictions, op=Slice)  

When I scan the nodes of the .pbtxt file, I can't find any nodes called ImageTensor or SemanticPredictions. I have tried with tensorboard, bazel's summarize_graph, and programmatically (e.g. here, here, or here). Summarize_graph says No inputs spotted and Found 664 possible outputs:.

This then leads to problems with freeze_graph.py. If I choose output_node_names from what I can see on tensorbord, then freeze_graph.py runs, and I am able to get a frozen graph. But running that model gives me

TypeError: Cannot interpret feed_dict key as Tensor: The name  'ImageTensor:0' refers to a Tensor which does not exist. The operation,  'ImageTensor', does not exist in the graph. 

I'm definitely doing something wrong here. The question is: what? I suspect it could be the arguments I supply to train.py, but really, that's just a shot in the dark. It could be that this is not how train.py is intended to be used, or deeplab's train.py is not compatible with MobileNetV2.

Edit: After a closer look at the options available in train.py, I have updated my command. Cleaning previous failed models from the TRAIN_DIR was also helpful to avoid the error:

Restoring from checkpoint failed. This is most likely due to a mismatch  between the current graph and the graph from the checkpoint. Please ensure  that you have not altered the graph expected based on the checkpoint. 

0 Answers

Read More

Sunday, September 16, 2018

Why does tensorflow/keras choke when I try to fit multiple models in parallel?

Leave a Comment

I'm trying to fit a finite mixture model, with the mixture models for each class being neural networks. It'd be super-useful for me to be able to be able to parallelize, because keras doesn't max out all of the available cores on my laptop, let alone a large cluster.

But when I try to set different learning rates for different models inside of a parallel foreach loop the whole thing chokes.

What is going on? I suspect that it has something to do with scope -- the workers aren't running on separate instantiations of tensorflow, maybe. But I really don't know. How can I make this work? And what do I need to understand to know why this doesn't work?

Here's a MWE. Set the foreach loop to %do% and it works fine. Set it to %dopar% and it chokes on the fitting stage.

library(foreach) library(doParallel) registerDoParallel(2) library(keras) library(tensorflow) mnist <- dataset_mnist() x_train <- mnist$train$x y_train <- mnist$train$y x_test <- mnist$test$x y_test <- mnist$test$y  x_train <- array_reshape(x_train, c(nrow(x_train), 784)) x_test <- array_reshape(x_test, c(nrow(x_test), 784)) # rescale x_train <- x_train / 255 x_test <- x_test / 255  y_train <- to_categorical(y_train, 10) y_test <- to_categorical(y_test, 10)  # make tensorflow run single-threaded session_conf <- tf$ConfigProto(intra_op_parallelism_threads = 1L,                                inter_op_parallelism_threads = 1L) # Create the session using the custom configuration sess <- tf$Session(config = session_conf) K <- backend() K$set_session(sess)   models <- foreach(i = 1:2) %dopar%{   model <- keras_model_sequential()    model %>%      layer_dense(units = 256/i, activation = 'relu', input_shape = c(784)) %>%      layer_dropout(rate = 0.4) %>%      layer_dense(units = 128/i, activation = 'relu') %>%     layer_dropout(rate = 0.3) %>%     layer_dense(units = 10, activation = 'softmax')    print("A")   model %>% compile(     loss = 'categorical_crossentropy',     optimizer = optimizer_rmsprop(),     metrics = c('accuracy')   )   print("B")   history <- model %>% fit(     x_train, y_train,      epochs = 3, batch_size = 128,      validation_split = 0.2, verbose = 0   )   print("done")   } 

Here's sessionInfo():

R version 3.5.1 (2018-07-02) Platform: x86_64-pc-linux-gnu (64-bit) Running under: Ubuntu 18.04.1 LTS  Matrix products: default BLAS: /usr/lib/x86_64-linux-gnu/blas/libblas.so.3.7.1 LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.7.1  locale:  [1] LC_CTYPE=en_US.UTF-8       LC_NUMERIC=C               LC_TIME=en_US.UTF-8        LC_COLLATE=en_US.UTF-8     LC_MONETARY=en_US.UTF-8     [6] LC_MESSAGES=en_US.UTF-8    LC_PAPER=en_US.UTF-8       LC_NAME=C                  LC_ADDRESS=C               LC_TELEPHONE=C             [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C         attached base packages: [1] splines   parallel  stats     graphics  grDevices utils     datasets  methods   base       other attached packages:  [1] panelNNET_1.0       matrixStats_0.54.0  MASS_7.3-50         lfe_2.8-2           tensorflow_1.9      keras_2.1.6.9005     [7] mgcv_1.8-24         nlme_3.1-137        scales_1.0.0        forcats_0.3.0       stringr_1.3.1       purrr_0.2.5         [13] readr_1.1.1         tidyr_0.8.1         tibble_1.4.2        tidyverse_1.2.1     maptools_0.9-3      rgeos_0.3-28        [19] rgdal_1.3-4         sp_1.3-1            broom_0.5.0         ggplot2_3.0.0       randomForest_4.6-14 dplyr_0.7.6         [25] glmnet_2.0-16       Matrix_1.2-14       doBy_4.6-2          doParallel_1.0.11   iterators_1.0.10    foreach_1.4.4        loaded via a namespace (and not attached):  [1] httr_1.3.1          jsonlite_1.5        modelr_0.1.2        Formula_1.2-3       assertthat_0.2.0    cellranger_1.1.0     [7] yaml_2.2.0          pillar_1.3.0        backports_1.1.2     lattice_0.20-35     glue_1.3.0          reticulate_1.10     [13] digest_0.6.15       RcppEigen_0.3.3.4.0 rvest_0.3.2         colorspace_1.3-2    sandwich_2.5-0      plyr_1.8.4          [19] pkgconfig_2.0.1     haven_1.1.2         xtable_1.8-2        whisker_0.3-2       withr_2.1.2         lazyeval_0.2.1      [25] cli_1.0.0           magrittr_1.5        crayon_1.3.4        readxl_1.1.0        xml2_1.2.0          foreign_0.8-70      [31] tools_3.5.1         hms_0.4.2           munsell_0.5.0       bindrcpp_0.2.2      compiler_3.5.1      rlang_0.2.2         [37] grid_3.5.1          rstudioapi_0.7      base64enc_0.1-3     labeling_0.3        gtable_0.2.0        codetools_0.2-15    [43] R6_2.2.2            tfruns_1.3          zoo_1.8-3           lubridate_1.7.4     zeallot_0.1.0       bindr_0.1.1         [49] stringi_1.2.4       Rcpp_0.12.18        tidyselect_0.2.4 

1 Answers

Answers 1

Keras requires there is only one training in a given session. I would try to create a different session for each model.

I would insert this part of the code inside the %dopar%, to create a different session per model

sess <- tf$Session(config = session_conf) K <- backend() K$set_session(sess) 
Read More

Thursday, September 13, 2018

“ValueError: Trying to share variable $var, but specified dtype float32 and found dtype float64_ref” when trying to use get_variable

Leave a Comment

I am trying to build a custom variational autoencoder network, where in I'm initializing the decoder weights using the transpose of the weights from the encoder layer, I couldn't find something native to tf.contrib.layers.fully_connected so I used tf.assign instead, here's my code for the layers:

def inference_network(inputs, hidden_units, n_outputs):     """Layer definition for the encoder layer."""     net = inputs     with tf.variable_scope('inference_network', reuse=tf.AUTO_REUSE):         for layer_idx, hidden_dim in enumerate(hidden_units):             net = layers.fully_connected(                 net,                 num_outputs=hidden_dim,                 weights_regularizer=layers.l2_regularizer(training_params.weight_decay),                 scope='inf_layer_{}'.format(layer_idx))             add_layer_summary(net)         z_mean = layers.fully_connected(net, num_outputs=n_outputs, activation_fn=None)         z_log_sigma = layers.fully_connected(             net, num_outputs=n_outputs, activation_fn=None)      return z_mean, z_log_sigma   def generation_network(inputs, decoder_units, n_x):     """Define the decoder network."""     net = inputs  # inputs here is the latent representation.     with tf.variable_scope("generation_network", reuse=tf.AUTO_REUSE):         assert(len(decoder_units) >= 2)         # First layer does not have a regularizer         net = layers.fully_connected(             net,             decoder_units[0],             scope="gen_layer_0",         )         for idx, decoder_unit in enumerate([decoder_units[1], n_x], 1):             net = layers.fully_connected(                 net,                 decoder_unit,                 scope="gen_layer_{}".format(idx),                 weights_regularizer=layers.l2_regularizer(training_params.weight_decay)             )     # Assign the transpose of weights to the respective layers     tf.assign(tf.get_variable("generation_network/gen_layer_1/weights"),               tf.transpose(tf.get_variable("inference_network/inf_layer_1/weights")))     tf.assign(tf.get_variable("generation_network/gen_layer_1/bias"),               tf.get_variable("generation_network/inf_layer_0/bias"))     tf.assign(tf.get_variable("generation_network/gen_layer_2/weights"),               tf.transpose(tf.get_variable("inference_network/inf_layer_0/weights")))     return net # x_recon 

It is wrapped using this tf.slim arg_scope:

def _autoencoder_arg_scope(activation_fn):     """Create an argument scope for the network based on its parameters."""      with slim.arg_scope([layers.fully_connected],                         weights_initializer=layers.xavier_initializer(),                         biases_initializer=tf.initializers.constant(0.0),                         activation_fn=activation_fn) as arg_sc:         return arg_sc 

However I'm getting the error: ValueError: Trying to share variable VarAutoEnc/generation_network/gen_layer_1/weights, but specified dtype float32 and found dtype float64_ref. I have narrowed this down to the get_variablecall, but I don't know why it's failing.

If there is a way where you can initialize a tf.contrib.layers.fully_connected from another fully connected layer without a tf.assign operation, that solution is fine with me.

1 Answers

Answers 1

I can't reproduce your error. Here is a minimalistic runnable example that does the same as your code:

import tensorflow as tf  with tf.contrib.slim.arg_scope([tf.contrib.layers.fully_connected],                                weights_initializer=tf.contrib.layers.xavier_initializer(),                                biases_initializer=tf.initializers.constant(0.0)):    i = tf.placeholder(tf.float32, [1, 30])    with tf.variable_scope("inference_network", reuse=tf.AUTO_REUSE):     tf.contrib.layers.fully_connected(i, 30, scope="gen_layer_0")    with tf.variable_scope("generation_network", reuse=tf.AUTO_REUSE):     tf.contrib.layers.fully_connected(i, 30, scope="gen_layer_0",       weights_regularizer=tf.contrib.layers.l2_regularizer(0.01))    with tf.variable_scope("", reuse=tf.AUTO_REUSE):     tf.assign(tf.get_variable("generation_network/gen_layer_0/weights"),               tf.transpose(tf.get_variable("inference_network/gen_layer_0/weights"))) 

The code runs without a ValueError. If you get a ValueError running this, then it is probably a bug that has been fixed in a later tensorflow version (I tested on 1.9). Otherwise the error is part of your code that you don't show in the question.

By the way, assign will return an op that will perform the assignment once the returned op is run in a session. So you will want to return the output of all assign calls in the generation_network function. You can bundle all assign ops into one using tf.group.

Read More

Tuesday, September 11, 2018

Tensorflow Estimator: Cache bottlenecks

Leave a Comment

When following the tensorflow image classification tutorial, at first it caches the bottleneck of each image:

def: cache_bottlenecks())

I have rewritten the training using tensorflow's Estimator. This really simplified all the code. However I want to cache the bottleneck features here.

Here is my model_fn. I want to cache the results of the dense layer so I can make changes to the actual training without having to compute the bottlenecks each time.

How can I accomplish that?

def model_fn(features, labels, mode, params):     is_training = mode == tf.estimator.ModeKeys.TRAIN      num_classes = len(params['label_vocab'])      module = hub.Module(params['module_spec'], trainable=is_training and params['train_module'])     bottleneck_tensor = module(features['image'])      with tf.name_scope('final_retrain_ops'):         logits = tf.layers.dense(bottleneck_tensor, units=num_classes, trainable=is_training)  # save this?      def train_op_fn(loss):         optimizer = tf.train.AdamOptimizer()         return optimizer.minimize(loss, global_step=tf.train.get_global_step())      head = tf.contrib.estimator.multi_class_head(n_classes=num_classes, label_vocabulary=params['label_vocab'])      return head.create_estimator_spec(         features, mode, logits, labels, train_op_fn=train_op_fn     ) 

0 Answers

Read More

Wednesday, August 29, 2018

Meaning of TensorFlow operation ` IsExpensive()`?

Leave a Comment

There is a method in OpKernel

 // Returns true iff this op kernel is considered "expensive". The  // runtime may use this flag to optimize graph execution for example  // to "inline" inexpensive kernels.  virtual bool IsExpensive() { return expensive_; } 

It seems that by default all operations on the GPU are considered as inexpensive whilst CPU, SYSL are flagged as expensive.

It is a bit hard to figure out the definition and effect of expensive. The is no information in the guide.

  1. Is there any specific guideline when IsExpensive should be false, true?
  2. What's the effect if an operation is flagged as expensive? So far I can only tell, that active profiling uses this just as a hint ? The only place querying this property is in the scheduler but without explaining what being inline means.
  3. In conjunction with "1." should I care about it in my custom Ops?
  4. While it makes sense, that any AsyncOp (like RemoteFusedGraphExecuteOp) is expensive, MPIAllgatherOp seems to be defined as not expensive. Isn't this a contradiction?

I am asking, because the IdentityOp is explicitly marked as inexpensive. I wonder, if I should override this method in my custom ops as well, since each CPU version (even any custom code) is flagged as expensive.

The entire logic of XLA seems to be about wether an instruction is expensive or not. So it might be an important part to consider. Therefore, a coin-toss about true/false might be not the best way to decide the return value in my custom op.

1 Answers

Answers 1

Before answering your questions I think it is worth trying to understand how TensorFlow uses threads in order to get your work done. For this, I suggest you read this related and very good SO post.

You will find that TensorFlow uses a thread-pool in order to get you work done. The expensive Ops are being scheduled for execution on the thread-pool, whereas the cheap Ops are executed "inline" meaning by the same thread which schedules the tasks (Sidenote: from the source file you have linked you find only one exception, i.e. when the inline_ready queue is empty the thread can execute the last expensive Op by itself.).

With this in mind, let us try to answer your questions.

  1. Is there any specific guideline when IsExpensive should be false, true?

I could not find a specific guideline in the TensorFlow manual, however, from the internals of what we discussed above a Op should be marked to be expensive, when the offset of scheduling a task to the thread pool is neglectable in comparison to the time the task needs to be executed.

  1. What's the effect if an operation is flagged as expensive? So far I can only tell, that active profiling uses this just as a hint ? The only place querying this property is in the scheduler but without explaining what being inline means.

The effect is the following, everytime an Ops IsExpensive method returns false it might be pushed to the inline_ready queue and may block the thread from performing further tasks hence stalling your programm. In contrast, if the Ops IsExpensive method returns true, it will be scheduled for execution on the thread pool and the scheduling thread is free to continue doing its tasks in the process loop.

  1. In conjunction with "1." should I care about it in my custom Ops?

I think you should care and try to reason as much as possible about the execution time of you Op. After that decide how you implement the IsExpensive method.

  1. While it makes sense, that any AsyncOp (like RemoteFusedGraphExecuteOp) is expensive, MPIAllgatherOp seems to be defined as not expensive. Isn't this a contradiction?

No, it is not a contradiction. If you read the comment of MPIAllgatherOp you will find the following:

// Although this op is handled asynchronously, the ComputeAsync call is // very inexpensive. It only sets up a CollectiveOpRecord and places it // in the table for the background thread to handle. Thus, we do not need // a TF pool thread to perform the op. 

Which clearly states, that scheduling this task for the thread pool would almost only be overhead. Therefore performing it inline makes a lot of sense.

Read More

Sunday, August 26, 2018

Opening Keras model with embedding layer in Tensorflow in Golang

Leave a Comment

I am trying to implement my Keras neural network in Go using the tfgo package. The model includes 2 regular inputs and two Keras embedding layers. It looks like this:

embedding_layer = Embedding(vocab_size,                             100,                             weights=[embedding_matrix],                             input_length=100,                             trainable=False)  sequence_input = Input(shape=(max_length,), dtype='int32') embedded_sequences = embedding_layer(sequence_input) text_lstm = Bidirectional(LSTM(256))(embedded_sequences) text_lstm = Dropout(0.5)(text_lstm) text_lstm  = Dense(512, activation='relu')(text_lstm ) text_lstm = Dropout(0.5)(text_lstm) text_lstm  = Dense(256, activation='relu')(text_lstm) text_lstm = Dropout(0.5)(text_lstm) text_lstm  = Dense(128, activation='relu')(text_lstm) text_lstm = Dropout(0.5)(text_lstm)  title_input = Input(shape=(max_title_length,), dtype='int32') title_embed = Embedding(vocab_size, embedding_vector_length, input_length=max_title_length)(title_input) title_lstm = Bidirectional(LSTM(128))(title_embed) title_lstm = Dropout(0.5)(title_lstm) title_lstm  = Dense(512, activation='relu')(title_lstm ) title_lstm = Dropout(0.5)(title_lstm) title_lstm  = Dense(256, activation='relu')(title_lstm) title_lstm = Dropout(0.5)(title_lstm) title_lstm  = Dense(128, activation='relu')(title_lstm) title_lstm = Dropout(0.5)(title_lstm)   merged = concatenate([text_lstm, title_lstm])   merged_d1 = Dense(1024, activation='relu')(merged) merged_d1 = Dropout(0.5)(merged_d1) merged_d1 = Dense(512, activation='relu')(merged_d1) merged_d1 = Dropout(0.5)(merged_d1)   text_class = Dense(num_classes, activation='sigmoid')(merged_d1) model = Model([sequence_input, title_input], text_class) 

I'm trying to load the model in Go, so far I think I've been able to include the regular input layers like this:

s := make([]int32, 100) s1 := make([]int32, 15) model := tg.LoadModel("myModel3", []string{"myTag"}, nil) tensor1, _ := tf.NewTensor(s) tensor2, _ := tf.NewTensor(s1)  result := model.Exec([]tf.Output{     model.Op("dense_18/Sigmoid", 0), }, map[tf.Output]*tf.Tensor{     model.Op("input_1", 0): tensor1,     model.Op("input_3", 0): tensor2, }) 

But when I run the code, it reminds me that there are actually two more "inputs":

panic: You must feed a value for placeholder tensor 'input_4' with dtype int32 and shape [?,15]      [[Node: input_4 = Placeholder[_output_shapes=[[?,15]], dtype=DT_INT32, shape=[?,15], _device="/job:localhost/replica:0/task:0/device:CPU:0"]()]] 

I would imagine that these would be "input2" and "input4" and that they would need to be initialized somehow with the embeddings from the model, but I have no idea how I could do this in Tensorflow in Go (I am new to Go).

I tried the following:

    s := make([]int32, 100) s1 := make([]int32, 15)  tensor1e, _ := tf.NewTensor([1][100][2]float32{}) tensor2e, _ := tf.NewTensor([1][15][2]float32{})  tensor1, _ := tf.NewTensor(s) tensor2, _ := tf.NewTensor(s1)  result := model.Exec([]tf.Output{     model.Op("dense_18/Sigmoid", 0), }, map[tf.Output]*tf.Tensor{     model.Op("input_3", 0):                tensor1,     model.Op("embedding_2/embeddings", 0): tensor2e,     model.Op("embedding_1/embeddings", 0): tensor1e,     model.Op("input_4", 0):                tensor2, })  But this  

produced the following, error:

2018-08-17 19:50:00.543771: W tensorflow/core/framework  /op_kernel.cc:1275] OP_REQUIRES failed at transpose_op.cc:157 : Invalid argument: transpose expects a vector of size 2. But input(1) is a vector of size 3 2018-08-17 19:50:00.543792: W tensorflow/core/framework/op_kernel.cc:1275] OP_REQUIRES failed at reduction_ops_common.h:155 : Invalid argument: Invalid reduction dimension (2 for input with 2 dimension(s) panic: Invalid reduction dimension (2 for input with 2 dimension(s)      [[Node: bidirectional_4/Sum = Sum[T=DT_FLOAT, Tidx=DT_INT32, _output_shapes=[[?]], keep_dims=false, _device="/job:localhost/replica:0/task:0/device:CPU:0"](bidirectional_4/zeros_like, bidirectional_3/Sum/reduction_indices)]] 

Can anyone point me in the right direction on how to complete this operation? Any help would be much appreciated!

1 Answers

Answers 1

So, it turns out that I did not need to specify the inputs for the embedding layers. I was actually structuring the input incorrectly. It should look like this:

tensor1, _ := tf.NewTensor([][]int32{tokes_text}) tensor2, _ := tf.NewTensor([][]int32{tokes_title})   result := model.Exec([]tf.Output{             model.Op("dense_18/Sigmoid", 0),         }, map[tf.Output]*tf.Tensor{             model.Op("input_3", 0): tensor1,             model.Op("input_4", 0): tensor2,         }) 
Read More

Wednesday, August 22, 2018

Keras: Masking and Flattening

Leave a Comment

I'm having difficulty building a straightforward model that deals with masked input values. My training data consists of variable-length lists of GPS traces, i.e. lists where each element contains Latitude and Longitude.

There are 70 training examples

enter image description here

Since they have variable lengths I am padding them with zeros, with the aim of then telling Keras to ignore these zero-values.

train_data = keras.preprocessing.sequence.pad_sequences(train_data, maxlen=max_sequence_len, dtype='float32',                                             padding='pre', truncating='pre', value=0) 

enter image description here

I then build a very basic model like so

model = Sequential() model.add(Dense(16, activation='relu',input_shape=(max_sequence_len, 2))) model.add(Flatten()) model.add(Dense(2, activation='sigmoid')) 

After some previous trial and error I realised that I need the Flatten layer or fitting the model would throw the error

ValueError: Error when checking target: expected dense_87 to have 3 dimensions, but got array with shape (70, 2) 

By including this Flatten layer, however, I can not use a Masking layer (to ignore the padded zeros) or Keras throws this error

TypeError: Layer flatten_31 does not support masking, but was passed an input_mask: Tensor("masking_9/Any_1:0", shape=(?, 48278), dtype=bool) 

I have searched extensively, reading GitHub issues and plenty of Q/A here but I can't figure it out.

2 Answers

Answers 1

Masking does seem bugged. But do not worry: the 0s are not going to make your model worse; at most less efficient.

I would recommend using a Convolutional approach instead of pure Dense or perhaps RNN. I think this will work really well for GPS data.

Please try the following code:

from keras.preprocessing.sequence import pad_sequences from keras import Sequential from keras.layers import Dense, Flatten, Masking, LSTM, GRU, Conv1D, Dropout, MaxPooling1D import numpy as np import random  max_sequence_len = 70  n_samples = 100 num_coordinates = 2 # lat/long  data = [[[random.random() for _ in range(num_coordinates)]          for y in range(min(x, max_sequence_len))]         for x in range(n_samples)]  train_y = np.random.random((n_samples, 2))  train_data = pad_sequences(data, maxlen=max_sequence_len, dtype='float32',                            padding='pre', truncating='pre', value=0)  model = Sequential() model.add(Conv1D(32, (5, ), input_shape=(max_sequence_len, num_coordinates))) model.add(Dropout(0.5)) model.add(MaxPooling1D()) model.add(Flatten()) model.add(Dense(2, activation='relu')) model.compile(loss='mean_squared_error', optimizer="adam") model.fit(train_data, train_y) 

Answers 2

Instead of using a Flatten layer, you could use a Global Pooling layer.

These are suited to collapse the length/time dimension without losing the capability of using variable lengths.

So, instead of Flatten(), you can try a GlobalAveragePooling1D or GlobalMaxPooling1D.

None of them use supports_masking in their code, so they must be used with care.

The average one will consider more inputs than the max (thus the values that should be masked).

The max will take only one from the length. With luck, if all your useful values are higher than the ones in the masked position, it will indirectly preserve the mask. It will probably need even more input neurons than the other.

That said, yes, try the Conv1D or RNN (LSTM) appoaches suggested.


Creating a custom pooling layer with mask

You can also create your own pooling layer (needs a functional API model where you pass both the model's inputs and the tensor which you want to pool)

Below, a working example with average pooling applying a mask based on the inputs:

def customPooling(maskVal):     def innerFunc(x):         inputs = x[0]         target = x[1]          #getting the mask by observing the model's inputs         mask = K.equal(inputs, maskVal)         mask = K.all(mask, axis=-1, keepdims=True)          #inverting the mask for getting the valid steps for each sample         mask = 1 - K.cast(mask, K.floatx())          #summing the valid steps for each sample         stepsPerSample = K.sum(mask, axis=1, keepdims=False)          #applying the mask to the target (to make sure you are summing zeros below)         target = target * mask          #calculating the mean of the steps (using our sum of valid steps as averager)         means = K.sum(target, axis=1, keepdims=False) / stepsPerSample          return means      return innerFunc   x = np.ones((2,5,3)) x[0,3:] = 0. x[1,1:] = 0.   print(x)  inputs = Input((5,3)) out = Lambda(lambda x: x*4)(inputs) out = Lambda(customPooling(0))([inputs,out])  model = Model(inputs,out) model.predict(x) 
Read More

Thursday, August 9, 2018

Distributed Tensorflow: who applies the parameter update?

Leave a Comment

I've used TensorFlow but am new to distributed TensorFlow for training models. My understanding is that current best practices favor the data parallel model with asynchronous updates:

A paper published by the Google Brain team in April 2016 benchmarked various approaches and found that data parallelism with synchronous updates using a few spare replicas was the most efficient, not only converging faster but also producing a better model. -- Chapter 12 of Hands-On Machine Learning with Scikit-Learn and Tensorflow.

Now, my confusion from reading further about this architecture is figuring out which component applies the parameter updates: the workers or the parameter server?

In my illustration below, it's clear to me that the workers compute the gradients dJ/dw (the gradient of the loss J with respect to the parameter weights w). But who applies the gradient descent update rule?

enter image description here

What's a bit confusing is that this O'Reilly article on Distributed TensorFlow states the following:

In the more centralized architecture, the devices send their output in the form of gradients to the parameter servers. These servers collect and aggregate the gradients. In synchronous training, the parameter servers compute the latest up-to-date version of the model, and send it back to devices. In asynchronous training, parameter servers send gradients to devices that locally compute the new model. In both architectures, the loop repeats until training terminates.

The above paragraph suggests that in asynchronous training:

  1. The workers compute gradients and send it to the parameter server.
  2. The parameter server broadcasts the gradients to the workers.
  3. Each worker receives the broadcasted gradients and applies the update rule.

Is my understanding correct? If it is, then that doesn't seem very asynchronous to me because the workers have to wait for the parameter server to broadcast the gradients. Any explanation would be appreciated.

1 Answers

Answers 1

Usually the parameter servers only store the global parameters and the workers directly apply their gradients to the global parameters (which are stored on the parameter servers). In asynchronous training no broadcasting takes place! The workers do the following in a loop:

  1. Get current global parameters from PS
  2. Calculate gradient
  3. Apply gradient to global parameters (after applying the gradients to variables stored on the parameter servers, tensorflow will send the gradients to the parameter servers and apply them there)

In between of the steps 1 and 3 the global parameters change because other workers apply their gradients. The applying of the gradients is usually hogwild.

In asynchronous training, parameter servers send gradients to devices

I don't think this happens in any asynchronous implementations. Don't know what the author tried to say here.

Read More

Saturday, August 4, 2018

Why is TensorFlow's `tf.data` package slowing down my code?

Leave a Comment

I'm just learning to use TensorFlow's tf.data API, and I've found that it is slowing my code down a lot, measured in time per epoch. This is the opposite of what it's supposed to do, I thought. I wrote a simple linear regression program to test it out.

Tl;Dr: With 100,000 training data, tf.data slows time per epoch down by about a factor of ten, if you're using full batch training. Worse if you use smaller batches. The opposite is true with 500 training data.

My question: What is going on? Is my implementation flawed? Other sources I've read have tf.data improving speeds by about 30%.

import tensorflow as tf  import numpy as np import timeit  import os os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2' tf.logging.set_verbosity(tf.logging.ERROR)  n_epochs = 10 input_dimensions_list = [10]  def function_to_approximate(x):     return np.dot(x, random_covector).astype(np.float32) + np.float32(.01) * np.random.randn(1,1).astype(np.float32)  def regress_without_tfData(n_epochs, input_dimension, training_inputs, training_labels):     tf.reset_default_graph()     weights = tf.get_variable("weights", initializer=np.random.randn(input_dimension, 1).astype(np.float32))      X = tf.placeholder(tf.float32, shape=(None, input_dimension), name='X')     Y = tf.placeholder(tf.float32, shape=(None, 1), name='Y')     prediction = tf.matmul(X,weights)     loss = tf.reduce_mean(tf.square(tf.subtract(prediction, Y)))     loss_op = tf.train.AdamOptimizer(.01).minimize(loss)      init = tf.global_variables_initializer()      with tf.Session() as sess:         sess.run(init)         for _ in range(n_epochs):             sess.run(loss_op, feed_dict={X: training_inputs, Y:training_labels})  def regress_with_tfData(n_epochs, input_dimension, training_inputs, training_labels, batch_size):     tf.reset_default_graph()     weights = tf.get_variable("weights", initializer=np.random.randn(input_dimension, 1).astype(np.float32))      X,Y = data_set.make_one_shot_iterator().get_next()      prediction = tf.matmul(X, weights)     loss = tf.reduce_mean(tf.square(tf.subtract(prediction, Y)))     loss_op = tf.train.AdamOptimizer(.01).minimize(loss)      init = tf.global_variables_initializer()      with tf.Session() as sess:         sess.run(init)         while True:             try:                  sess.run(loss_op)             except tf.errors.OutOfRangeError:                 break  for input_dimension in input_dimensions_list:     for data_size in [500, 100000]:          training_inputs = np.random.randn(data_size, input_dimension).astype(np.float32)         random_covector = np.random.randint(-5, 5, size=(input_dimension, 1))         training_labels = function_to_approximate(training_inputs)          print("Not using tf.data, with data size "         "{}, input dimension {} and training with "         "a full batch, it took an average of "         "{} seconds to run {} epochs.\n".             format(                 data_size,                 input_dimension,                 timeit.timeit(                     lambda: regress_without_tfData(                         n_epochs, input_dimension,                          training_inputs, training_labels                     ),                      number=3                 ),                 n_epochs))  for input_dimension in input_dimensions_list:     for data_size, batch_size in [(500, 50), (500, 500), (100000, 50), (100000, 100000)]:          training_inputs = np.random.randn(data_size, input_dimension).astype(np.float32)         random_covector = np.random.randint(-5, 5, size=(input_dimension, 1))         training_labels = function_to_approximate(training_inputs)          data_set = tf.data.Dataset.from_tensor_slices((training_inputs, training_labels))         data_set = data_set.repeat(n_epochs)         data_set = data_set.batch(batch_size)          print("Using tf.data, with data size "         "{}, and input dimension {}, and training with "         "batch size {}, it took an average of {} seconds "         "to run {} epochs.\n".             format(                 data_size,                 input_dimension,                 batch_size,                 timeit.timeit(                     lambda: regress_with_tfData(                         n_epochs, input_dimension,                          training_inputs, training_labels,                          batch_size                     ),                     number=3                 )/3,                 n_epochs             )) 

This outputs for me:

Not using tf.data, with data size 500, input dimension 10 and training with a full batch, it took an average of 0.20243382899980134 seconds to run 10 epochs.

Not using tf.data, with data size 100000, input dimension 10 and training with a full batch, it took an average of 0.2431719040000644 seconds to run 10 epochs.

Using tf.data, with data size 500, and input dimension 10, and training with batch size 50, it took an average of 0.09512088866661846 seconds to run 10 epochs.

Using tf.data, with data size 500, and input dimension 10, and training with batch size 500, it took an average of 0.07286913600000844 seconds to run 10 epochs.

Using tf.data, with data size 100000, and input dimension 10, and training with batch size 50, it took an average of 4.421892363666605 seconds to run 10 epochs.

Using tf.data, with data size 100000, and input dimension 10, and training with batch size 100000, it took an average of 2.2555197536667038 seconds to run 10 epochs.

Edit: Fixed an important issue that Fred Guth pointed out. It didn't much affect the results, though.

1 Answers

Answers 1

First:

You are recreating the dataset unnecessarily.

data_set = tf.data.Dataset.from_tensor_slices((training_inputs, training_labels))

Create the dataset prior to the loop and change the regress_with_tfData input signature to use dataset instead of training_inputs and training_labels.

Second:

The problem here is that minibatches of size 50 or even 500 are too small to compensate the cost of td.data building latency. You should increase the minibatch size. Interestingly you did so with a minibatch of size 100000, but then maybe it is too big ( I am not certain of this, I think it would need more tests).

There are a couple of things you could try:

1) Increase the minibatch size to something like 10000 and see if you get an improvement 2) Change your pipeline to use an iterator, example:

    data_set = tf.data.Dataset.from_tensor_slices((training_inputs, training_labels))     data_set = data_set.repeat(n_epochs)     data_set = data_set.batch(batch_size)     iterator = data_set.make_one_shot_iterator()     ....     next_element = iterator.get_next() 
Read More

Monday, July 30, 2018

How to forecast using the Tensorflow model?

Leave a Comment

I have created tensorflow program in order to for the close prices of the forex. I have successfully created the predcitions but failed understand the way to forecast the values for the future. See the following is my prediction function:

test_pred_list = []  def testAndforecast(xTest1,yTest1): #     test_pred_list = 0     truncated_backprop_length = 3     with tf.Session() as sess:     #     train_writer = tf.summary.FileWriter('logs', sess.graph)         tf.global_variables_initializer().run()         counter = 0 #         saver.restore(sess, "models\\model2298.ckpt")         try:             with open ("Checkpointcounter.txt","r") as file:                 value = file.read()         except FileNotFoundError:             print("First Time Running Training!....")           if(tf.train.checkpoint_exists("models\\model"+value+".ckpt")):             saver.restore(sess, "models\\model"+value+".ckpt")             print("models\\model"+value+".ckpt Session Loaded for Testing")         for test_idx in range(len(xTest1) - truncated_backprop_length):              testBatchX = xTest1[test_idx:test_idx+truncated_backprop_length,:].reshape((1,truncated_backprop_length,num_features))                     testBatchY = yTest1[test_idx:test_idx+truncated_backprop_length].reshape((1,truncated_backprop_length,1))               #_current_state = np.zeros((batch_size,state_size))             feed = {batchX_placeholder : testBatchX,                 batchY_placeholder : testBatchY}              #Test_pred contains 'window_size' predictions, we want the last one             _last_state,_last_label,test_pred = sess.run([last_state,last_label,prediction],feed_dict=feed)             test_pred_list.append(test_pred[-1][-1]) #The last one 

Here is the complete jupyter and datasets for test and train:
My repository with code.

Kindly, help me how I can forecast the close values for the future. Please do not share something related to predictions as I have tried. Kindly, let me know something that will forecast without any support just on the basis of training what I have given.

I hope to hear soon.

1 Answers

Answers 1

If I understand your question correctly, by forecasting you mean predicting multiple closing prices in future (for example next 5 closing prices from current state). I went through your jupyter notebook. In short, you can not easily do that.

Right now your code takes the last three positions defined by multiple futures (open/low/high/close prices and some indicators values). Based on that you predict next closing price. If you would like to predict even further position, you would have to create an "artificial" position based on the predicted closing price. Here you can approximate that open price is same as previous closing, but you can only guess high and low prices. Then you would calculate other futures/values (from indicators) and use this position with previous two to predict next closing price. You can continue like this for future steps.

The issue is in the open/low/high prices because you can only approximate them. You could remove them from data, retrain the model, and make predictions without them, but they may be necessary for indicators calculations.


I somehow compressed your code here to show the approach of predicting all OHLC prices:

# Data xTrain = datasetTrain[     ["open", "high", "low", "close", "k",      "d", "atr", "macdmain", "macdsgnal",      "bbup", "bbmid", "bblow"]].as_matrix() yTrain = datasetTrain[["open", "high", "low", "close"]].as_matrix()  # Settings batch_size = 1 num_batches = 1000 truncated_backprop_length = 3 state_size = 12  num_features = 12 num_classes = 4  # Graph batchX_placeholder = tf.placeholder(     dtype=tf.float32,     shape=[None, truncated_backprop_length, num_features],     name='data_ph') batchY_placeholder = tf.placeholder(     dtype=tf.float32,     shape=[None, num_classes],     name='target_ph')   cell = tf.contrib.rnn.BasicRNNCell(num_units=state_size) states_series, current_state = tf.nn.dynamic_rnn(     cell=cell,     inputs=batchX_placeholder,     dtype=tf.float32)  states_series = tf.transpose(states_series, [1,0,2])  last_state = tf.gather(     params=states_series,     indices=states_series.get_shape()[0]-1)  weight = tf.Variable(tf.truncated_normal([state_size, num_classes])) bias = tf.Variable(tf.constant(0.1, shape=[num_classes]))  prediction = tf.matmul(last_state, weight) + bias   loss = tf.reduce_mean(tf.squared_difference(last_label, prediction)) train_step = tf.train.AdamOptimizer(learning_rate=0.001).minimize(loss)  # Training for batch_idx in range(num_batches):     start_idx = batch_idx     end_idx = start_idx + truncated_backprop_length       batchX = xTrain[start_idx:end_idx,:].reshape(batch_size, truncated_backprop_length, num_features)     batchY = yTrain[end_idx].reshape(batch_size, truncated_backprop_length, num_classes)       feed = {batchX_placeholder: batchX, batchY_placeholder: batchY}      _loss, _train_step, _pred, _last_label,_prediction = sess.run(         fetches=[loss, train_step, prediction, last_label, prediction],         feed_dict=feed) 

I think it is not important to write the whole code plus I don't know how are the indicators calculated. Also you should change way of data feeding because right now it only works with batches os size 1.

Read More

How to quantize all nodes except a particular one?

Leave a Comment

I am using tensorflow Graph Transform Tool to quantize the graph using

input_names = ["prefix/input"] output_names = ["final_result"]   transforms1 = ["strip_unused_nodes","fold_constants(ignore_errors=true)",  "fold_batch_norms",  "fold_old_batch_norms","quantize_weights" ]  transformed_graph_def = TransformGraph(graph.as_graph_def(), input_names,output_names, transforms1) 

I use the option quantize_weights to quantize the weights in graph, I know that certain nodes can remain unquantized by changing threshold minimum_size in quantize_weights, so leaving some nodes unquantized is certainly possible.

I want to quantize the weights of all nodes except a particular node with name K or a set of nodes that have name in K(set). How can this be achieved ?

0 Answers

Read More

Friday, July 27, 2018

How to improve accuracy of Tensorflow camera demo on iOS for retrained graph

Leave a Comment

I have an Android app that was modeled after the Tensorflow Android demo for classifying images,

https://github.com/tensorflow/tensorflow/tree/master/tensorflow/examples/android

The original app uses a tensorflow graph (.pb) file to classify a generic set of images from Inception v3 (I think)

I then trained my own graph for my own images following the instruction in Tensorflow for Poets blog,

https://petewarden.com/2016/02/28/tensorflow-for-poets/

and this worked in the Android app very well, after changing the settings in,

ClassifierActivity

private static final int INPUT_SIZE = 299; private static final int IMAGE_MEAN = 128; private static final float IMAGE_STD = 128.0f; private static final String INPUT_NAME = "Mul"; private static final String OUTPUT_NAME = "final_result"; private static final String MODEL_FILE = "file:///android_asset/optimized_graph.pb"; private static final String LABEL_FILE =  "file:///android_asset/retrained_labels.txt"; 

To port the app to iOS, I then used the iOS camera demo, https://github.com/tensorflow/tensorflow/tree/master/tensorflow/examples/ios/camera

and used the same graph file and changed the settings in,

CameraExampleViewController.mm

// If you have your own model, modify this to the file name, and make sure // you've added the file to your app resources too. static NSString* model_file_name = @"tensorflow_inception_graph"; static NSString* model_file_type = @"pb"; // This controls whether we'll be loading a plain GraphDef proto, or a // file created by the convert_graphdef_memmapped_format utility that wraps a // GraphDef and parameter file that can be mapped into memory from file to // reduce overall memory usage. const bool model_uses_memory_mapping = false; // If you have your own model, point this to the labels file. static NSString* labels_file_name = @"imagenet_comp_graph_label_strings"; static NSString* labels_file_type = @"txt"; // These dimensions need to match those the model was trained with. const int wanted_input_width = 299; const int wanted_input_height = 299; const int wanted_input_channels = 3; const float input_mean = 128f; const float input_std = 128.0f; const std::string input_layer_name = "Mul"; const std::string output_layer_name = "final_result"; 

After this the app is working on iOS, however...

The app on Android performs much better than iOS in detecting classified images. If I fill the camera's view port with the image, both perform similar. But normally the image to detect is only part of the camera view port, on Android this doesn't seem to impact much, but on iOS it impacts a lot, so iOS cannot classify the image.

My guess is that Android is cropping if camera view port to the central 299x299 area, where as iOS is scaling its camera view port to the central 299x299 area.

Can anyone confirm this? and does anyone know how to fix the iOS demo to better detect focused images? (make it crop)

In the demo Android class,

ClassifierActivity.onPreviewSizeChosen()

rgbFrameBitmap = Bitmap.createBitmap(previewWidth, previewHeight, Config.ARGB_8888);     croppedBitmap = Bitmap.createBitmap(INPUT_SIZE, INPUT_SIZE, Config.ARGB_8888);  frameToCropTransform =         ImageUtils.getTransformationMatrix(             previewWidth, previewHeight,             INPUT_SIZE, INPUT_SIZE,             sensorOrientation, MAINTAIN_ASPECT);  cropToFrameTransform = new Matrix(); frameToCropTransform.invert(cropToFrameTransform); 

and on iOS is has,

CameraExampleViewController.runCNNOnFrame()

const int sourceRowBytes = (int)CVPixelBufferGetBytesPerRow(pixelBuffer);   const int image_width = (int)CVPixelBufferGetWidth(pixelBuffer);   const int fullHeight = (int)CVPixelBufferGetHeight(pixelBuffer);    CVPixelBufferLockFlags unlockFlags = kNilOptions;   CVPixelBufferLockBaseAddress(pixelBuffer, unlockFlags);    unsigned char *sourceBaseAddr =       (unsigned char *)(CVPixelBufferGetBaseAddress(pixelBuffer));   int image_height;   unsigned char *sourceStartAddr;   if (fullHeight <= image_width) {     image_height = fullHeight;     sourceStartAddr = sourceBaseAddr;   } else {     image_height = image_width;     const int marginY = ((fullHeight - image_width) / 2);     sourceStartAddr = (sourceBaseAddr + (marginY * sourceRowBytes));   }   const int image_channels = 4;    assert(image_channels >= wanted_input_channels);   tensorflow::Tensor image_tensor(       tensorflow::DT_FLOAT,       tensorflow::TensorShape(           {1, wanted_input_height, wanted_input_width, wanted_input_channels}));   auto image_tensor_mapped = image_tensor.tensor<float, 4>();   tensorflow::uint8 *in = sourceStartAddr;   float *out = image_tensor_mapped.data();   for (int y = 0; y < wanted_input_height; ++y) {     float *out_row = out + (y * wanted_input_width * wanted_input_channels);     for (int x = 0; x < wanted_input_width; ++x) {       const int in_x = (y * image_width) / wanted_input_width;       const int in_y = (x * image_height) / wanted_input_height;       tensorflow::uint8 *in_pixel =           in + (in_y * image_width * image_channels) + (in_x * image_channels);       float *out_pixel = out_row + (x * wanted_input_channels);       for (int c = 0; c < wanted_input_channels; ++c) {         out_pixel[c] = (in_pixel[c] - input_mean) / input_std;       }     }   }    CVPixelBufferUnlockBaseAddress(pixelBuffer, unlockFlags); 

I think the issue is here,

tensorflow::uint8 *in_pixel =           in + (in_y * image_width * image_channels) + (in_x * image_channels);       float *out_pixel = out_row + (x * wanted_input_channels); 

My understanding is this is just scaling to the 299 size by pick every xth pixel instead of scaling the original image to the 299 size. So this leads to poor scaling and poor image recognition.

The solution is to first scale to pixelBuffer to size 299. I tried this,

UIImage *uiImage = [self uiImageFromPixelBuffer: pixelBuffer]; float scaleFactor = (float)wanted_input_height / (float)fullHeight; float newWidth = image_width * scaleFactor; NSLog(@"width: %d, height: %d, scale: %f, height: %f", image_width, fullHeight, scaleFactor, newWidth); CGSize size = CGSizeMake(wanted_input_width, wanted_input_height); UIGraphicsBeginImageContext(size); [uiImage drawInRect:CGRectMake(0, 0, newWidth, size.height)]; UIImage *destImage = UIGraphicsGetImageFromCurrentImageContext(); UIGraphicsEndImageContext(); pixelBuffer = [self pixelBufferFromCGImage: destImage.CGImage]; 

and to convert image to pixle buffer,

- (CVPixelBufferRef) pixelBufferFromCGImage: (CGImageRef) image {     NSDictionary *options = @{                               (NSString*)kCVPixelBufferCGImageCompatibilityKey : @YES,                               (NSString*)kCVPixelBufferCGBitmapContextCompatibilityKey : @YES,                               };      CVPixelBufferRef pxbuffer = NULL;     CVReturn status = CVPixelBufferCreate(kCFAllocatorDefault, CGImageGetWidth(image),                                           CGImageGetHeight(image), kCVPixelFormatType_32ARGB, (__bridge CFDictionaryRef) options,                                           &pxbuffer);     if (status!=kCVReturnSuccess) {         NSLog(@"Operation failed");     }     NSParameterAssert(status == kCVReturnSuccess && pxbuffer != NULL);      CVPixelBufferLockBaseAddress(pxbuffer, 0);     void *pxdata = CVPixelBufferGetBaseAddress(pxbuffer);      CGColorSpaceRef rgbColorSpace = CGColorSpaceCreateDeviceRGB();     CGContextRef context = CGBitmapContextCreate(pxdata, CGImageGetWidth(image),                                                  CGImageGetHeight(image), 8, 4*CGImageGetWidth(image), rgbColorSpace,                                                  kCGImageAlphaNoneSkipFirst);     NSParameterAssert(context);      CGContextConcatCTM(context, CGAffineTransformMakeRotation(0));     CGAffineTransform flipVertical = CGAffineTransformMake( 1, 0, 0, -1, 0, CGImageGetHeight(image) );     CGContextConcatCTM(context, flipVertical);     CGAffineTransform flipHorizontal = CGAffineTransformMake( -1.0, 0.0, 0.0, 1.0, CGImageGetWidth(image), 0.0 );     CGContextConcatCTM(context, flipHorizontal);      CGContextDrawImage(context, CGRectMake(0, 0, CGImageGetWidth(image),                                            CGImageGetHeight(image)), image);     CGColorSpaceRelease(rgbColorSpace);     CGContextRelease(context);      CVPixelBufferUnlockBaseAddress(pxbuffer, 0);     return pxbuffer; }  - (UIImage*) uiImageFromPixelBuffer: (CVPixelBufferRef) pixelBuffer {     CIImage *ciImage = [CIImage imageWithCVPixelBuffer: pixelBuffer];      CIContext *temporaryContext = [CIContext contextWithOptions:nil];     CGImageRef videoImage = [temporaryContext                              createCGImage:ciImage                              fromRect:CGRectMake(0, 0,                                                  CVPixelBufferGetWidth(pixelBuffer),                                                  CVPixelBufferGetHeight(pixelBuffer))];      UIImage *uiImage = [UIImage imageWithCGImage:videoImage];     CGImageRelease(videoImage);     return uiImage; } 

Not sure if this is the best way to resize, but this worked. But it seemed to make image classification even worse, not better...

Any ideas, or issues with the image conversion/resize?

3 Answers

Answers 1

Since you are not using YOLO Detector the MAINTAIN_ASPECT flag is set to false. Hence the image on Android app is not getting cropped, but it's scaled. However, in the code snippet provided I don't see the actual initialisation of the flag. Confirm that the value of the flag is actually false in your app.

I know this isn't a complete solution but hope this helps you in debugging the issue.

Answers 2

Please change at this code:

// If you have your own model, modify this to the file name, and make sure // you've added the file to your app resources too. static NSString* model_file_name = @"tensorflow_inception_graph"; static NSString* model_file_type = @"pb"; // This controls whether we'll be loading a plain GraphDef proto, or a // file created by the convert_graphdef_memmapped_format utility that wraps a // GraphDef and parameter file that can be mapped into memory from file to // reduce overall memory usage. const bool model_uses_memory_mapping = false; // If you have your own model, point this to the labels file. static NSString* labels_file_name = @"imagenet_comp_graph_label_strings"; static NSString* labels_file_type = @"txt"; // These dimensions need to match those the model was trained with. const int wanted_input_width = 299; const int wanted_input_height = 299; const int wanted_input_channels = 3; const float input_mean = 128f; const float input_std = 1.0f; const std::string input_layer_name = "Mul"; const std::string output_layer_name = "final_result"; 

Here change : const float input_std = 1.0f;

Answers 3

Tensorflow Object detection have default and standard configurations, below is the list of settings,

Important things you need to check based on your input ML model,

-> model_file_name - This according to your .pb file name,

-> model_uses_memory_mapping - It's up to you to reduce overall memory usage.

-> labels_file_name - This varies based on our label file name,

-> input_layer_name/output_layer_name - Make sure you are using your own layer input/output names which you are using during graph(.pb) file creation.

snippet:

// If you have your own model, modify this to the file name, and make sure // you've added the file to your app resources too. static NSString* model_file_name = @"graph";//@"tensorflow_inception_graph"; static NSString* model_file_type = @"pb"; // This controls whether we'll be loading a plain GraphDef proto, or a // file created by the convert_graphdef_memmapped_format utility that wraps a // GraphDef and parameter file that can be mapped into memory from file to // reduce overall memory usage. const bool model_uses_memory_mapping = true; // If you have your own model, point this to the labels file. static NSString* labels_file_name = @"labels";//@"imagenet_comp_graph_label_strings"; static NSString* labels_file_type = @"txt"; // These dimensions need to match those the model was trained with. const int wanted_input_width = 224; const int wanted_input_height = 224; const int wanted_input_channels = 3; const float input_mean = 117.0f; const float input_std = 1.0f; const std::string input_layer_name = "input"; const std::string output_layer_name = "final_result"; 

Custom Image Tensorflow detection, you can use below working snippet:

-> For this process you just need to pass the UIImage.CGImage object,

NSString* RunInferenceOnImageResult(CGImageRef image) {     tensorflow::SessionOptions options;      tensorflow::Session* session_pointer = nullptr;     tensorflow::Status session_status = tensorflow::NewSession(options, &session_pointer);     if (!session_status.ok()) {         std::string status_string = session_status.ToString();         return [NSString stringWithFormat: @"Session create failed - %s",                 status_string.c_str()];     }     std::unique_ptr<tensorflow::Session> session(session_pointer);     LOG(INFO) << "Session created.";      tensorflow::GraphDef tensorflow_graph;     LOG(INFO) << "Graph created.";      NSString* network_path = FilePathForResourceNames(@"tensorflow_inception_graph", @"pb");     PortableReadFileToProtol([network_path UTF8String], &tensorflow_graph);      LOG(INFO) << "Creating session.";     tensorflow::Status s = session->Create(tensorflow_graph);     if (!s.ok()) {         LOG(ERROR) << "Could not create TensorFlow Graph: " << s;         return @"";     }      // Read the label list     NSString* labels_path = FilePathForResourceNames(@"imagenet_comp_graph_label_strings", @"txt");     std::vector<std::string> label_strings;     std::ifstream t;     t.open([labels_path UTF8String]);     std::string line;     while(t){         std::getline(t, line);         label_strings.push_back(line);     }     t.close();      // Read the Grace Hopper image.     //NSString* image_path = FilePathForResourceNames(@"grace_hopper", @"jpg");     int image_width;     int image_height;     int image_channels; //    std::vector<tensorflow::uint8> image_data = LoadImageFromFile( //                                                                  [image_path UTF8String], &image_width, &image_height, &image_channels);     std::vector<tensorflow::uint8> image_data = LoadImageFromImage(image,&image_width, &image_height, &image_channels);     const int wanted_width = 224;     const int wanted_height = 224;     const int wanted_channels = 3;     const float input_mean = 117.0f;     const float input_std = 1.0f;     assert(image_channels >= wanted_channels);     tensorflow::Tensor image_tensor(                                     tensorflow::DT_FLOAT,                                     tensorflow::TensorShape({         1, wanted_height, wanted_width, wanted_channels}));     auto image_tensor_mapped = image_tensor.tensor<float, 4>();     tensorflow::uint8* in = image_data.data();     // tensorflow::uint8* in_end = (in + (image_height * image_width * image_channels));     float* out = image_tensor_mapped.data();     for (int y = 0; y < wanted_height; ++y) {         const int in_y = (y * image_height) / wanted_height;         tensorflow::uint8* in_row = in + (in_y * image_width * image_channels);         float* out_row = out + (y * wanted_width * wanted_channels);         for (int x = 0; x < wanted_width; ++x) {             const int in_x = (x * image_width) / wanted_width;             tensorflow::uint8* in_pixel = in_row + (in_x * image_channels);             float* out_pixel = out_row + (x * wanted_channels);             for (int c = 0; c < wanted_channels; ++c) {                 out_pixel[c] = (in_pixel[c] - input_mean) / input_std;             }         }     }      NSString* result; //    result = [NSString stringWithFormat: @"%@ - %lu, %s - %dx%d", result, //              label_strings.size(), label_strings[0].c_str(), image_width, image_height];      std::string input_layer = "input";     std::string output_layer = "output";     std::vector<tensorflow::Tensor> outputs;     tensorflow::Status run_status = session->Run({{input_layer, image_tensor}},                                                  {output_layer}, {}, &outputs);     if (!run_status.ok()) {         LOG(ERROR) << "Running model failed: " << run_status;         tensorflow::LogAllRegisteredKernels();         result = @"Error running model";         return result;     }     tensorflow::string status_string = run_status.ToString();     result = [NSString stringWithFormat: @"Status :%s\n",               status_string.c_str()];      tensorflow::Tensor* output = &outputs[0];     const int kNumResults = 5;     const float kThreshold = 0.1f;     std::vector<std::pair<float, int> > top_results;     GetTopN(output->flat<float>(), kNumResults, kThreshold, &top_results);      std::stringstream ss;     ss.precision(3);     for (const auto& result : top_results) {         const float confidence = result.first;         const int index = result.second;          ss << index << " " << confidence << "  ";          // Write out the result as a string         if (index < label_strings.size()) {             // just for safety: theoretically, the output is under 1000 unless there             // is some numerical issues leading to a wrong prediction.             ss << label_strings[index];         } else {             ss << "Prediction: " << index;         }          ss << "\n";     }      LOG(INFO) << "Predictions: " << ss.str();      tensorflow::string predictions = ss.str();     result = [NSString stringWithFormat: @"%@ - %s", result,               predictions.c_str()];      return result; } 

Scaling Image for custom width and height - C++ code snippet,

std::vector<uint8> LoadImageFromImage(CGImageRef image,                                      int* out_width, int* out_height,                                      int* out_channels) {      const int width = (int)CGImageGetWidth(image);     const int height = (int)CGImageGetHeight(image);     const int channels = 4;     CGColorSpaceRef color_space = CGColorSpaceCreateDeviceRGB();     const int bytes_per_row = (width * channels);     const int bytes_in_image = (bytes_per_row * height);     std::vector<uint8> result(bytes_in_image);     const int bits_per_component = 8;     CGContextRef context = CGBitmapContextCreate(result.data(), width, height,                                                  bits_per_component, bytes_per_row, color_space,                                                  kCGImageAlphaPremultipliedLast | kCGBitmapByteOrder32Big);     CGColorSpaceRelease(color_space);     CGContextDrawImage(context, CGRectMake(0, 0, width, height), image);     CGContextRelease(context);     CFRelease(image);      *out_width = width;     *out_height = height;     *out_channels = channels;     return result; } 

Above function helps you to load the image data based on your custom ratio. High accurate image pixel ratio for both Width and height during tensorflow classification is 224 x 224.

You need to call above LoadImage function from RunInferenceOnImageResult, with actual custom width and height arguments along with Image reference.

Read More

Monday, July 23, 2018

Tensorflow Custom TFLite java.lang.NullPointerException: Can not allocate memory for the interpreter

Leave a Comment

I have created a custom tensorflow lite model using retrain.py from https://github.com/tensorflow/hub/blob/master/examples/image_retraining/retrain.py using the following command

python retrain.py --image_dir newImageDirectory --tfhub_module https://tfhub.dev/google/imagenet/mobilenet_v2_100_224/feature_vector/1 

Then I convert using toco the output_graph.pb file to a lite file. Using the below command

bazel run tensorflow/contrib/lite/toco:toco -- --input_file=/tmp/output_graph.pb --output_file=/tmp/optimized.lite --input_format=TENSORFLOW_GRAPHDEF --output_format=TFLITE --inpute_shape=1,224,224,3 --input_array=input --output_array=final_result --inference_type=FLOAT --input_data_type=FLOAT 

Then I take the new lite file and the labels.txt file and put them in tensorflow for poets 2 https://github.com/googlecodelabs/tensorflow-for-poets-2 to see if I can have it start to classify new categories. When the application launches I receive the following error.

Caused by: java.lang.NullPointerException: Can not allocate memory for the interpreter                                                                                            at org.tensorflow.lite.NativeInterpreterWrapper.createInterpreter(Native Method)                                                                       at org.tensorflow.lite.NativeInterpreterWrapper.<init>(NativeInterpreterWrapper.java:63)                                                                                         at org.tensorflow.lite.NativeInterpreterWrapper.<init>(NativeInterpreterWrapper.java:51)                                                                                         at org.tensorflow.lite.Interpreter.<init>(Interpreter.java:90)                                                                                         at com.example.android.tflitecamerademo.ImageClassifier.<init>(ImageClassifier.java:97) 

0 Answers

Read More

Friday, July 20, 2018

Error Trying to Convert TensorFlow Saved Model to TensorFlow.js Model

Leave a Comment

I have successfully trained a DNNClassifier to classify texts (posts from an online discussion board). I've created and saved my model using this code:

embedded_text_feature_column = hub.text_embedding_column(     key="sentence",     module_spec="https://tfhub.dev/google/nnlm-de-dim128/1") feature_columns = [embedded_text_feature_column] estimator = tf.estimator.DNNClassifier(     hidden_units=[500, 100],     feature_columns=feature_columns,     n_classes=2,     optimizer=tf.train.AdagradOptimizer(learning_rate=0.003)) feature_spec = tf.feature_column.make_parse_example_spec(feature_columns) serving_input_receiver_fn = tf.estimator.export.build_parsing_serving_input_receiver_fn(feature_spec) estimator.export_savedmodel(export_dir_base="/my/dir/base", serving_input_receiver_fn=serving_input_receiver_fn) 

Now I want to convert my saved model to use it with the JavaScript version of TensorFlow, tf.js, using the tfjs-converter.

When I issue the following command:

tensorflowjs_converter --input_format=tf_saved_model --output_node_names='dnn/head/predictions/str_classes,dnn/head/predictions/probabilities' --saved_model_tags=serve /my/dir/base /my/export/dir 

…I get this error message:

ValueError: Node 'dnn/input_from_feature_columns/input_layer/sentence_hub_module_embedding/module_apply_default/embedding_lookup_sparse/embedding_lookup' expects to be colocated with unknown node 'dnn/input_from_feature_columns/input_layer/sentence_hub_module_embedding

I assume I'm doing something wrong when saving the model.

What is the correct way to save an estimator model so that it can be converted with tfjs-converter?

The source code of my project can be found on GitHub.

0 Answers

Read More

Monday, July 16, 2018

How to Do a Simple CLI Query for a Saved Estimator Model?

Leave a Comment

I have successfully trained a DNNClassifier to classify texts (posts from an online discussion board). I've saved the model and I now want to classify texts using the TensorFlow CLI.

When I run saved_model_cli show for my saved model, I get this output:

saved_model_cli show --dir /my/model --tag_set serve --signature_def predict The given SavedModel SignatureDef contains the following input(s):   inputs['examples'] tensor_info:       dtype: DT_STRING       shape: (-1)       name: input_example_tensor:0 The given SavedModel SignatureDef contains the following output(s):   outputs['class_ids'] tensor_info:       dtype: DT_INT64       shape: (-1, 1)       name: dnn/head/predictions/ExpandDims:0   outputs['classes'] tensor_info:       dtype: DT_STRING       shape: (-1, 1)       name: dnn/head/predictions/str_classes:0   outputs['logistic'] tensor_info:       dtype: DT_FLOAT       shape: (-1, 1)       name: dnn/head/predictions/logistic:0   outputs['logits'] tensor_info:       dtype: DT_FLOAT       shape: (-1, 1)       name: dnn/logits/BiasAdd:0   outputs['probabilities'] tensor_info:       dtype: DT_FLOAT       shape: (-1, 2)       name: dnn/head/predictions/probabilities:0 Method name is: tensorflow/serving/predict 

I cannot figure out the correct parameters for saved_model_cli run to get a prediction.

I have tried several approaches, for example:

saved_model_cli run --dir /my/model --tag_set serve --signature_def predict --input_exprs='examples=["klassifiziere mich bitte"]' 

Which gives me this error message:

InvalidArgumentError (see above for traceback): Could not parse example input, value: 'klassifiziere mich bitte'  [[Node: ParseExample/ParseExample = ParseExample[Ndense=1, Nsparse=0, Tdense=[DT_STRING], dense_shapes=[[1]], sparse_types=[], _device="/job:localhost/replica:0/task:0/device:CPU:0"](_arg_input_example_tensor_0_0, ParseExample/ParseExample/names, ParseExample/ParseExample/dense_keys_0, ParseExample/ParseExample/names)]] 

What is the correct way to pass my input string to the CLI to get a classification?

You can find the code of my project, including the training data, on GitHub: https://github.com/pahund/beitragstuev

I'm building and saving my model like this (simplified, see GitHub for original code):

embedded_text_feature_column = hub.text_embedding_column(     key="sentence",     module_spec="https://tfhub.dev/google/nnlm-de-dim128/1") feature_columns = [embedded_text_feature_column] estimator = tf.estimator.DNNClassifier(     hidden_units=[500, 100],     feature_columns=feature_columns,     n_classes=2,     optimizer=tf.train.AdagradOptimizer(learning_rate=0.003)) feature_spec = tf.feature_column.make_parse_example_spec(feature_columns) serving_input_receiver_fn = tf.estimator.export.build_parsing_serving_input_receiver_fn(feature_spec) estimator.export_savedmodel(export_dir_base="/my/dir/base", serving_input_receiver_fn=serving_input_receiver_fn) 

1 Answers

Answers 1

The ServingInputReceiver you're creating for the model export is is telling the saved model to expect serialized tf.Example protos instead of the raw strings you wish to classify.

From the Save and Restore documentation:

A typical pattern is that inference requests arrive in the form of serialized tf.Examples, so the serving_input_receiver_fn() creates a single string placeholder to receive them. The serving_input_receiver_fn() is then also responsible for parsing the tf.Examples by adding a tf.parse_example op to the graph.

....

The tf.estimator.export.build_parsing_serving_input_receiver_fn utility function provides that input receiver for the common case.

So your exported model contains a tf.parse_example op that expects to receive serialized tf.Example protos satisfying the feature specification you passed to build_parsing_serving_input_receiver_fn, i.e. in your case it expects serialized examples that have the sentence feature. To predict with the model, you have to provide those serialized protos.

Fortunately, Tensorflow makes it fairly easy to construct these. Here's one possible function to return an expression mapping the examples input key to a batch of strings, which you can then pass to the CLI:

import tensorflow as tf  def serialize_example_string(strings):    serialized_examples = []   for s in strings:     try:       value = [bytes(s, "utf-8")]     except TypeError:  # python 2       value = [bytes(s)]      example = tf.train.Example(                 features=tf.train.Features(                   feature={                     "sentence": tf.train.Feature(bytes_list=tf.train.BytesList(value=value))                   }                 )               )     serialized_examples.append(example.SerializeToString())    return "examples=" + repr(serialized_examples).replace("'", "\"") 

So using some strings pulled from your examples:

strings = ["klassifiziere mich bitte",            "Das Paket „S Line Competition“ umfasst unter anderem optische Details, eine neue Farbe (Turboblau), 19-Zöller und LED-Lampen.",            "(pro Stimme geht 1 Euro Spende von Pfuscher ans Forum) ah du sack, also so gehts ja net :D:D:D"]  print (serialize_example_string(strings)) 

the CLI command would be:

saved_model_cli run --dir /path/to/model --tag_set serve --signature_def predict --input_exprs='examples=[b"\n*\n(\n\x08sentence\x12\x1c\n\x1a\n\x18klassifiziere mich bitte", b"\n\x98\x01\n\x95\x01\n\x08sentence\x12\x88\x01\n\x85\x01\n\x82\x01Das Paket \xe2\x80\x9eS Line Competition\xe2\x80\x9c umfasst unter anderem optische Details, eine neue Farbe (Turboblau), 19-Z\xc3\xb6ller und LED-Lampen.", b"\np\nn\n\x08sentence\x12b\n`\n^(pro Stimme geht 1 Euro Spende von Pfuscher ans Forum) ah du sack, also so gehts ja net :D:D:D"]' 

which should give you the desired results:

Result for output key class_ids: [[0]  [1]  [0]] Result for output key classes: [[b'0']  [b'1']  [b'0']] Result for output key logistic: [[0.05852016]  [0.88453305]  [0.04373989]] Result for output key logits: [[-2.7780817]  [ 2.0360758]  [-3.0847695]] Result for output key probabilities: [[0.94147986 0.05852016]  [0.11546692 0.88453305]  [0.9562601  0.04373989]] 
Read More