{
  "id": 71760,
  "title": "Multi GPU experience",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/71760",
  "author_name": "beluga",
  "post_date": "2018-11-16T08:42:03.466000",
  "votes": 4,
  "comment_count": 15,
  "views": 0,
  "content": "<p>What setup do you use for multi GPU training?</p>\n\n<p>I tried the default recommended tf/keras model-&gt; multi_gpu_model way. It was not really faster.\nI rent GCP GPUs so I have to find a good cost&amp;performance solution</p>",
  "messages": [
    {
      "id": 422467,
      "postDate": "2018-11-16T09:00:53.350Z",
      "content": "<p>I'm using keras, multiGPU on GCP. It works fine for me.    Two GPUs can double my batch_size.</p>\n\n<p>A few days ago, I solved my performance issue. The GPU usage is usually 0%, and one CPU process is very busy. After using keras data generator with multi_processing, more CPUs are reading csv and rendering pictures,  then the GPU is much busier than before.</p>\n\n<p>By the way, I set workers=8 in the parallel_model.fit_generator, this can keep 10 CPU and 2 P100 GPU busy most of the time.</p>\n\n<pre><code>parallel_model.fit_generator(\n    tr_generator, \n    steps_per_epoch=STEPS, \n    epochs=1, \n    verbose=1,\n    validation_data=va_generator,\n    validation_steps=STEPS_VALID,\n    callbacks = callbacks,\n    use_multiprocessing=True,\n    workers=8\n)\n</code></pre>",
      "rawMarkdown": "I'm using keras, multiGPU on GCP. It works fine for me.    Two GPUs can double my batch_size.\n\nA few days ago, I solved my performance issue. The GPU usage is usually 0%, and one CPU process is very busy. After using keras data generator with multi_processing, more CPUs are reading csv and rendering pictures,  then the GPU is much busier than before.\n\nBy the way, I set workers=8 in the parallel_model.fit_generator, this can keep 10 CPU and 2 P100 GPU busy most of the time.\n\n    parallel_model.fit_generator(\n        tr_generator, \n        steps_per_epoch=STEPS, \n        epochs=1, \n        verbose=1,\n        validation_data=va_generator,\n        validation_steps=STEPS_VALID,\n        callbacks = callbacks,\n        use_multiprocessing=True,\n        workers=8\n    )",
      "votes": 5,
      "replies": [
        {
          "id": 422620,
          "postDate": "2018-11-16T14:13:24.213Z",
          "content": "<p>Is your <code>tr_generator</code>  an instance of <code>keras.utils.Sequence</code> ?  There will exist a warning about \"data copy\" in my model and be slower when I keep <code>use_multiprocessing=True</code>,  did you have the same warning? </p>",
          "rawMarkdown": "Is your `tr_generator`  an instance of `keras.utils.Sequence` ?  There will exist a warning about \"data copy\" in my model and be slower when I keep `use_multiprocessing=True`,  did you have the same warning? "
        },
        {
          "id": 422713,
          "postDate": "2018-11-16T16:57:56.043Z",
          "content": "<p>yes, it's a keras.utils.Sequence.  no warning about \"data copy\".</p>",
          "rawMarkdown": "yes, it's a keras.utils.Sequence.  no warning about \"data copy\"."
        },
        {
          "id": 423061,
          "postDate": "2018-11-17T11:25:51.153Z",
          "content": "<p>Did you use the shuffled csv to generate training data ? I don't know how to implement the <code>Sequence</code> method <code>__getitem__</code> which can generate the <code>idx</code>th batch data randomly like this: <br>\n<code>\ndef __getitem__(self, idx): <br>\n        batch_x = .... <br>\n        batch_y = .... <br>\n        return batch_x, batch_y\n</code> <br>\nSo, I implement a generator which can only generate sequential batched of data: <br>\n<code>\nfor df in pd.read_csv(filename, chunksize=batchsize): <br>\n                batch_x = df_to_gray_image_array(...) <br>\n                batch_y = keras.utils.to_categorical(df.y, num_classes=NCATS) <br>\n                yield x, y\n</code> <br>\nCould you tell me how to implement <code>__getitem__</code> which can generate batches randomly? Thanks.</p>",
          "rawMarkdown": "Did you use the shuffled csv to generate training data ? I don't know how to implement the `Sequence` method `__getitem__` which can generate the `idx`th batch data randomly like this:    \n```\ndef __getitem__(self, idx):      \n        batch_x = ....    \n        batch_y = ....    \n        return batch_x, batch_y\n```     \nSo, I implement a generator which can only generate sequential batched of data:   \n```\nfor df in pd.read_csv(filename, chunksize=batchsize):      \n                batch_x = df_to_gray_image_array(...)    \n                batch_y = keras.utils.to_categorical(df.y, num_classes=NCATS)    \n                yield x, y\n```   \nCould you tell me how to implement `__getitem__` which can generate batches randomly? Thanks."
        },
        {
          "id": 423222,
          "postDate": "2018-11-17T18:13:35.213Z",
          "content": "<p>yes, I used beluga's shuffling code. When create my generate, I didn't generate batches randomly, my code generate those batches according to file sequence.  For example, we get batch_1 to batch_N from file_1,  we get batch_N+1 to batch_2N from file_2</p>",
          "rawMarkdown": "yes, I used beluga's shuffling code. When create my generate, I didn't generate batches randomly, my code generate those batches according to file sequence.  For example, we get batch_1 to batch_N from file_1,  we get batch_N+1 to batch_2N from file_2"
        },
        {
          "id": 423391,
          "postDate": "2018-11-18T05:41:36.730Z",
          "content": "<p>Ok, thanks!</p>",
          "rawMarkdown": "Ok, thanks!"
        },
        {
          "id": 425935,
          "postDate": "2018-11-22T10:35:19.403Z",
          "content": "<p>Do you mean that you load all files into memory?</p>",
          "rawMarkdown": "Do you mean that you load all files into memory?"
        },
        {
          "id": 426117,
          "postDate": "2018-11-22T16:44:36.320Z",
          "content": "<p>no, load data once a batch, with keras sequence, <a href=\"https://keras.io/utils/#sequence\">https://keras.io/utils/#sequence</a>,</p>",
          "rawMarkdown": "no, load data once a batch, with keras sequence, https://keras.io/utils/#sequence,"
        },
        {
          "id": 427465,
          "postDate": "2018-11-25T15:17:25.357Z",
          "content": "<p>Hey <a href=\"/yyqing\">@yyqing</a> .. I tried to create a keras.utils.Sequence class which I can use for multi gpus in keras. See below for my implementation              </p>\n\n<pre><code>class quickdrawSequence(Sequence):\n    def __init__(self, size, batch_size, total_len):\n        self.size = size\n        self.total_len = total_len\n        self.batch_size = batch_size\n\n    def __len__(self):\n        return self.total_len//self.batch_size\n\n    def __draw(raw_strokes, size=256, lw=6, time_color=True):\n        img = np.zeros((BASE_SIZE, BASE_SIZE), np.uint8)\n        for t, stroke in enumerate(raw_strokes):\n            for i in range(len(stroke[0]) - 1):\n                color = 255 - min(t, 10) * 13 if time_color else 255\n                _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                             (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n        if size != BASE_SIZE:\n            return cv2.resize(img, (self.size, self.size))\n        else:\n            return img\n\n    def __data_generation(self, skiprows):\n        train_filename = os.path.join(BASE_DIR, 'data', 'train.csv')\n        df = pd.read_csv(train_filename, \n                         header=None, \n                         skiprows=skiprows, \n                         nrows=self.batch_size)\n        df[1] = df[1].apply(ast.literal_eval)\n        x = np.zeros((len(df), self.size, self.size, 1))\n        for i, raw_strokes in enumerate(df[1].values):\n          x[i, :, :, 0] = self.__draw(raw_strokes)\n        print('After rawstrokes processing:', x.shape)\n        x = preprocess_input(x).astype(np.float32)\n        y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n        print(x.shape, y.shape)\n        return x,y\n\n    def __getitem__(self, idx):\n        skiprows = 1+(idx*self.batch_size)\n        print(idx, skiprows)\n        x,y = self.__data_generation(skiprows)\n        return x, y\n</code></pre>\n\n<p>This is the <code>fit_generator</code></p>\n\n<pre><code>hist = model.fit_generator(generator=quickdrawSequence(size, batchsize, 49673597),\n                           steps_per_epoch=STEPS,\n                           epochs=EPOCHS,\n                           verbose=1,\n                           validation_data=(x_valid, y_valid),\n                           callbacks=callbacks,\n                           max_queue_size=2,\n                           workers=2,\n                           use_multiprocessing=True,\n                           shuffle=True,\n                           )\n</code></pre>\n\n<p>While running this with 2 gpus..The training doesn't start..and generator keeps on sending idx=0 indefinitely. Below is what it prints.  </p>\n\n<pre><code>Epoch 1/1\n74133 37956097\n36175 18521601\n0 1\n0 1\n0 1\n</code></pre>\n\n<p>Any idea what lacks in my code. Also I have a single shuffled train file <code>train.csv</code> with all the images (except 100 per class that goes in validation) and trying to read each batch from this file using <code>pd.read_csv</code> with <code>skiprows</code> &amp; <code>nrows</code>.  </p>",
          "rawMarkdown": "Hey @yyqing .. I tried to create a keras.utils.Sequence class which I can use for multi gpus in keras. See below for my implementation              \n\n    class quickdrawSequence(Sequence):\n        def __init__(self, size, batch_size, total_len):\n            self.size = size\n            self.total_len = total_len\n            self.batch_size = batch_size\n    \n        def __len__(self):\n            return self.total_len//self.batch_size\n    \n        def __draw(raw_strokes, size=256, lw=6, time_color=True):\n            img = np.zeros((BASE_SIZE, BASE_SIZE), np.uint8)\n            for t, stroke in enumerate(raw_strokes):\n                for i in range(len(stroke[0]) - 1):\n                    color = 255 - min(t, 10) * 13 if time_color else 255\n                    _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                                 (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n            if size != BASE_SIZE:\n                return cv2.resize(img, (self.size, self.size))\n            else:\n                return img\n    \n        def __data_generation(self, skiprows):\n            train_filename = os.path.join(BASE_DIR, 'data', 'train.csv')\n            df = pd.read_csv(train_filename, \n                             header=None, \n                             skiprows=skiprows, \n                             nrows=self.batch_size)\n            df[1] = df[1].apply(ast.literal_eval)\n            x = np.zeros((len(df), self.size, self.size, 1))\n            for i, raw_strokes in enumerate(df[1].values):\n              x[i, :, :, 0] = self.__draw(raw_strokes)\n            print('After rawstrokes processing:', x.shape)\n            x = preprocess_input(x).astype(np.float32)\n            y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n            print(x.shape, y.shape)\n            return x,y\n    \n        def __getitem__(self, idx):\n            skiprows = 1+(idx*self.batch_size)\n            print(idx, skiprows)\n            x,y = self.__data_generation(skiprows)\n            return x, y\n\nThis is the `fit_generator`\n\n    hist = model.fit_generator(generator=quickdrawSequence(size, batchsize, 49673597),\n                               steps_per_epoch=STEPS,\n                               epochs=EPOCHS,\n                               verbose=1,\n                               validation_data=(x_valid, y_valid),\n                               callbacks=callbacks,\n                               max_queue_size=2,\n                               workers=2,\n                               use_multiprocessing=True,\n                               shuffle=True,\n                               )\n\nWhile running this with 2 gpus..The training doesn't start..and generator keeps on sending idx=0 indefinitely. Below is what it prints.  \n\n    Epoch 1/1\n    74133 37956097\n    36175 18521601\n    0 1\n    0 1\n    0 1\n\nAny idea what lacks in my code. Also I have a single shuffled train file `train.csv` with all the images (except 100 per class that goes in validation) and trying to read each batch from this file using `pd.read_csv` with `skiprows` &amp; `nrows`.  "
        },
        {
          "id": 427495,
          "postDate": "2018-11-25T16:01:13.300Z",
          "content": "<p>I think your code looks fine, mine looks similar like these.\nYou can try the example code of keras document, try to print the idx, see if it's always 0.</p>\n\n<pre><code>class CIFAR10Sequence(Sequence):\n    def __init__(self, x_set, y_set, batch_size):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n\n    def __len__(self):\n        return int(np.ceil(len(self.x) / float(self.batch_size)))\n\n    def __getitem__(self, idx):\n        batch_x = self.x[idx * self.batch_size:(idx + 1) * self.batch_size]\n        batch_y = self.y[idx * self.batch_size:(idx + 1) * self.batch_size]\n\n        return np.array([\n            resize(imread(file_name), (200, 200))\n               for file_name in batch_x]), np.array(batch_y)\n</code></pre>",
          "rawMarkdown": "I think your code looks fine, mine looks similar like these.\nYou can try the example code of keras document, try to print the idx, see if it's always 0.\n\n    class CIFAR10Sequence(Sequence):\n        def __init__(self, x_set, y_set, batch_size):\n            self.x, self.y = x_set, y_set\n            self.batch_size = batch_size\n    \n        def __len__(self):\n            return int(np.ceil(len(self.x) / float(self.batch_size)))\n    \n        def __getitem__(self, idx):\n            batch_x = self.x[idx * self.batch_size:(idx + 1) * self.batch_size]\n            batch_y = self.y[idx * self.batch_size:(idx + 1) * self.batch_size]\n    \n            return np.array([\n                resize(imread(file_name), (200, 200))\n                   for file_name in batch_x]), np.array(batch_y)"
        },
        {
          "id": 427604,
          "postDate": "2018-11-25T20:15:29.220Z",
          "content": "<p>I use a different method to get training on more than one GPU for a keras model.  I thought that the \n                            workers=2,\n                           use_multiprocessing=True,\nwas telling the fit generator to use multiple cores on my CPU rather than multiple GPU.</p>\n\n<p>I create a model\n                           model1 = MobileNet(input_shape=(size, size, 1), alpha=1, weights=None, classes=NCATS)</p>\n\n<p>than I setup for the multliple GPU</p>\n\n<p>from keras.utils import multi_gpu_model\n                            model = multi_gpu_model(model1, gpus=2)</p>\n\n<p>than compile and  fit generator\n                             model.compile(optimizer=Adam(lr=0.002), loss='categorical_crossentropy',\n              metrics=[categorical_crossentropy, categorical_accuracy, top_3_accuracy])</p>\n\n<pre><code>hist = model.fit_generator(\ntrain_datagen, steps_per_epoch=STEPS, epochs=EPOCHS, verbose=1,\nvalidation_data=(x_valid, y_valid), \ncallbacks = callbacks\n</code></pre>\n\n<p>)</p>",
          "rawMarkdown": "I use a different method to get training on more than one GPU for a keras model.  I thought that the \n                            workers=2,\n                           use_multiprocessing=True,\nwas telling the fit generator to use multiple cores on my CPU rather than multiple GPU.\n\nI create a model\n                           model1 = MobileNet(input_shape=(size, size, 1), alpha=1, weights=None, classes=NCATS)\n\nthan I setup for the multliple GPU\n\nfrom keras.utils import multi_gpu_model\n                            model = multi_gpu_model(model1, gpus=2)\n\nthan compile and  fit generator\n                             model.compile(optimizer=Adam(lr=0.002), loss='categorical_crossentropy',\n              metrics=[categorical_crossentropy, categorical_accuracy, top_3_accuracy])\n\n    hist = model.fit_generator(\n    train_datagen, steps_per_epoch=STEPS, epochs=EPOCHS, verbose=1,\n    validation_data=(x_valid, y_valid), \n    callbacks = callbacks\n)",
          "votes": 1
        },
        {
          "id": 427670,
          "postDate": "2018-11-25T23:54:55.183Z",
          "content": "<p>agree, if you want to use multiple GPU in keras, this is a must: <a href=\"https://keras.io/utils/#multi_gpu_model\">https://keras.io/utils/#multi_gpu_model</a></p>",
          "rawMarkdown": "agree, if you want to use multiple GPU in keras, this is a must: https://keras.io/utils/#multi_gpu_model"
        }
      ]
    },
    {
      "id": 422454,
      "postDate": "2018-11-16T08:42:03.467Z",
      "content": "<p>What setup do you use for multi GPU training?</p>\n\n<p>I tried the default recommended tf/keras model-&gt; multi_gpu_model way. It was not really faster.\nI rent GCP GPUs so I have to find a good cost&amp;performance solution</p>",
      "rawMarkdown": "What setup do you use for multi GPU training?\n\nI tried the default recommended tf/keras model-&gt; multi_gpu_model way. It was not really faster.\nI rent GCP GPUs so I have to find a good cost&amp;performance solution\n",
      "votes": 4
    },
    {
      "id": 422926,
      "postDate": "2018-11-17T04:40:43.467Z",
      "content": "<p>Running your MobileNet kernel (with some modifications) on two PC's.  One has two GPU and the other has three GPU.  But I have only a little experience regarding things running faster.  I have always increased batch sizes or steps to try for training with more data in the same time frame rather than trying to improve speed.</p>\n\n<p>Your kernel did not run successfully with multi GPU until I changed the \"from\" statements, so I did get a little peek at speed difference at same settings on the first successful run with two GPU.  It's a little tricky getting the GPU's to run at steady rates - watching GPU usage on MSI Afterburner I saw that a huge sawtooth plot present.  The GPU's not getting fed fast enough - when that has happened in past, speed with two GPU's was often slower than speed with a single.\nI have had issues with using the workers=x - I am too much of a noobie to get the generator thread safe.  So I play with batch size bigger which kind of matches my desire to run more data in the same time.</p>\n\n<p>So to answer your question - use enough workers to keep CPU at 50% or higher, and large enough batches to avoid big sawtooth usage plots on GPU usage.  Like most things in deep learning - you got to tweak.  </p>",
      "rawMarkdown": "Running your MobileNet kernel (with some modifications) on two PC's.  One has two GPU and the other has three GPU.  But I have only a little experience regarding things running faster.  I have always increased batch sizes or steps to try for training with more data in the same time frame rather than trying to improve speed.\n\nYour kernel did not run successfully with multi GPU until I changed the \"from\" statements, so I did get a little peek at speed difference at same settings on the first successful run with two GPU.  It's a little tricky getting the GPU's to run at steady rates - watching GPU usage on MSI Afterburner I saw that a huge sawtooth plot present.  The GPU's not getting fed fast enough - when that has happened in past, speed with two GPU's was often slower than speed with a single.\nI have had issues with using the workers=x - I am too much of a noobie to get the generator thread safe.  So I play with batch size bigger which kind of matches my desire to run more data in the same time.\n\nSo to answer your question - use enough workers to keep CPU at 50% or higher, and large enough batches to avoid big sawtooth usage plots on GPU usage.  Like most things in deep learning - you got to tweak.  \n",
      "votes": 1
    },
    {
      "id": 985216,
      "postDate": "2020-08-25T15:11:49.060Z",
      "content": "<p>If you are planning to do get best bang for buck in terms of GPU training then you can either setup a local on-premise system with multiple GPUs or go on a peer to peer GPU cloud to get the best costing for GPU instances.</p>\n<p>Building an on-prem system can be very costly upfront. Going on cloud is the best bet in that case. But primitive clouds are pretty costly for GPU instances. In that case, you can explore P2p GPU platform like <a href=\"https://www.qblocks.cloud\" target=\"_blank\">Q Blocks</a> to get gpu instances at upto 10x lower costs for your AI workloads.</p>",
      "rawMarkdown": "If you are planning to do get best bang for buck in terms of GPU training then you can either setup a local on-premise system with multiple GPUs or go on a peer to peer GPU cloud to get the best costing for GPU instances.\n\nBuilding an on-prem system can be very costly upfront. Going on cloud is the best bet in that case. But primitive clouds are pretty costly for GPU instances. In that case, you can explore P2p GPU platform like [Q Blocks](https://www.qblocks.cloud) to get gpu instances at upto 10x lower costs for your AI workloads."
    },
    {
      "id": 428380,
      "postDate": "2018-11-27T07:22:12.770Z",
      "content": "<p>It works for me. I use multi-gpu with MobileNet / Xception in Keras. Both of them shows faster running for the same batch size. If I double my batchsize, the speed of each fit-generator epoch will not increase (but that's mean you can process double the data, so effectively it is faster) </p>",
      "rawMarkdown": "It works for me. I use multi-gpu with MobileNet / Xception in Keras. Both of them shows faster running for the same batch size. If I double my batchsize, the speed of each fit-generator epoch will not increase (but that's mean you can process double the data, so effectively it is faster) "
    }
  ],
  "comments": [
    {
      "id": 422467,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-11-16T09:00:53.350000",
      "content": "<p>I'm using keras, multiGPU on GCP. It works fine for me.    Two GPUs can double my batch_size.</p>\n\n<p>A few days ago, I solved my performance issue. The GPU usage is usually 0%, and one CPU process is very busy. After using keras data generator with multi_processing, more CPUs are reading csv and rendering pictures,  then the GPU is much busier than before.</p>\n\n<p>By the way, I set workers=8 in the parallel_model.fit_generator, this can keep 10 CPU and 2 P100 GPU busy most of the time.</p>\n\n<pre><code>parallel_model.fit_generator(\n    tr_generator, \n    steps_per_epoch=STEPS, \n    epochs=1, \n    verbose=1,\n    validation_data=va_generator,\n    validation_steps=STEPS_VALID,\n    callbacks = callbacks,\n    use_multiprocessing=True,\n    workers=8\n)\n</code></pre>",
      "votes": 5,
      "replies": [
        {
          "id": 422620,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-16T14:13:24.213000",
          "content": "<p>Is your <code>tr_generator</code>  an instance of <code>keras.utils.Sequence</code> ?  There will exist a warning about \"data copy\" in my model and be slower when I keep <code>use_multiprocessing=True</code>,  did you have the same warning? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422713,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-16T16:57:56.043000",
          "content": "<p>yes, it's a keras.utils.Sequence.  no warning about \"data copy\".</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423061,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-17T11:25:51.153000",
          "content": "<p>Did you use the shuffled csv to generate training data ? I don't know how to implement the <code>Sequence</code> method <code>__getitem__</code> which can generate the <code>idx</code>th batch data randomly like this: <br>\n<code>\ndef __getitem__(self, idx): <br>\n        batch_x = .... <br>\n        batch_y = .... <br>\n        return batch_x, batch_y\n</code> <br>\nSo, I implement a generator which can only generate sequential batched of data: <br>\n<code>\nfor df in pd.read_csv(filename, chunksize=batchsize): <br>\n                batch_x = df_to_gray_image_array(...) <br>\n                batch_y = keras.utils.to_categorical(df.y, num_classes=NCATS) <br>\n                yield x, y\n</code> <br>\nCould you tell me how to implement <code>__getitem__</code> which can generate batches randomly? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423222,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-17T18:13:35.213000",
          "content": "<p>yes, I used beluga's shuffling code. When create my generate, I didn't generate batches randomly, my code generate those batches according to file sequence.  For example, we get batch_1 to batch_N from file_1,  we get batch_N+1 to batch_2N from file_2</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423391,
          "author_name": "upup",
          "author_url": "",
          "post_date": "2018-11-18T05:41:36.730000",
          "content": "<p>Ok, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 425935,
          "author_name": "Wolverine",
          "author_url": "",
          "post_date": "2018-11-22T10:35:19.403000",
          "content": "<p>Do you mean that you load all files into memory?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 426117,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-22T16:44:36.320000",
          "content": "<p>no, load data once a batch, with keras sequence, <a href=\"https://keras.io/utils/#sequence\">https://keras.io/utils/#sequence</a>,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427465,
          "author_name": "Abhilash Awasthi",
          "author_url": "",
          "post_date": "2018-11-25T15:17:25.357000",
          "content": "<p>Hey <a href=\"/yyqing\">@yyqing</a> .. I tried to create a keras.utils.Sequence class which I can use for multi gpus in keras. See below for my implementation              </p>\n\n<pre><code>class quickdrawSequence(Sequence):\n    def __init__(self, size, batch_size, total_len):\n        self.size = size\n        self.total_len = total_len\n        self.batch_size = batch_size\n\n    def __len__(self):\n        return self.total_len//self.batch_size\n\n    def __draw(raw_strokes, size=256, lw=6, time_color=True):\n        img = np.zeros((BASE_SIZE, BASE_SIZE), np.uint8)\n        for t, stroke in enumerate(raw_strokes):\n            for i in range(len(stroke[0]) - 1):\n                color = 255 - min(t, 10) * 13 if time_color else 255\n                _ = cv2.line(img, (stroke[0][i], stroke[1][i]),\n                             (stroke[0][i + 1], stroke[1][i + 1]), color, lw)\n        if size != BASE_SIZE:\n            return cv2.resize(img, (self.size, self.size))\n        else:\n            return img\n\n    def __data_generation(self, skiprows):\n        train_filename = os.path.join(BASE_DIR, 'data', 'train.csv')\n        df = pd.read_csv(train_filename, \n                         header=None, \n                         skiprows=skiprows, \n                         nrows=self.batch_size)\n        df[1] = df[1].apply(ast.literal_eval)\n        x = np.zeros((len(df), self.size, self.size, 1))\n        for i, raw_strokes in enumerate(df[1].values):\n          x[i, :, :, 0] = self.__draw(raw_strokes)\n        print('After rawstrokes processing:', x.shape)\n        x = preprocess_input(x).astype(np.float32)\n        y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n        print(x.shape, y.shape)\n        return x,y\n\n    def __getitem__(self, idx):\n        skiprows = 1+(idx*self.batch_size)\n        print(idx, skiprows)\n        x,y = self.__data_generation(skiprows)\n        return x, y\n</code></pre>\n\n<p>This is the <code>fit_generator</code></p>\n\n<pre><code>hist = model.fit_generator(generator=quickdrawSequence(size, batchsize, 49673597),\n                           steps_per_epoch=STEPS,\n                           epochs=EPOCHS,\n                           verbose=1,\n                           validation_data=(x_valid, y_valid),\n                           callbacks=callbacks,\n                           max_queue_size=2,\n                           workers=2,\n                           use_multiprocessing=True,\n                           shuffle=True,\n                           )\n</code></pre>\n\n<p>While running this with 2 gpus..The training doesn't start..and generator keeps on sending idx=0 indefinitely. Below is what it prints.  </p>\n\n<pre><code>Epoch 1/1\n74133 37956097\n36175 18521601\n0 1\n0 1\n0 1\n</code></pre>\n\n<p>Any idea what lacks in my code. Also I have a single shuffled train file <code>train.csv</code> with all the images (except 100 per class that goes in validation) and trying to read each batch from this file using <code>pd.read_csv</code> with <code>skiprows</code> &amp; <code>nrows</code>.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427495,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-25T16:01:13.300000",
          "content": "<p>I think your code looks fine, mine looks similar like these.\nYou can try the example code of keras document, try to print the idx, see if it's always 0.</p>\n\n<pre><code>class CIFAR10Sequence(Sequence):\n    def __init__(self, x_set, y_set, batch_size):\n        self.x, self.y = x_set, y_set\n        self.batch_size = batch_size\n\n    def __len__(self):\n        return int(np.ceil(len(self.x) / float(self.batch_size)))\n\n    def __getitem__(self, idx):\n        batch_x = self.x[idx * self.batch_size:(idx + 1) * self.batch_size]\n        batch_y = self.y[idx * self.batch_size:(idx + 1) * self.batch_size]\n\n        return np.array([\n            resize(imread(file_name), (200, 200))\n               for file_name in batch_x]), np.array(batch_y)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 427604,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2018-11-25T20:15:29.220000",
          "content": "<p>I use a different method to get training on more than one GPU for a keras model.  I thought that the \n                            workers=2,\n                           use_multiprocessing=True,\nwas telling the fit generator to use multiple cores on my CPU rather than multiple GPU.</p>\n\n<p>I create a model\n                           model1 = MobileNet(input_shape=(size, size, 1), alpha=1, weights=None, classes=NCATS)</p>\n\n<p>than I setup for the multliple GPU</p>\n\n<p>from keras.utils import multi_gpu_model\n                            model = multi_gpu_model(model1, gpus=2)</p>\n\n<p>than compile and  fit generator\n                             model.compile(optimizer=Adam(lr=0.002), loss='categorical_crossentropy',\n              metrics=[categorical_crossentropy, categorical_accuracy, top_3_accuracy])</p>\n\n<pre><code>hist = model.fit_generator(\ntrain_datagen, steps_per_epoch=STEPS, epochs=EPOCHS, verbose=1,\nvalidation_data=(x_valid, y_valid), \ncallbacks = callbacks\n</code></pre>\n\n<p>)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 427670,
          "author_name": "yyqing",
          "author_url": "",
          "post_date": "2018-11-25T23:54:55.183000",
          "content": "<p>agree, if you want to use multiple GPU in keras, this is a must: <a href=\"https://keras.io/utils/#multi_gpu_model\">https://keras.io/utils/#multi_gpu_model</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422926,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-11-17T04:40:43.467000",
      "content": "<p>Running your MobileNet kernel (with some modifications) on two PC's.  One has two GPU and the other has three GPU.  But I have only a little experience regarding things running faster.  I have always increased batch sizes or steps to try for training with more data in the same time frame rather than trying to improve speed.</p>\n\n<p>Your kernel did not run successfully with multi GPU until I changed the \"from\" statements, so I did get a little peek at speed difference at same settings on the first successful run with two GPU.  It's a little tricky getting the GPU's to run at steady rates - watching GPU usage on MSI Afterburner I saw that a huge sawtooth plot present.  The GPU's not getting fed fast enough - when that has happened in past, speed with two GPU's was often slower than speed with a single.\nI have had issues with using the workers=x - I am too much of a noobie to get the generator thread safe.  So I play with batch size bigger which kind of matches my desire to run more data in the same time.</p>\n\n<p>So to answer your question - use enough workers to keep CPU at 50% or higher, and large enough batches to avoid big sawtooth usage plots on GPU usage.  Like most things in deep learning - you got to tweak.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 985216,
      "author_name": "Gaurav",
      "author_url": "",
      "post_date": "2020-08-25T15:11:49.060000",
      "content": "<p>If you are planning to do get best bang for buck in terms of GPU training then you can either setup a local on-premise system with multiple GPUs or go on a peer to peer GPU cloud to get the best costing for GPU instances.</p>\n<p>Building an on-prem system can be very costly upfront. Going on cloud is the best bet in that case. But primitive clouds are pretty costly for GPU instances. In that case, you can explore P2p GPU platform like <a href=\"https://www.qblocks.cloud\" target=\"_blank\">Q Blocks</a> to get gpu instances at upto 10x lower costs for your AI workloads.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 428380,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-11-27T07:22:12.770000",
      "content": "<p>It works for me. I use multi-gpu with MobileNet / Xception in Keras. Both of them shows faster running for the same batch size. If I double my batchsize, the speed of each fit-generator epoch will not increase (but that's mean you can process double the data, so effectively it is faster) </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422467": "I'm using keras, multiGPU on GCP. It works fine for me.    Two GPUs can double my batch_size.\n\nA few days ago, I solved my performance issue. The GPU usage is usually 0%, and one CPU process is very busy. After using keras data generator with multi_processing, more CPUs are reading csv and rendering pictures,  then the GPU is much busier than before.\n\nBy the way, I set workers=8 in the parallel_model.fit_generator, this can keep 10 CPU and 2 P100 GPU busy most of the time.\n\n    parallel_model.fit_generator(\n        tr_generator, \n        steps_per_epoch=STEPS, \n        epochs=1, \n        verbose=1,\n        validation_data=va_generator,\n        validation_steps=STEPS_VALID,\n        callbacks = callbacks,\n        use_multiprocessing=True,\n        workers=8\n    )",
    "422454": "What setup do you use for multi GPU training?\n\nI tried the default recommended tf/keras model-&gt; multi_gpu_model way. It was not really faster.\nI rent GCP GPUs so I have to find a good cost&amp;performance solution\n",
    "422926": "Running your MobileNet kernel (with some modifications) on two PC's.  One has two GPU and the other has three GPU.  But I have only a little experience regarding things running faster.  I have always increased batch sizes or steps to try for training with more data in the same time frame rather than trying to improve speed.\n\nYour kernel did not run successfully with multi GPU until I changed the \"from\" statements, so I did get a little peek at speed difference at same settings on the first successful run with two GPU.  It's a little tricky getting the GPU's to run at steady rates - watching GPU usage on MSI Afterburner I saw that a huge sawtooth plot present.  The GPU's not getting fed fast enough - when that has happened in past, speed with two GPU's was often slower than speed with a single.\nI have had issues with using the workers=x - I am too much of a noobie to get the generator thread safe.  So I play with batch size bigger which kind of matches my desire to run more data in the same time.\n\nSo to answer your question - use enough workers to keep CPU at 50% or higher, and large enough batches to avoid big sawtooth usage plots on GPU usage.  Like most things in deep learning - you got to tweak.  \n",
    "985216": "If you are planning to do get best bang for buck in terms of GPU training then you can either setup a local on-premise system with multiple GPUs or go on a peer to peer GPU cloud to get the best costing for GPU instances.\n\nBuilding an on-prem system can be very costly upfront. Going on cloud is the best bet in that case. But primitive clouds are pretty costly for GPU instances. In that case, you can explore P2p GPU platform like [Q Blocks](https://www.qblocks.cloud) to get gpu instances at upto 10x lower costs for your AI workloads.",
    "428380": "It works for me. I use multi-gpu with MobileNet / Xception in Keras. Both of them shows faster running for the same batch size. If I double my batchsize, the speed of each fit-generator epoch will not increase (but that's mean you can process double the data, so effectively it is faster) "
  }
}