{
  "id": 72110,
  "title": "Training MobileNet on multi GPUs is slow?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/72110",
  "author_name": "good good study",
  "post_date": "2018-11-20T12:48:39.565000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was trying to train MobileNet on multi gpus using pytorch. From <code>watch nvidia-smi</code>, I see that GPUs are sometimes working, and sometimes not working(GPU util is 0%). This slows down training speed a lot.</p>\n\n<p>But training MobileNet  on a single GPU and training ResNet50 on multi GPUs do not have such issue. I was wondering what is going wrong. Is there someone used to meet this problem?</p>\n\n<p>PS: I read all training data into memory.  Pytorch version is 0.4.0</p>",
  "messages": [
    {
      "id": 424622,
      "postDate": "2018-11-20T12:48:39.567Z",
      "content": "<p>I was trying to train MobileNet on multi gpus using pytorch. From <code>watch nvidia-smi</code>, I see that GPUs are sometimes working, and sometimes not working(GPU util is 0%). This slows down training speed a lot.</p>\n\n<p>But training MobileNet  on a single GPU and training ResNet50 on multi GPUs do not have such issue. I was wondering what is going wrong. Is there someone used to meet this problem?</p>\n\n<p>PS: I read all training data into memory.  Pytorch version is 0.4.0</p>",
      "rawMarkdown": "I was trying to train MobileNet on multi gpus using pytorch. From `watch nvidia-smi`, I see that GPUs are sometimes working, and sometimes not working(GPU util is 0%). This slows down training speed a lot.\n\nBut training MobileNet  on a single GPU and training ResNet50 on multi GPUs do not have such issue. I was wondering what is going wrong. Is there someone used to meet this problem?\n\nPS: I read all training data into memory.  Pytorch version is 0.4.0",
      "votes": 1
    },
    {
      "id": 426511,
      "postDate": "2018-11-23T11:32:33.557Z",
      "content": "<p>Was facing the same problem. I was using 8 V100s and time for one epoch was nearly same as 4 V100s. I was using data loaders and had to increasse my num_workers(V100s are very fast), after that still usage was dropping to zero.\nI decreased from 8 to 4. Now utilization is alteast not zero at any time. \nSince you are loading all of data into memory, so my solution won't help you. Are you monitoring cpu usage?? The weights are updated on the cpu, try debuggin there.</p>",
      "rawMarkdown": "Was facing the same problem. I was using 8 V100s and time for one epoch was nearly same as 4 V100s. I was using data loaders and had to increasse my num_workers(V100s are very fast), after that still usage was dropping to zero.\nI decreased from 8 to 4. Now utilization is alteast not zero at any time. \nSince you are loading all of data into memory, so my solution won't help you. Are you monitoring cpu usage?? The weights are updated on the cpu, try debuggin there."
    },
    {
      "id": 424923,
      "postDate": "2018-11-20T22:15:39.423Z",
      "content": "<p>Do most of my stuff on two or three GPU local PC.  Have occasionally seen what your describing.  When its extreme two GPU's end up slower than one.  I am pretty much a noobie with Python, so I seldom can fix this issue.  To get maximum GPU usage I have success with things that increase the CPU usage.\nThings that seem to work.\n1.  Larger batch size -  make it just short of the size that creates OOM.   This seems to do the best job when CPU usage is low.\n2. Make validation size smaller.\n3.  Use generator. <br>\n4.  Play with max_q_size callback.</p>",
      "rawMarkdown": "Do most of my stuff on two or three GPU local PC.  Have occasionally seen what your describing.  When its extreme two GPU's end up slower than one.  I am pretty much a noobie with Python, so I seldom can fix this issue.  To get maximum GPU usage I have success with things that increase the CPU usage.\nThings that seem to work.\n1.  Larger batch size -  make it just short of the size that creates OOM.   This seems to do the best job when CPU usage is low.\n2. Make validation size smaller.\n3.  Use generator.  \n4.  Play with max_q_size callback.",
      "replies": [
        {
          "id": 424983,
          "postDate": "2018-11-21T01:04:55.150Z",
          "content": "<p>Thanks for you reply. I think you are using Keras. I have also tried Keras, using multi gpus is slower, but in Pytorch, using multi gpus is much more slower than Keras.</p>",
          "rawMarkdown": "Thanks for you reply. I think you are using Keras. I have also tried Keras, using multi gpus is slower, but in Pytorch, using multi gpus is much more slower than Keras."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 426511,
      "author_name": "VikasSangwan",
      "author_url": "",
      "post_date": "2018-11-23T11:32:33.557000",
      "content": "<p>Was facing the same problem. I was using 8 V100s and time for one epoch was nearly same as 4 V100s. I was using data loaders and had to increasse my num_workers(V100s are very fast), after that still usage was dropping to zero.\nI decreased from 8 to 4. Now utilization is alteast not zero at any time. \nSince you are loading all of data into memory, so my solution won't help you. Are you monitoring cpu usage?? The weights are updated on the cpu, try debuggin there.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 424923,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2018-11-20T22:15:39.423000",
      "content": "<p>Do most of my stuff on two or three GPU local PC.  Have occasionally seen what your describing.  When its extreme two GPU's end up slower than one.  I am pretty much a noobie with Python, so I seldom can fix this issue.  To get maximum GPU usage I have success with things that increase the CPU usage.\nThings that seem to work.\n1.  Larger batch size -  make it just short of the size that creates OOM.   This seems to do the best job when CPU usage is low.\n2. Make validation size smaller.\n3.  Use generator. <br>\n4.  Play with max_q_size callback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 424983,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2018-11-21T01:04:55.150000",
          "content": "<p>Thanks for you reply. I think you are using Keras. I have also tried Keras, using multi gpus is slower, but in Pytorch, using multi gpus is much more slower than Keras.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "424622": "I was trying to train MobileNet on multi gpus using pytorch. From `watch nvidia-smi`, I see that GPUs are sometimes working, and sometimes not working(GPU util is 0%). This slows down training speed a lot.\n\nBut training MobileNet  on a single GPU and training ResNet50 on multi GPUs do not have such issue. I was wondering what is going wrong. Is there someone used to meet this problem?\n\nPS: I read all training data into memory.  Pytorch version is 0.4.0",
    "426511": "Was facing the same problem. I was using 8 V100s and time for one epoch was nearly same as 4 V100s. I was using data loaders and had to increasse my num_workers(V100s are very fast), after that still usage was dropping to zero.\nI decreased from 8 to 4. Now utilization is alteast not zero at any time. \nSince you are loading all of data into memory, so my solution won't help you. Are you monitoring cpu usage?? The weights are updated on the cpu, try debuggin there.",
    "424923": "Do most of my stuff on two or three GPU local PC.  Have occasionally seen what your describing.  When its extreme two GPU's end up slower than one.  I am pretty much a noobie with Python, so I seldom can fix this issue.  To get maximum GPU usage I have success with things that increase the CPU usage.\nThings that seem to work.\n1.  Larger batch size -  make it just short of the size that creates OOM.   This seems to do the best job when CPU usage is low.\n2. Make validation size smaller.\n3.  Use generator.  \n4.  Play with max_q_size callback."
  }
}