{
  "id": 192674,
  "title": "Is the compute too big for this competition?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/192674",
  "author_name": "Mark P",
  "post_date": "2020-10-22T16:06:29.848000",
  "votes": 2,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>I struggled for a long time to train an Effnet on all the data (Ian Pan's Jpeg dataset) in pytorch.  I used the GPUs on kaggle and got it down to 45mins/epoch but thought this was still far too long.  I then tried using the TPUs and got it running but the times were the same so I think my bottleneck was the CPU.</p>\n<p>But reading in detail through <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> training notebooks I see his training took many many GPU hours too and I think he went to GCP to train some parts.</p>\n<p>My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models?  If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?</p>\n<p>or is everyone using TF and/or TPUs and its just me struggling to get pytorch working that is the problem?   [feeling it is probably&lt;---]</p>",
  "messages": [
    {
      "id": 1057411,
      "postDate": "2020-10-22T16:06:29.850Z",
      "content": "<p>Hi all,</p>\n<p>I struggled for a long time to train an Effnet on all the data (Ian Pan's Jpeg dataset) in pytorch.  I used the GPUs on kaggle and got it down to 45mins/epoch but thought this was still far too long.  I then tried using the TPUs and got it running but the times were the same so I think my bottleneck was the CPU.</p>\n<p>But reading in detail through <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> training notebooks I see his training took many many GPU hours too and I think he went to GCP to train some parts.</p>\n<p>My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models?  If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?</p>\n<p>or is everyone using TF and/or TPUs and its just me struggling to get pytorch working that is the problem?   [feeling it is probably&lt;---]</p>",
      "rawMarkdown": "Hi all,\n\nI struggled for a long time to train an Effnet on all the data (Ian Pan's Jpeg dataset) in pytorch.  I used the GPUs on kaggle and got it down to 45mins/epoch but thought this was still far too long.  I then tried using the TPUs and got it running but the times were the same so I think my bottleneck was the CPU.\n\nBut reading in detail through @khyeh0719 training notebooks I see his training took many many GPU hours too and I think he went to GCP to train some parts.\n\nMy question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models?  If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?\n\nor is everyone using TF and/or TPUs and its just me struggling to get pytorch working that is the problem?   [feeling it is probably<---]",
      "votes": 2
    },
    {
      "id": 1057625,
      "postDate": "2020-10-22T19:58:12.273Z",
      "content": "<blockquote>\n  <p>My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models? If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?</p>\n</blockquote>\n<p>You can still use Colab Pro,  it's way much cheaper than GCP;  That's what I do whenever I need extensive resources</p>\n<p>And for GPU, colab Pro has now  <code>Tesla V100-sxm2</code> option, which is way faster than kaggle P100</p>",
      "rawMarkdown": "> My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models? If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?\n\n\n\nYou can still use Colab Pro,  it's way much cheaper than GCP;  That's what I do whenever I need extensive resources\n\nAnd for GPU, colab Pro has now  `Tesla V100-sxm2` option, which is way faster than kaggle P100",
      "votes": 1,
      "replies": [
        {
          "id": 1057650,
          "postDate": "2020-10-22T20:25:27.230Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1057654,
          "postDate": "2020-10-22T20:30:38.010Z",
          "content": "<p>Yeah High RAM.</p>\n<p>Once I get P100, I keep retinitializing the notebook, until I get V100 ^^</p>",
          "rawMarkdown": "Yeah High RAM.\n\nOnce I get P100, I keep retinitializing the notebook, until I get V100 ^^",
          "votes": 1
        },
        {
          "id": 1057705,
          "postDate": "2020-10-22T22:38:15.227Z",
          "content": "<p>Colab pro is only available in US, so that option doesn't work for many of us. Is a great advantage for those who can use that opportunity. Just to wait and hope for a release in Europe.</p>",
          "rawMarkdown": "Colab pro is only available in US, so that option doesn't work for many of us. Is a great advantage for those who can use that opportunity. Just to wait and hope for a release in Europe."
        },
        {
          "id": 1057951,
          "postDate": "2020-10-23T07:20:14.053Z",
          "content": "<p>Yeah this is kind of my point - if you can only do well in this competition by using external resources (own GPUs, Colab Pro etc) that either cost money or are only available to USA participants is it really a fair competition?  Shouldn't it be about knowledge rather than availability of resources?</p>",
          "rawMarkdown": "Yeah this is kind of my point - if you can only do well in this competition by using external resources (own GPUs, Colab Pro etc) that either cost money or are only available to USA participants is it really a fair competition?  Shouldn't it be about knowledge rather than availability of resources?",
          "votes": 1
        },
        {
          "id": 1057978,
          "postDate": "2020-10-23T08:09:35.413Z",
          "content": "<p>It is quite common that $ prize/ gold medal winners trained their models entirely by Kaggle Kernel/ Colab. Hardware can make a difference but more important is your knowledge, experience and \"luck\".</p>\n<p>I think as long as there is no un-disclosed external data source and private sharing, the competition is considered as fair. </p>\n<p>Will become a philosophy question if drill too deep about \"fairness\". If both Tom and Jerry are joining this competition but Tom is a machine learning PhD while Jerry is just a chemistry graduate, is this fair even they have the same hardware?</p>",
          "rawMarkdown": "It is quite common that $ prize/ gold medal winners trained their models entirely by Kaggle Kernel/ Colab. Hardware can make a difference but more important is your knowledge, experience and \"luck\".\n\nI think as long as there is no un-disclosed external data source and private sharing, the competition is considered as fair. \n\nWill become a philosophy question if drill too deep about \"fairness\". If both Tom and Jerry are joining this competition but Tom is a machine learning PhD while Jerry is just a chemistry graduate, is this fair even they have the same hardware?\n",
          "votes": 1
        },
        {
          "id": 1058031,
          "postDate": "2020-10-23T09:09:57.507Z",
          "content": "<p>One can make it simple, in general, and to not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering.</p>",
          "rawMarkdown": "One can make it simple, in general, and to not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering."
        },
        {
          "id": 1058067,
          "postDate": "2020-10-23T09:53:35.093Z",
          "content": "<blockquote>\n  <p>Shouldn't it be about knowledge rather than availability of resources?</p>\n</blockquote>\n<p>And why not both ? </p>\n<p>As far as i'm concerned, I don't have even own single GPU and Colab Pro is available everywhere not only the USA (I suscribed from Paris in the early days). </p>\n<p>As for fairness, it's quite relative . Because you can also say even Kaggle itself is not fair,  as it gives +30 hours of GPU/TPU  to each team member for people teaming up. </p>\n<p>Anyway this competition requires extensive resources by its nature and by the size of the dataset.  You can't help but finding solutions like teaming up, using online services or buying your own solid infrastructure.  Otherwise there are still other competitions that require much less resources</p>",
          "rawMarkdown": "> Shouldn't it be about knowledge rather than availability of resources?\n\nAnd why not both ? \n\nAs far as i'm concerned, I don't have even own single GPU and Colab Pro is available everywhere not only the USA (I suscribed from Paris in the early days). \n\nAs for fairness, it's quite relative . Because you can also say even Kaggle itself is not fair,  as it gives +30 hours of GPU/TPU  to each team member for people teaming up. \n\nAnyway this competition requires extensive resources by its nature and by the size of the dataset.  You can't help but finding solutions like teaming up, using online services or buying your own solid infrastructure.  Otherwise there are still other competitions that require much less resources",
          "votes": -1
        },
        {
          "id": 1058071,
          "postDate": "2020-10-23T10:00:31.127Z",
          "content": "<blockquote>\n  <p>One can make it simple, in general, and do not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering.</p>\n</blockquote>\n<p>I propose we all use 2 cores CPU only and build every thing from scratch. Because, even using Pytorch instead of writing own CNN and Backpropagation algorithms, give some advantages. </p>",
          "rawMarkdown": ">  One can make it simple, in general, and do not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering.\n\nI propose we all use 2 cores CPU only and build every thing from scratch. Because, even using Pytorch instead of writing own CNN and Backpropagation algorithms, give some advantages. \n"
        },
        {
          "id": 1058111,
          "postDate": "2020-10-23T10:59:28.167Z",
          "content": "<p>I thought we were talking about hw. I can only talk for my self, if I had the possible to use more TPU and GPU time through Colab Pro (V100 is really a good option that they have added, using Volta+mixed precision 👍) I would use it to help tuning and experiment more and wider, but unfortunately it says only available in US and Canada. Would it help me make a better model or getting a better scores? Who knows, that's more a factor of knowledge but I get that knowledge through testing and experimenting, thats why one loves Kaggle and Colab, resources and challenges helping one getting better in different AI areas through helping and contribute solving a problem and with the goal to help and contribute solving problems even better :)<br>\nI hope they make Colab Pro more available, that's my point.</p>",
          "rawMarkdown": "I thought we were talking about hw. I can only talk for my self, if I had the possible to use more TPU and GPU time through Colab Pro (V100 is really a good option that they have added, using Volta+mixed precision 👍) I would use it to help tuning and experiment more and wider, but unfortunately it says only available in US and Canada. Would it help me make a better model or getting a better scores? Who knows, that's more a factor of knowledge but I get that knowledge through testing and experimenting, thats why one loves Kaggle and Colab, resources and challenges helping one getting better in different AI areas through helping and contribute solving a problem and with the goal to help and contribute solving problems even better :)\nI hope they make Colab Pro more available, that's my point.",
          "votes": 1
        },
        {
          "id": 1058953,
          "postDate": "2020-10-24T12:53:53.883Z",
          "content": "<p>As I said, Colab Pro is available outside of these countries. Don't mind what is written at the homepage. </p>\n<p>I'm just giving advices of the most affordable  resources to common kagglers in such highly demanding competition.  You can be sure many people at (top) LB use much more extensive resources. </p>\n<p>It's (nearly) impossible to win such competition by using Kaggle only, just like the annual Youtube-8m video competition (+ 1.5 TB of data last year) and many others google competitions. <br>\nThat's why kaggle (and sponsors)  offer GCP credits on most of these competitions, in particular for those who can't afford high computational power</p>",
          "rawMarkdown": "As I said, Colab Pro is available outside of these countries. Don't mind what is written at the homepage. \n\nI'm just giving advices of the most affordable  resources to common kagglers in such highly demanding competition.  You can be sure many people at (top) LB use much more extensive resources. \n\nIt's (nearly) impossible to win such competition by using Kaggle only, just like the annual Youtube-8m video competition (+ 1.5 TB of data last year) and many others google competitions. \nThat's why kaggle (and sponsors)  offer GCP credits on most of these competitions, in particular for those who can't afford high computational power\n\n",
          "votes": 1
        },
        {
          "id": 1059061,
          "postDate": "2020-10-24T15:09:50.630Z",
          "content": "<p>Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. </p>",
          "rawMarkdown": "Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. \n\n",
          "votes": 1
        },
        {
          "id": 1059084,
          "postDate": "2020-10-24T15:40:51.823Z",
          "content": "<p>Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. </p>",
          "rawMarkdown": "Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. \n\n"
        },
        {
          "id": 1059176,
          "postDate": "2020-10-24T18:04:53.123Z",
          "content": "<p>Yes and it's a great advice Serigne 👍 So you think it's OK to use outside US and Canada? I have to look into it again then, maybe something has changed :)<br>\nAnd I'm also greatfull that we can use Kaggle and Colab, opens the door for everyone to learn and contribute.  So I don't want to sound ungrateful but it's just that the Colab Pro seems like a great subscribe service and open up for even more powerful testing and training for everyone. They have now added another country to the list, so maybe more will come soon if it is not possible to use.</p>",
          "rawMarkdown": "Yes and it's a great advice Serigne 👍 So you think it's OK to use outside US and Canada? I have to look into it again then, maybe something has changed :)\nAnd I'm also greatfull that we can use Kaggle and Colab, opens the door for everyone to learn and contribute.  So I don't want to sound ungrateful but it's just that the Colab Pro seems like a great subscribe service and open up for even more powerful testing and training for everyone. They have now added another country to the list, so maybe more will come soon if it is not possible to use."
        }
      ]
    },
    {
      "id": 1058657,
      "postDate": "2020-10-24T04:05:39.717Z",
      "content": "<p>You can try rapid from nvidia for using gpu power while matrix calculations and preporcessing.</p>",
      "rawMarkdown": "You can try rapid from nvidia for using gpu power while matrix calculations and preporcessing."
    },
    {
      "id": 1057436,
      "postDate": "2020-10-22T16:32:39.863Z",
      "content": "<p>We <em>were</em> silver before the public notebook (30th or so). We used Efficientnetb0 trained on tfrecords and TPU, and had an epoch time of about 5 minutes with 5 fold, which was spectacular. Did everything on TPU, but at this point, I think it would be outstandingly difficult to beat the 0.233 baseline with only Kaggle resources. The baseline has a lot to tune, but without insane resources, it is ridiculously difficult to improve it (my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch). Try out Colab if you're running out of GPU hours though.</p>",
      "rawMarkdown": "We *were* silver before the public notebook (30th or so). We used Efficientnetb0 trained on tfrecords and TPU, and had an epoch time of about 5 minutes with 5 fold, which was spectacular. Did everything on TPU, but at this point, I think it would be outstandingly difficult to beat the 0.233 baseline with only Kaggle resources. The baseline has a lot to tune, but without insane resources, it is ridiculously difficult to improve it (my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch). Try out Colab if you're running out of GPU hours though.",
      "replies": [
        {
          "id": 1057456,
          "postDate": "2020-10-22T16:52:56.187Z",
          "content": "<p>Wow. Great work Stanley. I’m sure pytorch must be able to get close to those speeds on TPU right? Or are TFrecords really so much faster?</p>",
          "rawMarkdown": "Wow. Great work Stanley. I’m sure pytorch must be able to get close to those speeds on TPU right? Or are TFrecords really so much faster?",
          "votes": 1
        },
        {
          "id": 1057468,
          "postDate": "2020-10-22T17:07:52.020Z",
          "content": "<p>I'm not sure, I really haven't done an insanely large competition like this with pytorch. I believe that they are quite a bit faster than pytorch, since I/O is decreased (reading 30 2gb files) instead of thousands of 1mb files. Preprocessing and stratified validation is already completed, so any CPU bottlenecks are minimized also. All the CPU needs to do is split the tfrecords into groups and read them. As far as I know, Kaggle TPU's are mainly CPU bottlenecked, since feeding 128gb of RAM is quite intensive. </p>",
          "rawMarkdown": "I'm not sure, I really haven't done an insanely large competition like this with pytorch. I believe that they are quite a bit faster than pytorch, since I/O is decreased (reading 30 2gb files) instead of thousands of 1mb files. Preprocessing and stratified validation is already completed, so any CPU bottlenecks are minimized also. All the CPU needs to do is split the tfrecords into groups and read them. As far as I know, Kaggle TPU's are mainly CPU bottlenecked, since feeding 128gb of RAM is quite intensive. ",
          "votes": 1
        },
        {
          "id": 1057573,
          "postDate": "2020-10-22T19:08:31.730Z",
          "content": "<p>Have you used one stage, or two stage model?</p>",
          "rawMarkdown": "Have you used one stage, or two stage model?",
          "votes": 1
        },
        {
          "id": 1057579,
          "postDate": "2020-10-22T19:13:57.103Z",
          "content": "<p>This was only with regular CNN trained on image and exam level - stacking with GRU/LSTM would take longer, and separating them would reduce individual training time slightly. The big limitation here is the max TPU time of 3 hours, so even if you had 5 minutes per epoch, you still only get 36 epochs, so 7 epochs for 5 fold, which is not very much. Better to just do each fold separately</p>",
          "rawMarkdown": "This was only with regular CNN trained on image and exam level - stacking with GRU/LSTM would take longer, and separating them would reduce individual training time slightly. The big limitation here is the max TPU time of 3 hours, so even if you had 5 minutes per epoch, you still only get 36 epochs, so 7 epochs for 5 fold, which is not very much. Better to just do each fold separately",
          "votes": 2,
          "replies": [
            {
              "id": 1057615,
              "postDate": "2020-10-22T19:52:08.747Z",
              "content": "<p>I have never  used Kaggle TPU for pytorch.  Colab Pro has much better config for Pytorch.  And of course no time limitation. </p>\n<p>I am training on TPU with a non effnet based model but bigger than b0, b1, b2. </p>\n<p>But I think it's a big mistake.  I won't likely be able to do all inference in 9 hours.</p>\n<blockquote>\n  <p>my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch</p>\n</blockquote>\n<p>TPU is the way to go on colab Pro in this case.  As I said , I'm training a bigger model than b0 with less than 2 hour per epoch on all the data and heavy augmentations. <br>\nBig batch size shared between 8 TPU cores and you have more ram/CPU/disk space on TPU than GPU</p>",
              "rawMarkdown": "I have never  used Kaggle TPU for pytorch.  Colab Pro has much better config for Pytorch.  And of course no time limitation. \n\nI am training on TPU with a non effnet based model but bigger than b0, b1, b2. \n\nBut I think it's a big mistake.  I won't likely be able to do all inference in 9 hours.\n\n\n> my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch\n\nTPU is the way to go on colab Pro in this case.  As I said , I'm training a bigger model than b0 with less than 2 hour per epoch on all the data and heavy augmentations. \nBig batch size shared between 8 TPU cores and you have more ram/CPU/disk space on TPU than GPU",
              "votes": 1
            }
          ]
        },
        {
          "id": 1057628,
          "postDate": "2020-10-22T20:00:36.450Z",
          "content": "<p>I agree with your opinion about TPU -- 3 hours is too little for such a big task.</p>",
          "rawMarkdown": "I agree with your opinion about TPU -- 3 hours is too little for such a big task.",
          "votes": 1
        },
        {
          "id": 1057629,
          "postDate": "2020-10-22T20:01:10.853Z",
          "content": "<p>Have you tried stacking with LGBM trained on meta features?</p>",
          "rawMarkdown": "Have you tried stacking with LGBM trained on meta features?",
          "votes": 1
        },
        {
          "id": 1058557,
          "postDate": "2020-10-23T21:32:19.370Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1057625,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-10-22T19:58:12.273000",
      "content": "<blockquote>\n  <p>My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models? If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?</p>\n</blockquote>\n<p>You can still use Colab Pro,  it's way much cheaper than GCP;  That's what I do whenever I need extensive resources</p>\n<p>And for GPU, colab Pro has now  <code>Tesla V100-sxm2</code> option, which is way faster than kaggle P100</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1057650,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-22T20:25:27.230000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057654,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-22T20:30:38.010000",
          "content": "<p>Yeah High RAM.</p>\n<p>Once I get P100, I keep retinitializing the notebook, until I get V100 ^^</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057705,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-10-22T22:38:15.227000",
          "content": "<p>Colab pro is only available in US, so that option doesn't work for many of us. Is a great advantage for those who can use that opportunity. Just to wait and hope for a release in Europe.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057951,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-10-23T07:20:14.053000",
          "content": "<p>Yeah this is kind of my point - if you can only do well in this competition by using external resources (own GPUs, Colab Pro etc) that either cost money or are only available to USA participants is it really a fair competition?  Shouldn't it be about knowledge rather than availability of resources?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057978,
          "author_name": "Salaryman",
          "author_url": "",
          "post_date": "2020-10-23T08:09:35.413000",
          "content": "<p>It is quite common that $ prize/ gold medal winners trained their models entirely by Kaggle Kernel/ Colab. Hardware can make a difference but more important is your knowledge, experience and \"luck\".</p>\n<p>I think as long as there is no un-disclosed external data source and private sharing, the competition is considered as fair. </p>\n<p>Will become a philosophy question if drill too deep about \"fairness\". If both Tom and Jerry are joining this competition but Tom is a machine learning PhD while Jerry is just a chemistry graduate, is this fair even they have the same hardware?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1058031,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-10-23T09:09:57.507000",
          "content": "<p>One can make it simple, in general, and to not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058067,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-23T09:53:35.093000",
          "content": "<blockquote>\n  <p>Shouldn't it be about knowledge rather than availability of resources?</p>\n</blockquote>\n<p>And why not both ? </p>\n<p>As far as i'm concerned, I don't have even own single GPU and Colab Pro is available everywhere not only the USA (I suscribed from Paris in the early days). </p>\n<p>As for fairness, it's quite relative . Because you can also say even Kaggle itself is not fair,  as it gives +30 hours of GPU/TPU  to each team member for people teaming up. </p>\n<p>Anyway this competition requires extensive resources by its nature and by the size of the dataset.  You can't help but finding solutions like teaming up, using online services or buying your own solid infrastructure.  Otherwise there are still other competitions that require much less resources</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1058071,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-23T10:00:31.127000",
          "content": "<blockquote>\n  <p>One can make it simple, in general, and do not include the extremes, having access to unlimited TPU and V100, gives some advantage in machine learning engineering.</p>\n</blockquote>\n<p>I propose we all use 2 cores CPU only and build every thing from scratch. Because, even using Pytorch instead of writing own CNN and Backpropagation algorithms, give some advantages. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058111,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-10-23T10:59:28.167000",
          "content": "<p>I thought we were talking about hw. I can only talk for my self, if I had the possible to use more TPU and GPU time through Colab Pro (V100 is really a good option that they have added, using Volta+mixed precision 👍) I would use it to help tuning and experiment more and wider, but unfortunately it says only available in US and Canada. Would it help me make a better model or getting a better scores? Who knows, that's more a factor of knowledge but I get that knowledge through testing and experimenting, thats why one loves Kaggle and Colab, resources and challenges helping one getting better in different AI areas through helping and contribute solving a problem and with the goal to help and contribute solving problems even better :)<br>\nI hope they make Colab Pro more available, that's my point.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1058953,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-10-24T12:53:53.883000",
          "content": "<p>As I said, Colab Pro is available outside of these countries. Don't mind what is written at the homepage. </p>\n<p>I'm just giving advices of the most affordable  resources to common kagglers in such highly demanding competition.  You can be sure many people at (top) LB use much more extensive resources. </p>\n<p>It's (nearly) impossible to win such competition by using Kaggle only, just like the annual Youtube-8m video competition (+ 1.5 TB of data last year) and many others google competitions. <br>\nThat's why kaggle (and sponsors)  offer GCP credits on most of these competitions, in particular for those who can't afford high computational power</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1059061,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-10-24T15:09:50.630000",
          "content": "<p>Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1059084,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-10-24T15:40:51.823000",
          "content": "<p>Thanks Serigne and all.  I think its fair overall.  Those at the very top are probably those who have the combination of knowledge and resources plus the technical skills to properly utilise those resources. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1059176,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-10-24T18:04:53.123000",
          "content": "<p>Yes and it's a great advice Serigne 👍 So you think it's OK to use outside US and Canada? I have to look into it again then, maybe something has changed :)<br>\nAnd I'm also greatfull that we can use Kaggle and Colab, opens the door for everyone to learn and contribute.  So I don't want to sound ungrateful but it's just that the Colab Pro seems like a great subscribe service and open up for even more powerful testing and training for everyone. They have now added another country to the list, so maybe more will come soon if it is not possible to use.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1058657,
      "author_name": "Vivek",
      "author_url": "",
      "post_date": "2020-10-24T04:05:39.717000",
      "content": "<p>You can try rapid from nvidia for using gpu power while matrix calculations and preporcessing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1057436,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2020-10-22T16:32:39.863000",
      "content": "<p>We <em>were</em> silver before the public notebook (30th or so). We used Efficientnetb0 trained on tfrecords and TPU, and had an epoch time of about 5 minutes with 5 fold, which was spectacular. Did everything on TPU, but at this point, I think it would be outstandingly difficult to beat the 0.233 baseline with only Kaggle resources. The baseline has a lot to tune, but without insane resources, it is ridiculously difficult to improve it (my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch). Try out Colab if you're running out of GPU hours though.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1057456,
          "author_name": "Mark P",
          "author_url": "",
          "post_date": "2020-10-22T16:52:56.187000",
          "content": "<p>Wow. Great work Stanley. I’m sure pytorch must be able to get close to those speeds on TPU right? Or are TFrecords really so much faster?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057468,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-22T17:07:52.020000",
          "content": "<p>I'm not sure, I really haven't done an insanely large competition like this with pytorch. I believe that they are quite a bit faster than pytorch, since I/O is decreased (reading 30 2gb files) instead of thousands of 1mb files. Preprocessing and stratified validation is already completed, so any CPU bottlenecks are minimized also. All the CPU needs to do is split the tfrecords into groups and read them. As far as I know, Kaggle TPU's are mainly CPU bottlenecked, since feeding 128gb of RAM is quite intensive. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057573,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-10-22T19:08:31.730000",
          "content": "<p>Have you used one stage, or two stage model?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057579,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2020-10-22T19:13:57.103000",
          "content": "<p>This was only with regular CNN trained on image and exam level - stacking with GRU/LSTM would take longer, and separating them would reduce individual training time slightly. The big limitation here is the max TPU time of 3 hours, so even if you had 5 minutes per epoch, you still only get 36 epochs, so 7 epochs for 5 fold, which is not very much. Better to just do each fold separately</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1057615,
              "author_name": "Serigne ",
              "author_url": "",
              "post_date": "2020-10-22T19:52:08.747000",
              "content": "<p>I have never  used Kaggle TPU for pytorch.  Colab Pro has much better config for Pytorch.  And of course no time limitation. </p>\n<p>I am training on TPU with a non effnet based model but bigger than b0, b1, b2. </p>\n<p>But I think it's a big mistake.  I won't likely be able to do all inference in 9 hours.</p>\n<blockquote>\n  <p>my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch</p>\n</blockquote>\n<p>TPU is the way to go on colab Pro in this case.  As I said , I'm training a bigger model than b0 with less than 2 hour per epoch on all the data and heavy augmentations. <br>\nBig batch size shared between 8 TPU cores and you have more ram/CPU/disk space on TPU than GPU</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1057628,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-10-22T20:00:36.450000",
          "content": "<p>I agree with your opinion about TPU -- 3 hours is too little for such a big task.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057629,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2020-10-22T20:01:10.853000",
          "content": "<p>Have you tried stacking with LGBM trained on meta features?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1058557,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-23T21:32:19.370000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1057411": "Hi all,\n\nI struggled for a long time to train an Effnet on all the data (Ian Pan's Jpeg dataset) in pytorch.  I used the GPUs on kaggle and got it down to 45mins/epoch but thought this was still far too long.  I then tried using the TPUs and got it running but the times were the same so I think my bottleneck was the CPU.\n\nBut reading in detail through @khyeh0719 training notebooks I see his training took many many GPU hours too and I think he went to GCP to train some parts.\n\nMy question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models?  If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?\n\nor is everyone using TF and/or TPUs and its just me struggling to get pytorch working that is the problem?   [feeling it is probably<---]",
    "1057625": "> My question is then is everybody in the gold and silver positions having to resort to using compute outside of kaggle to train their models? If so doesn't that rather undermine the purpose of the kaggle platform if we have to do that?\n\n\n\nYou can still use Colab Pro,  it's way much cheaper than GCP;  That's what I do whenever I need extensive resources\n\nAnd for GPU, colab Pro has now  `Tesla V100-sxm2` option, which is way faster than kaggle P100",
    "1058657": "You can try rapid from nvidia for using gpu power while matrix calculations and preporcessing.",
    "1057436": "We *were* silver before the public notebook (30th or so). We used Efficientnetb0 trained on tfrecords and TPU, and had an epoch time of about 5 minutes with 5 fold, which was spectacular. Did everything on TPU, but at this point, I think it would be outstandingly difficult to beat the 0.233 baseline with only Kaggle resources. The baseline has a lot to tune, but without insane resources, it is ridiculously difficult to improve it (my teammate tried to train on Colab Pro GPU and it predicted 8 hours per epoch). Try out Colab if you're running out of GPU hours though."
  }
}