{
  "id": 410548,
  "title": "How to speed things up ?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/410548",
  "author_name": "iceman273k",
  "post_date": "2023-05-15T18:14:49.205000",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone !<br>\nthis is my very first feature competition, I'm brand new to machine and deep learning, and I'm beginning to realize that this particular competition is a hell of a start… (dataset size, semantic segmentation + time frames, not forgetting the training time limit… and the terribly low number of participants is kind of confirming the size of the challenge)<br>\nI would really appreciate some piece of advice, at least to be sure I'm not on the wrong track.<br>\nI'm going with tensorflow, planning on training a Unet (I'll start with the one provided on Keras as an example by F. Chollet). To cut down complexity a bit, I decided not to take into account the 4+3 timeframes around the masked one.<br>\nRight now, my main concern is the time it takes to train a model :<br>\nI already have a custom Sequence to load the data in batches of 32 instances, and I'm just fiddling with a one layer model just to see if everything fits together, using a custom Dice loss function.<br>\nAnd the average training time over one epoch is… almost 2 hours. RAM is almost full, and CPU usage is curiously low.<br>\nI fear the moment when I'll start training a real Unet model with 30+ layers over 10s of epochs…</p>\n<p>Anyway, for all the TF users here, what did you do (plan to do) to boost the training time ?</p>",
  "messages": [
    {
      "id": 2276863,
      "postDate": "2023-05-27T09:11:24.037Z",
      "content": "<p>Many people including myself started with Keras. Learn from the mistakes of others and switch pytorch. This comp already has a strong baseline torch code LB0.494. Just trying to reduce suffering in this word…</p>",
      "rawMarkdown": "Many people including myself started with Keras. Learn from the mistakes of others and switch pytorch. This comp already has a strong baseline torch code LB0.494. Just trying to reduce suffering in this word..."
    },
    {
      "id": 2260745,
      "postDate": "2023-05-15T20:44:58.817Z",
      "content": "<p>Hey iceman, <br>\nwelcome to this challenge! Currently, I am pursuing a similar approach than you do.<br>\nI am also training a UNET with tensorflow. For my pipeline, I am using tf.data.Dataset, since it allows to parallelize the data loading and training (If you are not familiar with this, check this link <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset)\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset)</a>. I am also only using the masked image, not the ones before and after. My training time per epoch lies between 5 and 15 minutes, depending on the \"depth\" (amount of filters and layers) I use for the UNET. </p>\n<p>And what do you mean with the \"preprint\" you mentioned in your comment? Where can i find this information?</p>\n<p>Thanks and good luck!</p>",
      "rawMarkdown": "Hey iceman, \nwelcome to this challenge! Currently, I am pursuing a similar approach than you do.\nI am also training a UNET with tensorflow. For my pipeline, I am using tf.data.Dataset, since it allows to parallelize the data loading and training (If you are not familiar with this, check this link https://www.tensorflow.org/api_docs/python/tf/data/Dataset). I am also only using the masked image, not the ones before and after. My training time per epoch lies between 5 and 15 minutes, depending on the \"depth\" (amount of filters and layers) I use for the UNET. \n\nAnd what do you mean with the \"preprint\" you mentioned in your comment? Where can i find this information?\n\nThanks and good luck!",
      "replies": [
        {
          "id": 2260778,
          "postDate": "2023-05-15T21:28:07.140Z",
          "content": "<p>Hey Jan,<br>\nthanks for your reply !<br>\nYou have the preprint of the research here : <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">https://arxiv.org/abs/2304.02122</a> (download the pdf at the top right corner of the page)<br>\nThey went with a modified ResNet.<br>\nConcerning your training time, this is what I expected too, but 2 hours for just a single layer, damn ! Something is clearly wrong somewhere. I was thinking about moving to a Dataset, but I don't see how to avoid going through a np.load() to get the images in the shape I want.<br>\nFor your Unet, you build it from scratch or do you do transfer learning ?</p>\n<p>Good luck to you too !</p>",
          "rawMarkdown": "Hey Jan,\nthanks for your reply !\nYou have the preprint of the research here : https://arxiv.org/abs/2304.02122 (download the pdf at the top right corner of the page)\nThey went with a modified ResNet.\nConcerning your training time, this is what I expected too, but 2 hours for just a single layer, damn ! Something is clearly wrong somewhere. I was thinking about moving to a Dataset, but I don't see how to avoid going through a np.load() to get the images in the shape I want.\nFor your Unet, you build it from scratch or do you do transfer learning ?\n\nGood luck to you too !",
          "replies": [
            {
              "id": 2260793,
              "postDate": "2023-05-15T22:02:36.457Z",
              "content": "<p>Hey,</p>\n<p>I used an implementation which was available publicly but trained it from scratch.<br>\nDid you already consider training on GPU? On CPU, cnn training is quite slow whereas on GPU it is a lot faster!</p>",
              "rawMarkdown": "Hey,\n\nI used an implementation which was available publicly but trained it from scratch.\nDid you already consider training on GPU? On CPU, cnn training is quite slow whereas on GPU it is a lot faster!"
            },
            {
              "id": 2260798,
              "postDate": "2023-05-15T22:07:10.033Z",
              "content": "<p>Well yes I tried, but it didn't seem to work, I had lots of warning telling me something like 'cuda not responding' (I honestly don't recall exactly the warnings)<br>\nAnd I'm going for a scratch training too ;-)</p>",
              "rawMarkdown": "Well yes I tried, but it didn't seem to work, I had lots of warning telling me something like 'cuda not responding' (I honestly don't recall exactly the warnings)\nAnd I'm going for a scratch training too ;-)",
              "votes": 1
            }
          ]
        },
        {
          "id": 2265567,
          "postDate": "2023-05-19T10:44:02.900Z",
          "content": "<p>Hey Jan,<br>\nI come back to you since I'm lost in implementation hell with these damn tf.dataset. I tried using from_generator, list_file (and then mapping a numpy_function or a py_function), and either this is still very slow, or I lose the shape of the tensors during the process… Could you give me a hint on how you worked your way around ? (congrats for your first results, by the way !)</p>",
          "rawMarkdown": "Hey Jan,\nI come back to you since I'm lost in implementation hell with these damn tf.dataset. I tried using from_generator, list_file (and then mapping a numpy_function or a py_function), and either this is still very slow, or I lose the shape of the tensors during the process... Could you give me a hint on how you worked your way around ? (congrats for your first results, by the way !)",
          "replies": [
            {
              "id": 2267132,
              "postDate": "2023-05-20T16:04:04.230Z",
              "content": "<p>Hey Iceman,</p>\n<p>I created the tf Dataset using the paths to the record ids folder and then make a \"from_tensor_sclices\" dataset. Then i mapped a function for loading the bands and masks (you need to use tf pyfunc) and processing them. Hope this helps!</p>",
              "rawMarkdown": "Hey Iceman,\n\nI created the tf Dataset using the paths to the record ids folder and then make a \"from_tensor_sclices\" dataset. Then i mapped a function for loading the bands and masks (you need to use tf pyfunc) and processing them. Hope this helps!",
              "votes": 1
            },
            {
              "id": 2269682,
              "postDate": "2023-05-22T16:13:41.673Z",
              "content": "<p>Hi Jan,<br>\nthanks for your reply ! (and sorry for the delay) Yes, it helps, by confirming I was on the right track, but still experiencing long loading time issues. I went with a dataset list_files, which I then map with a numpy_function to extract the data from each .npy file. I might try with a py_func instead, maybe it's more efficient.</p>",
              "rawMarkdown": "Hi Jan,\nthanks for your reply ! (and sorry for the delay) Yes, it helps, by confirming I was on the right track, but still experiencing long loading time issues. I went with a dataset list_files, which I then map with a numpy_function to extract the data from each .npy file. I might try with a py_func instead, maybe it's more efficient.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2260626,
      "postDate": "2023-05-15T18:59:28.387Z",
      "content": "<p>Little side note :<br>\nI just read the preprint (finally), and at least, the intuition of dropping the timeframes seems ok, at least for the beginning, since it boosts there AUC by just ~2 (on a score of ~70). And they conclude that the image BEFORE have more impact on the result than the image AFTER. And it seems that two of the bands are the difference between two couples of other bands. Wouldn't it be worth it of discarding them, then, to avoid processing the 'same' informations ?</p>",
      "rawMarkdown": "Little side note :\nI just read the preprint (finally), and at least, the intuition of dropping the timeframes seems ok, at least for the beginning, since it boosts there AUC by just ~2 (on a score of ~70). And they conclude that the image BEFORE have more impact on the result than the image AFTER. And it seems that two of the bands are the difference between two couples of other bands. Wouldn't it be worth it of discarding them, then, to avoid processing the 'same' informations ?"
    },
    {
      "id": 2260566,
      "postDate": "2023-05-15T18:14:49.207Z",
      "content": "<p>Hi everyone !<br>\nthis is my very first feature competition, I'm brand new to machine and deep learning, and I'm beginning to realize that this particular competition is a hell of a start… (dataset size, semantic segmentation + time frames, not forgetting the training time limit… and the terribly low number of participants is kind of confirming the size of the challenge)<br>\nI would really appreciate some piece of advice, at least to be sure I'm not on the wrong track.<br>\nI'm going with tensorflow, planning on training a Unet (I'll start with the one provided on Keras as an example by F. Chollet). To cut down complexity a bit, I decided not to take into account the 4+3 timeframes around the masked one.<br>\nRight now, my main concern is the time it takes to train a model :<br>\nI already have a custom Sequence to load the data in batches of 32 instances, and I'm just fiddling with a one layer model just to see if everything fits together, using a custom Dice loss function.<br>\nAnd the average training time over one epoch is… almost 2 hours. RAM is almost full, and CPU usage is curiously low.<br>\nI fear the moment when I'll start training a real Unet model with 30+ layers over 10s of epochs…</p>\n<p>Anyway, for all the TF users here, what did you do (plan to do) to boost the training time ?</p>",
      "rawMarkdown": "Hi everyone !\nthis is my very first feature competition, I'm brand new to machine and deep learning, and I'm beginning to realize that this particular competition is a hell of a start... (dataset size, semantic segmentation + time frames, not forgetting the training time limit... and the terribly low number of participants is kind of confirming the size of the challenge)\nI would really appreciate some piece of advice, at least to be sure I'm not on the wrong track.\nI'm going with tensorflow, planning on training a Unet (I'll start with the one provided on Keras as an example by F. Chollet). To cut down complexity a bit, I decided not to take into account the 4+3 timeframes around the masked one.\nRight now, my main concern is the time it takes to train a model :\nI already have a custom Sequence to load the data in batches of 32 instances, and I'm just fiddling with a one layer model just to see if everything fits together, using a custom Dice loss function.\nAnd the average training time over one epoch is... almost 2 hours. RAM is almost full, and CPU usage is curiously low.\nI fear the moment when I'll start training a real Unet model with 30+ layers over 10s of epochs...\n\nAnyway, for all the TF users here, what did you do (plan to do) to boost the training time ?"
    }
  ],
  "comments": [
    {
      "id": 2276863,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "2023-05-27T09:11:24.037000",
      "content": "<p>Many people including myself started with Keras. Learn from the mistakes of others and switch pytorch. This comp already has a strong baseline torch code LB0.494. Just trying to reduce suffering in this word…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2260745,
      "author_name": "Jan H",
      "author_url": "",
      "post_date": "2023-05-15T20:44:58.817000",
      "content": "<p>Hey iceman, <br>\nwelcome to this challenge! Currently, I am pursuing a similar approach than you do.<br>\nI am also training a UNET with tensorflow. For my pipeline, I am using tf.data.Dataset, since it allows to parallelize the data loading and training (If you are not familiar with this, check this link <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset)\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset)</a>. I am also only using the masked image, not the ones before and after. My training time per epoch lies between 5 and 15 minutes, depending on the \"depth\" (amount of filters and layers) I use for the UNET. </p>\n<p>And what do you mean with the \"preprint\" you mentioned in your comment? Where can i find this information?</p>\n<p>Thanks and good luck!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2260778,
          "author_name": "iceman273k",
          "author_url": "",
          "post_date": "2023-05-15T21:28:07.140000",
          "content": "<p>Hey Jan,<br>\nthanks for your reply !<br>\nYou have the preprint of the research here : <a href=\"https://arxiv.org/abs/2304.02122\" target=\"_blank\">https://arxiv.org/abs/2304.02122</a> (download the pdf at the top right corner of the page)<br>\nThey went with a modified ResNet.<br>\nConcerning your training time, this is what I expected too, but 2 hours for just a single layer, damn ! Something is clearly wrong somewhere. I was thinking about moving to a Dataset, but I don't see how to avoid going through a np.load() to get the images in the shape I want.<br>\nFor your Unet, you build it from scratch or do you do transfer learning ?</p>\n<p>Good luck to you too !</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2260793,
              "author_name": "Jan H",
              "author_url": "",
              "post_date": "2023-05-15T22:02:36.457000",
              "content": "<p>Hey,</p>\n<p>I used an implementation which was available publicly but trained it from scratch.<br>\nDid you already consider training on GPU? On CPU, cnn training is quite slow whereas on GPU it is a lot faster!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2260798,
              "author_name": "iceman273k",
              "author_url": "",
              "post_date": "2023-05-15T22:07:10.033000",
              "content": "<p>Well yes I tried, but it didn't seem to work, I had lots of warning telling me something like 'cuda not responding' (I honestly don't recall exactly the warnings)<br>\nAnd I'm going for a scratch training too ;-)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2265567,
          "author_name": "iceman273k",
          "author_url": "",
          "post_date": "2023-05-19T10:44:02.900000",
          "content": "<p>Hey Jan,<br>\nI come back to you since I'm lost in implementation hell with these damn tf.dataset. I tried using from_generator, list_file (and then mapping a numpy_function or a py_function), and either this is still very slow, or I lose the shape of the tensors during the process… Could you give me a hint on how you worked your way around ? (congrats for your first results, by the way !)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2267132,
              "author_name": "Jan H",
              "author_url": "",
              "post_date": "2023-05-20T16:04:04.230000",
              "content": "<p>Hey Iceman,</p>\n<p>I created the tf Dataset using the paths to the record ids folder and then make a \"from_tensor_sclices\" dataset. Then i mapped a function for loading the bands and masks (you need to use tf pyfunc) and processing them. Hope this helps!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2269682,
              "author_name": "iceman273k",
              "author_url": "",
              "post_date": "2023-05-22T16:13:41.673000",
              "content": "<p>Hi Jan,<br>\nthanks for your reply ! (and sorry for the delay) Yes, it helps, by confirming I was on the right track, but still experiencing long loading time issues. I went with a dataset list_files, which I then map with a numpy_function to extract the data from each .npy file. I might try with a py_func instead, maybe it's more efficient.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2260626,
      "author_name": "iceman273k",
      "author_url": "",
      "post_date": "2023-05-15T18:59:28.387000",
      "content": "<p>Little side note :<br>\nI just read the preprint (finally), and at least, the intuition of dropping the timeframes seems ok, at least for the beginning, since it boosts there AUC by just ~2 (on a score of ~70). And they conclude that the image BEFORE have more impact on the result than the image AFTER. And it seems that two of the bands are the difference between two couples of other bands. Wouldn't it be worth it of discarding them, then, to avoid processing the 'same' informations ?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2276863": "Many people including myself started with Keras. Learn from the mistakes of others and switch pytorch. This comp already has a strong baseline torch code LB0.494. Just trying to reduce suffering in this word...",
    "2260745": "Hey iceman, \nwelcome to this challenge! Currently, I am pursuing a similar approach than you do.\nI am also training a UNET with tensorflow. For my pipeline, I am using tf.data.Dataset, since it allows to parallelize the data loading and training (If you are not familiar with this, check this link https://www.tensorflow.org/api_docs/python/tf/data/Dataset). I am also only using the masked image, not the ones before and after. My training time per epoch lies between 5 and 15 minutes, depending on the \"depth\" (amount of filters and layers) I use for the UNET. \n\nAnd what do you mean with the \"preprint\" you mentioned in your comment? Where can i find this information?\n\nThanks and good luck!",
    "2260626": "Little side note :\nI just read the preprint (finally), and at least, the intuition of dropping the timeframes seems ok, at least for the beginning, since it boosts there AUC by just ~2 (on a score of ~70). And they conclude that the image BEFORE have more impact on the result than the image AFTER. And it seems that two of the bands are the difference between two couples of other bands. Wouldn't it be worth it of discarding them, then, to avoid processing the 'same' informations ?",
    "2260566": "Hi everyone !\nthis is my very first feature competition, I'm brand new to machine and deep learning, and I'm beginning to realize that this particular competition is a hell of a start... (dataset size, semantic segmentation + time frames, not forgetting the training time limit... and the terribly low number of participants is kind of confirming the size of the challenge)\nI would really appreciate some piece of advice, at least to be sure I'm not on the wrong track.\nI'm going with tensorflow, planning on training a Unet (I'll start with the one provided on Keras as an example by F. Chollet). To cut down complexity a bit, I decided not to take into account the 4+3 timeframes around the masked one.\nRight now, my main concern is the time it takes to train a model :\nI already have a custom Sequence to load the data in batches of 32 instances, and I'm just fiddling with a one layer model just to see if everything fits together, using a custom Dice loss function.\nAnd the average training time over one epoch is... almost 2 hours. RAM is almost full, and CPU usage is curiously low.\nI fear the moment when I'll start training a real Unet model with 30+ layers over 10s of epochs...\n\nAnyway, for all the TF users here, what did you do (plan to do) to boost the training time ?"
  }
}