{
  "id": 193183,
  "title": "the competition is closing to an end and my brain is getting hard fried.",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193183",
  "author_name": "DarkCube",
  "post_date": "2020-10-25T17:27:46.668000",
  "votes": 6,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I commit once, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. </p>\n<p>I run a code cell in colab, I go do some stuff. I come back. the code cell is requesting my input. There is no way for me to enter anything. I have to terminate the session and repeat a 45 min download.</p>\n<p>I load a model in a notebook and everything is fine. I load in another notebook and a warning shows up. I load from Colab and a warning and an error show up. And <strong><em>this</em></strong>, this really messed shit up. I knew I had no chance of training the feature extracted from scratch. I was counting on <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> pretrained models. (I was going to train a better model with a mix of softmax and sigmoid activations and undersampling the negative exams but it was better than nothing). but there is a catch. I use TensorFlow. I had to convert them back from torch which wasted an entire day. now that model doesn't work. So there is that.</p>\n<p>was so excited to participate in this competition. I had this idea of using an attention model that has been used in machine translation (which I found was nowhere discussed) to predict exam and image level labels and I was convinced that it would be a great solution, especially after the \"baseline\" that used a really simple GRU coming straight up from last year's solution. But it seems I won't be able to submit in time.</p>\n<p>There lies my worthless idea. It may help someone. Although I really doubt that.</p>\n<p>I have only skimmed through the rules so I'm really sorry if this kind of posts isn't allowed. I just had to post my frustration somewhere. Just tell me and I'll delete it.</p>",
  "messages": [
    {
      "id": 1060023,
      "postDate": "2020-10-25T17:27:46.667Z",
      "content": "<p>I commit once, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. </p>\n<p>I run a code cell in colab, I go do some stuff. I come back. the code cell is requesting my input. There is no way for me to enter anything. I have to terminate the session and repeat a 45 min download.</p>\n<p>I load a model in a notebook and everything is fine. I load in another notebook and a warning shows up. I load from Colab and a warning and an error show up. And <strong><em>this</em></strong>, this really messed shit up. I knew I had no chance of training the feature extracted from scratch. I was counting on <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> pretrained models. (I was going to train a better model with a mix of softmax and sigmoid activations and undersampling the negative exams but it was better than nothing). but there is a catch. I use TensorFlow. I had to convert them back from torch which wasted an entire day. now that model doesn't work. So there is that.</p>\n<p>was so excited to participate in this competition. I had this idea of using an attention model that has been used in machine translation (which I found was nowhere discussed) to predict exam and image level labels and I was convinced that it would be a great solution, especially after the \"baseline\" that used a really simple GRU coming straight up from last year's solution. But it seems I won't be able to submit in time.</p>\n<p>There lies my worthless idea. It may help someone. Although I really doubt that.</p>\n<p>I have only skimmed through the rules so I'm really sorry if this kind of posts isn't allowed. I just had to post my frustration somewhere. Just tell me and I'll delete it.</p>",
      "rawMarkdown": "I commit once, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. \n\nI run a code cell in colab, I go do some stuff. I come back. the code cell is requesting my input. There is no way for me to enter anything. I have to terminate the session and repeat a 45 min download.\n\nI load a model in a notebook and everything is fine. I load in another notebook and a warning shows up. I load from Colab and a warning and an error show up. And ***this***, this really messed shit up. I knew I had no chance of training the feature extracted from scratch. I was counting on @khyeh0719 pretrained models. (I was going to train a better model with a mix of softmax and sigmoid activations and undersampling the negative exams but it was better than nothing). but there is a catch. I use TensorFlow. I had to convert them back from torch which wasted an entire day. now that model doesn't work. So there is that.\n\nwas so excited to participate in this competition. I had this idea of using an attention model that has been used in machine translation (which I found was nowhere discussed) to predict exam and image level labels and I was convinced that it would be a great solution, especially after the \"baseline\" that used a really simple GRU coming straight up from last year's solution. But it seems I won't be able to submit in time.\n\nThere lies my worthless idea. It may help someone. Although I really doubt that.\n\nI have only skimmed through the rules so I'm really sorry if this kind of posts isn't allowed. I just had to post my frustration somewhere. Just tell me and I'll delete it.",
      "votes": 7
    },
    {
      "id": 1060044,
      "postDate": "2020-10-25T17:59:56Z",
      "content": "<p>Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow, or beating the baseline easily 😂<br>\nThere is a reason for most of Kaggle's opinions, like converting pytorch to tensorflow, or the baseline being difficult to train. </p>\n<blockquote>\n  <p>I commit one, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. </p>\n</blockquote>\n<p>Try random seed?</p>\n<p>Better luck next competition though, it is outstandingly hard to do a competition in a week, and I hope to see you in the next competitions. I echo the sentiment of the other commenter - Kaggle is often difficult and frustrating, but that's what makes it so rewarding when you do well.</p>",
      "rawMarkdown": "Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow, or beating the baseline easily 😂\nThere is a reason for most of Kaggle's opinions, like converting pytorch to tensorflow, or the baseline being difficult to train. \n\n> I commit one, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. \n\nTry random seed?\n\nBetter luck next competition though, it is outstandingly hard to do a competition in a week, and I hope to see you in the next competitions. I echo the sentiment of the other commenter - Kaggle is often difficult and frustrating, but that's what makes it so rewarding when you do well.",
      "votes": 3,
      "replies": [
        {
          "id": 1060051,
          "postDate": "2020-10-25T18:07:19.760Z",
          "content": "<blockquote>\n  <p>Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow</p>\n</blockquote>\n<p>I did convert the model. it just doesn't work without GPU. and It doesn't work in colab. And the input configuration is (C, H, W) as in torch instead of (H, W, C) as in TF.</p>\n<blockquote>\n  <p>Better luck next competition though, it is outstandingly hard to do a competition in a week</p>\n</blockquote>\n<p>Is that a challenge?</p>",
          "rawMarkdown": "> Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow\n\nI did convert the model. it just doesn't work without GPU. and It doesn't work in colab. And the input configuration is (C, H, W) as in torch instead of (H, W, C) as in TF.\n\n  > Better luck next competition though, it is outstandingly hard to do a competition in a week\n\nIs that a challenge?",
          "votes": 2
        },
        {
          "id": 1060075,
          "postDate": "2020-10-25T18:41:03.277Z",
          "content": "<p>thats the reason why everyone is shifting towards pytorch  these days .</p>",
          "rawMarkdown": "thats the reason why everyone is shifting towards pytorch  these days ."
        },
        {
          "id": 1060091,
          "postDate": "2020-10-25T18:55:07.060Z",
          "content": "<p>wdym? is it because you can't convert the model from torch to tensorflow? if anything that should be a reason for people to switch to torch. from what I understood torch is primarily for research whereas TF is for both research and production. </p>",
          "rawMarkdown": "wdym? is it because you can't convert the model from torch to tensorflow? if anything that should be a reason for people to switch to torch. from what I understood torch is primarily for research whereas TF is for both research and production. ",
          "votes": -3
        }
      ]
    },
    {
      "id": 1061036,
      "postDate": "2020-10-26T17:50:31.680Z",
      "content": "<p>While I can understand your 'frustrations' up to a certain point…some of the things you mention will/can occur in any competition you enter :-) See them as issues/frustrations or something you deal with and then solve them. That usually works best ;-)</p>\n<p>So the competition ends today … but why would you not try to finish your idea? It won't help you with the score or leaderboard. But the learning and fun of finishing something that you want to finish can be equally or may'be even more important. Late submissions are perfect to try out further stuff and see the full effect of them.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "While I can understand your 'frustrations' up to a certain point...some of the things you mention will/can occur in any competition you enter :-) See them as issues/frustrations or something you deal with and then solve them. That usually works best ;-)\n\nSo the competition ends today ... but why would you not try to finish your idea? It won't help you with the score or leaderboard. But the learning and fun of finishing something that you want to finish can be equally or may'be even more important. Late submissions are perfect to try out further stuff and see the full effect of them.\n\nGood luck!",
      "votes": 1
    },
    {
      "id": 1060324,
      "postDate": "2020-10-26T03:53:47.970Z",
      "content": "<p>I hear you and agree! for such competitions new high scoring kernels should not be published at least 2-3 weeks ahead ….</p>\n<p>however, i am happy that all the current Public leaderboard medal contenders are those who have done some effort, not just Forked the high scoring notebook.  TF to Pytorch and lack of GPU computing resources is a pain, getting the pipeline and understanding new approach and training time consuming… <br>\nI am out of GPU hours till the Saturday weekly refresh and fell why bother…. </p>",
      "rawMarkdown": "I hear you and agree! for such competitions new high scoring kernels should not be published at least 2-3 weeks ahead ....\n\nhowever, i am happy that all the current Public leaderboard medal contenders are those who have done some effort, not just Forked the high scoring notebook.  TF to Pytorch and lack of GPU computing resources is a pain, getting the pipeline and understanding new approach and training time consuming... \nI am out of GPU hours till the Saturday weekly refresh and fell why bother.... ",
      "votes": 1
    },
    {
      "id": 1060034,
      "postDate": "2020-10-25T17:46:52.883Z",
      "content": "<p>Welcome to data science and Kaggle competitions. Sometimes frustation is part of the game. I find that many of my own mistakes come from a lack of focus when coding very simple things, which could have a base of tiredness, long hours or repeating many things over and over. Some other times, frustration comes from a lack of hardware that add complexity in the form of data engineering. But the positive part is that all of this is just part of the game, and also a huge source of learning. Better to look at it that way…</p>",
      "rawMarkdown": "Welcome to data science and Kaggle competitions. Sometimes frustation is part of the game. I find that many of my own mistakes come from a lack of focus when coding very simple things, which could have a base of tiredness, long hours or repeating many things over and over. Some other times, frustration comes from a lack of hardware that add complexity in the form of data engineering. But the positive part is that all of this is just part of the game, and also a huge source of learning. Better to look at it that way...",
      "votes": 1
    },
    {
      "id": 1060151,
      "postDate": "2020-10-25T20:47:39.120Z",
      "content": "<p>Bro, I FEEEL you.</p>",
      "rawMarkdown": "Bro, I FEEEL you.",
      "votes": 2,
      "replies": [
        {
          "id": 1060178,
          "postDate": "2020-10-25T22:07:23.693Z",
          "content": "<p>Remember… your program never does what you think it does</p>\n<p>In most computer science, it's relatively obvious when your program is wrong. In deep learning, all you know is is doesn't work - or it does work and you don't know why. Or it might be working correctly, just not \"learning\" anything.</p>\n<p>A lesson we have to keep learning in every competition.</p>",
          "rawMarkdown": "Remember... your program never does what you think it does\n\nIn most computer science, it's relatively obvious when your program is wrong. In deep learning, all you know is is doesn't work - or it does work and you don't know why. Or it might be working correctly, just not \"learning\" anything.\n\nA lesson we have to keep learning in every competition.",
          "votes": 4
        },
        {
          "id": 1060180,
          "postDate": "2020-10-25T22:17:46.543Z",
          "content": "<blockquote>\n  <p>Or it might be working correctly, just not \"learning\" anything.</p>\n</blockquote>\n<p>I wasted two days of precious time, <em>2 days</em> trying to figure out why is my model not learning anything. It turns out you have to set training=True.</p>",
          "rawMarkdown": "> Or it might be working correctly, just not \"learning\" anything.\n\nI wasted two days of precious time, *2 days* trying to figure out why is my model not learning anything. It turns out you have to set training=True.",
          "votes": 2,
          "replies": [
            {
              "id": 1060184,
              "postDate": "2020-10-25T22:23:27.417Z",
              "content": "<p>I spent days creating TFRecords until I realized that the slow part wasn't pydicom reading the image, resize, cropping, window-leveling or encoding the jpeg.</p>\n<p>No, the slow part was the way I was looking up the metadata in a Panda DataFrame.</p>\n<p>Final (multithreaded) version can create TFRecords for the entire training set in about 8 hours (local resources).</p>\n<p>Live and learn…</p>",
              "rawMarkdown": "I spent days creating TFRecords until I realized that the slow part wasn't pydicom reading the image, resize, cropping, window-leveling or encoding the jpeg.\n\nNo, the slow part was the way I was looking up the metadata in a Panda DataFrame.\n\nFinal (multithreaded) version can create TFRecords for the entire training set in about 8 hours (local resources).\n\nLive and learn..."
            },
            {
              "id": 1060200,
              "postDate": "2020-10-25T23:17:34.670Z",
              "content": "<p>Oh yeah, that. Luckily for me I noticed it quite early. But oh boy, the retrieval went from 2 min to a few seconds</p>",
              "rawMarkdown": "Oh yeah, that. Luckily for me I noticed it quite early. But oh boy, the retrieval went from 2 min to a few seconds"
            },
            {
              "id": 1060204,
              "postDate": "2020-10-25T23:30:03.470Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1060044,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2020-10-25T17:59:56",
      "content": "<p>Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow, or beating the baseline easily 😂<br>\nThere is a reason for most of Kaggle's opinions, like converting pytorch to tensorflow, or the baseline being difficult to train. </p>\n<blockquote>\n  <p>I commit one, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. </p>\n</blockquote>\n<p>Try random seed?</p>\n<p>Better luck next competition though, it is outstandingly hard to do a competition in a week, and I hope to see you in the next competitions. I echo the sentiment of the other commenter - Kaggle is often difficult and frustrating, but that's what makes it so rewarding when you do well.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1060051,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-25T18:07:19.760000",
          "content": "<blockquote>\n  <p>Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow</p>\n</blockquote>\n<p>I did convert the model. it just doesn't work without GPU. and It doesn't work in colab. And the input configuration is (C, H, W) as in torch instead of (H, W, C) as in TF.</p>\n<blockquote>\n  <p>Better luck next competition though, it is outstandingly hard to do a competition in a week</p>\n</blockquote>\n<p>Is that a challenge?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1060075,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2020-10-25T18:41:03.277000",
          "content": "<p>thats the reason why everyone is shifting towards pytorch  these days .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1060091,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-25T18:55:07.060000",
          "content": "<p>wdym? is it because you can't convert the model from torch to tensorflow? if anything that should be a reason for people to switch to torch. from what I understood torch is primarily for research whereas TF is for both research and production. </p>",
          "votes": -3,
          "replies": []
        }
      ]
    },
    {
      "id": 1061036,
      "author_name": "Robin Smits",
      "author_url": "",
      "post_date": "2020-10-26T17:50:31.680000",
      "content": "<p>While I can understand your 'frustrations' up to a certain point…some of the things you mention will/can occur in any competition you enter :-) See them as issues/frustrations or something you deal with and then solve them. That usually works best ;-)</p>\n<p>So the competition ends today … but why would you not try to finish your idea? It won't help you with the score or leaderboard. But the learning and fun of finishing something that you want to finish can be equally or may'be even more important. Late submissions are perfect to try out further stuff and see the full effect of them.</p>\n<p>Good luck!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1060324,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2020-10-26T03:53:47.970000",
      "content": "<p>I hear you and agree! for such competitions new high scoring kernels should not be published at least 2-3 weeks ahead ….</p>\n<p>however, i am happy that all the current Public leaderboard medal contenders are those who have done some effort, not just Forked the high scoring notebook.  TF to Pytorch and lack of GPU computing resources is a pain, getting the pipeline and understanding new approach and training time consuming… <br>\nI am out of GPU hours till the Saturday weekly refresh and fell why bother…. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1060034,
      "author_name": "Alberto Benayas",
      "author_url": "",
      "post_date": "2020-10-25T17:46:52.883000",
      "content": "<p>Welcome to data science and Kaggle competitions. Sometimes frustation is part of the game. I find that many of my own mistakes come from a lack of focus when coding very simple things, which could have a base of tiredness, long hours or repeating many things over and over. Some other times, frustration comes from a lack of hardware that add complexity in the form of data engineering. But the positive part is that all of this is just part of the game, and also a huge source of learning. Better to look at it that way…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1060151,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-10-25T20:47:39.120000",
      "content": "<p>Bro, I FEEEL you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1060178,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-10-25T22:07:23.693000",
          "content": "<p>Remember… your program never does what you think it does</p>\n<p>In most computer science, it's relatively obvious when your program is wrong. In deep learning, all you know is is doesn't work - or it does work and you don't know why. Or it might be working correctly, just not \"learning\" anything.</p>\n<p>A lesson we have to keep learning in every competition.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1060180,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-25T22:17:46.543000",
          "content": "<blockquote>\n  <p>Or it might be working correctly, just not \"learning\" anything.</p>\n</blockquote>\n<p>I wasted two days of precious time, <em>2 days</em> trying to figure out why is my model not learning anything. It turns out you have to set training=True.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1060184,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-25T22:23:27.417000",
              "content": "<p>I spent days creating TFRecords until I realized that the slow part wasn't pydicom reading the image, resize, cropping, window-leveling or encoding the jpeg.</p>\n<p>No, the slow part was the way I was looking up the metadata in a Panda DataFrame.</p>\n<p>Final (multithreaded) version can create TFRecords for the entire training set in about 8 hours (local resources).</p>\n<p>Live and learn…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1060200,
              "author_name": "DarkCube",
              "author_url": "",
              "post_date": "2020-10-25T23:17:34.670000",
              "content": "<p>Oh yeah, that. Luckily for me I noticed it quite early. But oh boy, the retrieval went from 2 min to a few seconds</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1060204,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-10-25T23:30:03.470000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1060023": "I commit once, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. \n\nI run a code cell in colab, I go do some stuff. I come back. the code cell is requesting my input. There is no way for me to enter anything. I have to terminate the session and repeat a 45 min download.\n\nI load a model in a notebook and everything is fine. I load in another notebook and a warning shows up. I load from Colab and a warning and an error show up. And ***this***, this really messed shit up. I knew I had no chance of training the feature extracted from scratch. I was counting on @khyeh0719 pretrained models. (I was going to train a better model with a mix of softmax and sigmoid activations and undersampling the negative exams but it was better than nothing). but there is a catch. I use TensorFlow. I had to convert them back from torch which wasted an entire day. now that model doesn't work. So there is that.\n\nwas so excited to participate in this competition. I had this idea of using an attention model that has been used in machine translation (which I found was nowhere discussed) to predict exam and image level labels and I was convinced that it would be a great solution, especially after the \"baseline\" that used a really simple GRU coming straight up from last year's solution. But it seems I won't be able to submit in time.\n\nThere lies my worthless idea. It may help someone. Although I really doubt that.\n\nI have only skimmed through the rules so I'm really sorry if this kind of posts isn't allowed. I just had to post my frustration somewhere. Just tell me and I'll delete it.",
    "1060044": "Nice to see reality catch up to your previous forum posts about converting from pytorch to tensorflow, or beating the baseline easily 😂\nThere is a reason for most of Kaggle's opinions, like converting pytorch to tensorflow, or the baseline being difficult to train. \n\n> I commit one, and I get about 3min per step. I commit twice, the same f-ing code, and the loss shots from 3 to 6. I commit thrice and I get to 1-2 min per step. \n\nTry random seed?\n\nBetter luck next competition though, it is outstandingly hard to do a competition in a week, and I hope to see you in the next competitions. I echo the sentiment of the other commenter - Kaggle is often difficult and frustrating, but that's what makes it so rewarding when you do well.",
    "1061036": "While I can understand your 'frustrations' up to a certain point...some of the things you mention will/can occur in any competition you enter :-) See them as issues/frustrations or something you deal with and then solve them. That usually works best ;-)\n\nSo the competition ends today ... but why would you not try to finish your idea? It won't help you with the score or leaderboard. But the learning and fun of finishing something that you want to finish can be equally or may'be even more important. Late submissions are perfect to try out further stuff and see the full effect of them.\n\nGood luck!",
    "1060324": "I hear you and agree! for such competitions new high scoring kernels should not be published at least 2-3 weeks ahead ....\n\nhowever, i am happy that all the current Public leaderboard medal contenders are those who have done some effort, not just Forked the high scoring notebook.  TF to Pytorch and lack of GPU computing resources is a pain, getting the pipeline and understanding new approach and training time consuming... \nI am out of GPU hours till the Saturday weekly refresh and fell why bother.... ",
    "1060034": "Welcome to data science and Kaggle competitions. Sometimes frustation is part of the game. I find that many of my own mistakes come from a lack of focus when coding very simple things, which could have a base of tiredness, long hours or repeating many things over and over. Some other times, frustration comes from a lack of hardware that add complexity in the form of data engineering. But the positive part is that all of this is just part of the game, and also a huge source of learning. Better to look at it that way...",
    "1060151": "Bro, I FEEEL you."
  }
}