{
  "id": 447651,
  "title": "32nd place overview + mistakes",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447651",
  "author_name": "Optimo",
  "post_date": "2023-10-16T16:52:20.352000",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I know that 32nd place is not necessarily the dream place people will aim for in future competitions but it's always important to learn from mistakes and we've made a few so I'll share our solution here and share to my future self and others how to improve for the next time.</p>\n<h2>Overall solution</h2>\n<p><strong>First stage: organ segmentation</strong><br>\nIn order to go from weak labels to strong labels it's important to be able to segment the organs of interest, so that's what we did like many other teams.</p>\n<p>One potential mistake we did here is to keep the original labels of the annotated data and treat them as out-of-fold predictions for the next stages. Although it might seem a harmless decision, the fact that the annotated data all had injuries and that ground truth segmentation could be distinguishable by second stage models might have introduce a small data leakage.</p>\n<p><strong>Second stage: organ injury classification</strong><br>\nLike many others we used our segmentation mask to crop region of interests and feed them to a 2.5D + LSTM models (one for each of the segmented organs: bowel, liver, spleen and kidney). But we decided to also give the segmentation mask to the model as input, this is where we did something different which probably was not necessary and might have introduce a leak in our pipeline (not the only problem though but a point of concern).</p>\n<p>One thing we've done differently from most solution I've read so far is to train one single model for all targets but in the axis parallel to the z-axis: the CT scan is fully stacked to form a 3D image Z, W, H and then slices of data are generated by slicing across W and H -&gt; this allows to see all the organs within only few frames (16 or 24), to train end to end with all organs, to augment training by switching the slicing between W and H and to ensemble a very different approach.</p>\n<p>On both approaches we noticed that our models were bad for extravasation, so we simply ignored it and never had time to come back to it before end of the competition. Note to myself, joining a competition 3 weeks before the end is short.</p>\n<p><strong>Thrid stage: MLP for ensembling and competition metric optimization</strong></p>\n<p>Since the goal of a Kaggle competition is to optimize a metric it's important to have a final approach that tries to minimize that metric. So for each patient we took the predictions of our two approaches, stacked them and stacked the predictions for two series belonging to a same patient.<br>\nNow we just train an MLP to minimize the official competition metric.<br>\nNote that the final model is predicting extravasation class without any input information about it.</p>\n<p>Individually both of our second stage approaches reached about CV 0.42 - LB 0.52, when stacked together they reached CV 0.386 - LB 0.49.</p>\n<h2>Big mistake</h2>\n<p>When entering a competition you are always eager to know what your first solution is worth on the public LB, so you go for a quick and dirty inference notebook. For basic image competition everything goes fine, for 3 stages solution things can get ugly pretty fast.</p>\n<p>You don't care too much about your crappy inference code when your LB score is improving, so you add new stuff until your score does not improve anymore and your CV - LB gap is the worst among all participants! 😭</p>\n<p>In the end I've spent the last few days (+ the 2 extra days) trying to figure out from where this huge CV - LB gap came from…</p>\n<p>Score on fold 0 during CV 0.393 -&gt; score on fold 0 out of kaggle inference notebook 0.418… Silent errors in machine learning are frequent so be careful!</p>\n<p>Anyway it was a fun competition to participate in, congratulations to everyone and see you soon!</p>",
  "messages": [
    {
      "id": 2484759,
      "postDate": "2023-10-16T16:52:20.353Z",
      "content": "<p>I know that 32nd place is not necessarily the dream place people will aim for in future competitions but it's always important to learn from mistakes and we've made a few so I'll share our solution here and share to my future self and others how to improve for the next time.</p>\n<h2>Overall solution</h2>\n<p><strong>First stage: organ segmentation</strong><br>\nIn order to go from weak labels to strong labels it's important to be able to segment the organs of interest, so that's what we did like many other teams.</p>\n<p>One potential mistake we did here is to keep the original labels of the annotated data and treat them as out-of-fold predictions for the next stages. Although it might seem a harmless decision, the fact that the annotated data all had injuries and that ground truth segmentation could be distinguishable by second stage models might have introduce a small data leakage.</p>\n<p><strong>Second stage: organ injury classification</strong><br>\nLike many others we used our segmentation mask to crop region of interests and feed them to a 2.5D + LSTM models (one for each of the segmented organs: bowel, liver, spleen and kidney). But we decided to also give the segmentation mask to the model as input, this is where we did something different which probably was not necessary and might have introduce a leak in our pipeline (not the only problem though but a point of concern).</p>\n<p>One thing we've done differently from most solution I've read so far is to train one single model for all targets but in the axis parallel to the z-axis: the CT scan is fully stacked to form a 3D image Z, W, H and then slices of data are generated by slicing across W and H -&gt; this allows to see all the organs within only few frames (16 or 24), to train end to end with all organs, to augment training by switching the slicing between W and H and to ensemble a very different approach.</p>\n<p>On both approaches we noticed that our models were bad for extravasation, so we simply ignored it and never had time to come back to it before end of the competition. Note to myself, joining a competition 3 weeks before the end is short.</p>\n<p><strong>Thrid stage: MLP for ensembling and competition metric optimization</strong></p>\n<p>Since the goal of a Kaggle competition is to optimize a metric it's important to have a final approach that tries to minimize that metric. So for each patient we took the predictions of our two approaches, stacked them and stacked the predictions for two series belonging to a same patient.<br>\nNow we just train an MLP to minimize the official competition metric.<br>\nNote that the final model is predicting extravasation class without any input information about it.</p>\n<p>Individually both of our second stage approaches reached about CV 0.42 - LB 0.52, when stacked together they reached CV 0.386 - LB 0.49.</p>\n<h2>Big mistake</h2>\n<p>When entering a competition you are always eager to know what your first solution is worth on the public LB, so you go for a quick and dirty inference notebook. For basic image competition everything goes fine, for 3 stages solution things can get ugly pretty fast.</p>\n<p>You don't care too much about your crappy inference code when your LB score is improving, so you add new stuff until your score does not improve anymore and your CV - LB gap is the worst among all participants! 😭</p>\n<p>In the end I've spent the last few days (+ the 2 extra days) trying to figure out from where this huge CV - LB gap came from…</p>\n<p>Score on fold 0 during CV 0.393 -&gt; score on fold 0 out of kaggle inference notebook 0.418… Silent errors in machine learning are frequent so be careful!</p>\n<p>Anyway it was a fun competition to participate in, congratulations to everyone and see you soon!</p>",
      "rawMarkdown": "I know that 32nd place is not necessarily the dream place people will aim for in future competitions but it's always important to learn from mistakes and we've made a few so I'll share our solution here and share to my future self and others how to improve for the next time.\n\n## Overall solution\n\n**First stage: organ segmentation**\nIn order to go from weak labels to strong labels it's important to be able to segment the organs of interest, so that's what we did like many other teams.\n\nOne potential mistake we did here is to keep the original labels of the annotated data and treat them as out-of-fold predictions for the next stages. Although it might seem a harmless decision, the fact that the annotated data all had injuries and that ground truth segmentation could be distinguishable by second stage models might have introduce a small data leakage.\n\n**Second stage: organ injury classification**\nLike many others we used our segmentation mask to crop region of interests and feed them to a 2.5D + LSTM models (one for each of the segmented organs: bowel, liver, spleen and kidney). But we decided to also give the segmentation mask to the model as input, this is where we did something different which probably was not necessary and might have introduce a leak in our pipeline (not the only problem though but a point of concern).\n\nOne thing we've done differently from most solution I've read so far is to train one single model for all targets but in the axis parallel to the z-axis: the CT scan is fully stacked to form a 3D image Z, W, H and then slices of data are generated by slicing across W and H -> this allows to see all the organs within only few frames (16 or 24), to train end to end with all organs, to augment training by switching the slicing between W and H and to ensemble a very different approach.\n\nOn both approaches we noticed that our models were bad for extravasation, so we simply ignored it and never had time to come back to it before end of the competition. Note to myself, joining a competition 3 weeks before the end is short.\n\n**Thrid stage: MLP for ensembling and competition metric optimization**\n\nSince the goal of a Kaggle competition is to optimize a metric it's important to have a final approach that tries to minimize that metric. So for each patient we took the predictions of our two approaches, stacked them and stacked the predictions for two series belonging to a same patient.\nNow we just train an MLP to minimize the official competition metric.\nNote that the final model is predicting extravasation class without any input information about it.\n\nIndividually both of our second stage approaches reached about CV 0.42 - LB 0.52, when stacked together they reached CV 0.386 - LB 0.49.\n\n##Big mistake\nWhen entering a competition you are always eager to know what your first solution is worth on the public LB, so you go for a quick and dirty inference notebook. For basic image competition everything goes fine, for 3 stages solution things can get ugly pretty fast.\n\nYou don't care too much about your crappy inference code when your LB score is improving, so you add new stuff until your score does not improve anymore and your CV - LB gap is the worst among all participants! 😭\n\nIn the end I've spent the last few days (+ the 2 extra days) trying to figure out from where this huge CV - LB gap came from...\n\nScore on fold 0 during CV 0.393 -> score on fold 0 out of kaggle inference notebook 0.418... Silent errors in machine learning are frequent so be careful!\n\n\nAnyway it was a fun competition to participate in, congratulations to everyone and see you soon!",
      "votes": 18
    },
    {
      "id": 2509460,
      "postDate": "2023-11-02T11:38:53.697Z",
      "content": "<p>Thanks for sharing your insights!</p>",
      "rawMarkdown": "Thanks for sharing your insights!"
    }
  ],
  "comments": [
    {
      "id": 2509460,
      "author_name": "Pankaj Pansari",
      "author_url": "",
      "post_date": "2023-11-02T11:38:53.697000",
      "content": "<p>Thanks for sharing your insights!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2484759": "I know that 32nd place is not necessarily the dream place people will aim for in future competitions but it's always important to learn from mistakes and we've made a few so I'll share our solution here and share to my future self and others how to improve for the next time.\n\n## Overall solution\n\n**First stage: organ segmentation**\nIn order to go from weak labels to strong labels it's important to be able to segment the organs of interest, so that's what we did like many other teams.\n\nOne potential mistake we did here is to keep the original labels of the annotated data and treat them as out-of-fold predictions for the next stages. Although it might seem a harmless decision, the fact that the annotated data all had injuries and that ground truth segmentation could be distinguishable by second stage models might have introduce a small data leakage.\n\n**Second stage: organ injury classification**\nLike many others we used our segmentation mask to crop region of interests and feed them to a 2.5D + LSTM models (one for each of the segmented organs: bowel, liver, spleen and kidney). But we decided to also give the segmentation mask to the model as input, this is where we did something different which probably was not necessary and might have introduce a leak in our pipeline (not the only problem though but a point of concern).\n\nOne thing we've done differently from most solution I've read so far is to train one single model for all targets but in the axis parallel to the z-axis: the CT scan is fully stacked to form a 3D image Z, W, H and then slices of data are generated by slicing across W and H -> this allows to see all the organs within only few frames (16 or 24), to train end to end with all organs, to augment training by switching the slicing between W and H and to ensemble a very different approach.\n\nOn both approaches we noticed that our models were bad for extravasation, so we simply ignored it and never had time to come back to it before end of the competition. Note to myself, joining a competition 3 weeks before the end is short.\n\n**Thrid stage: MLP for ensembling and competition metric optimization**\n\nSince the goal of a Kaggle competition is to optimize a metric it's important to have a final approach that tries to minimize that metric. So for each patient we took the predictions of our two approaches, stacked them and stacked the predictions for two series belonging to a same patient.\nNow we just train an MLP to minimize the official competition metric.\nNote that the final model is predicting extravasation class without any input information about it.\n\nIndividually both of our second stage approaches reached about CV 0.42 - LB 0.52, when stacked together they reached CV 0.386 - LB 0.49.\n\n##Big mistake\nWhen entering a competition you are always eager to know what your first solution is worth on the public LB, so you go for a quick and dirty inference notebook. For basic image competition everything goes fine, for 3 stages solution things can get ugly pretty fast.\n\nYou don't care too much about your crappy inference code when your LB score is improving, so you add new stuff until your score does not improve anymore and your CV - LB gap is the worst among all participants! 😭\n\nIn the end I've spent the last few days (+ the 2 extra days) trying to figure out from where this huge CV - LB gap came from...\n\nScore on fold 0 during CV 0.393 -> score on fold 0 out of kaggle inference notebook 0.418... Silent errors in machine learning are frequent so be careful!\n\n\nAnyway it was a fun competition to participate in, congratulations to everyone and see you soon!",
    "2509460": "Thanks for sharing your insights!"
  }
}