{
  "id": 430611,
  "title": "2.5D Method Seems to Overfit",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/430611",
  "author_name": "Awsaf",
  "post_date": "2023-08-10T13:02:02.896000",
  "votes": 17,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I tried the 2.5D method, where I took 4 slices from each <code>series</code>. This approach improved the validation accuracy from <code>0.73+</code> to <code>0.93+</code>, and the validation loss from <code>3.4+</code> to <code>1.7+</code>. However, it performed poorly on the leaderboard with a score of <code>1.12</code>, which is worse than the 2D method (loss = <code>0.81</code>).</p>\n<p>I attempted to identify errors but couldn't find any. I also experimented with different models, numbers of slices, augmentations, and more, yet I consistently obtained the same overfitted result. I even considered the possibility of data leakage since the 2.5D method improved the score by <code>20%</code>. To mitigate this, I reduced the training epoch, which should have minimized leakage and improved the score. Paradoxically, it actually lowered the score.</p>\n<p>For the slices, I divided the series into 6 segments and selected the middle 4 slices for all cases. This leads me to wonder whether the test distribution differs significantly from the training distribution.</p>\n<p>Based on the leaderboard, it appears that the top scores are still achieved through the <code>weighted mean</code> submission. This suggests that the <code>3D</code> and <code>2.5D</code> methods have not proven successful for many participants. If anyone has encountered similar results, I kindly ask for confirmation.</p>",
  "messages": [
    {
      "id": 2383578,
      "postDate": "2023-08-10T13:02:02.897Z",
      "content": "<p>I tried the 2.5D method, where I took 4 slices from each <code>series</code>. This approach improved the validation accuracy from <code>0.73+</code> to <code>0.93+</code>, and the validation loss from <code>3.4+</code> to <code>1.7+</code>. However, it performed poorly on the leaderboard with a score of <code>1.12</code>, which is worse than the 2D method (loss = <code>0.81</code>).</p>\n<p>I attempted to identify errors but couldn't find any. I also experimented with different models, numbers of slices, augmentations, and more, yet I consistently obtained the same overfitted result. I even considered the possibility of data leakage since the 2.5D method improved the score by <code>20%</code>. To mitigate this, I reduced the training epoch, which should have minimized leakage and improved the score. Paradoxically, it actually lowered the score.</p>\n<p>For the slices, I divided the series into 6 segments and selected the middle 4 slices for all cases. This leads me to wonder whether the test distribution differs significantly from the training distribution.</p>\n<p>Based on the leaderboard, it appears that the top scores are still achieved through the <code>weighted mean</code> submission. This suggests that the <code>3D</code> and <code>2.5D</code> methods have not proven successful for many participants. If anyone has encountered similar results, I kindly ask for confirmation.</p>",
      "rawMarkdown": "I tried the 2.5D method, where I took 4 slices from each `series`. This approach improved the validation accuracy from `0.73+` to `0.93+`, and the validation loss from `3.4+` to `1.7+`. However, it performed poorly on the leaderboard with a score of `1.12`, which is worse than the 2D method (loss = `0.81`).\n\nI attempted to identify errors but couldn't find any. I also experimented with different models, numbers of slices, augmentations, and more, yet I consistently obtained the same overfitted result. I even considered the possibility of data leakage since the 2.5D method improved the score by `20%`. To mitigate this, I reduced the training epoch, which should have minimized leakage and improved the score. Paradoxically, it actually lowered the score.\n\nFor the slices, I divided the series into 6 segments and selected the middle 4 slices for all cases. This leads me to wonder whether the test distribution differs significantly from the training distribution.\n\nBased on the leaderboard, it appears that the top scores are still achieved through the `weighted mean` submission. This suggests that the `3D` and `2.5D` methods have not proven successful for many participants. If anyone has encountered similar results, I kindly ask for confirmation.",
      "votes": 16
    },
    {
      "id": 2390239,
      "postDate": "2023-08-14T13:03:56.557Z",
      "content": "<p>RSNA competitions are usually quite hard to approach. 3D data is tougher than 2D to tackle, but much more interesting.<br>\nAlso, labels may be weak and the log loss heavily penalizes mistakes. <br>\nSometimes a simple choice made when building the pipeline can be the difference between a model that will beat the baselines or not.</p>\n<p>Great to see that <a href=\"https://www.kaggle.com/fengqilong\" target=\"_blank\">@fengqilong</a> could make it work though. </p>",
      "rawMarkdown": "RSNA competitions are usually quite hard to approach. 3D data is tougher than 2D to tackle, but much more interesting.\nAlso, labels may be weak and the log loss heavily penalizes mistakes. \nSometimes a simple choice made when building the pipeline can be the difference between a model that will beat the baselines or not.\n\nGreat to see that @fengqilong could make it work though. ",
      "votes": 6,
      "replies": [
        {
          "id": 2390247,
          "postDate": "2023-08-14T13:12:28.743Z",
          "content": "<p>Weird Fact!!<br>\nI had a bug in inference code; when I fixed it lb dropped for <strong>2D</strong>  (<code>0.81</code> -&gt; <code>1.03</code>) but for <strong>2.5D</strong> it improved from (<code>1.21</code> -&gt; <code>0.84</code>) xD</p>",
          "rawMarkdown": "Weird Fact!!\nI had a bug in inference code; when I fixed it lb dropped for **2D**  (`0.81` -> `1.03`) but for **2.5D** it improved from (`1.21` -> `0.84`) xD",
          "votes": 3
        }
      ]
    },
    {
      "id": 2391267,
      "postDate": "2023-08-15T04:56:10.630Z",
      "content": "<p>In case anyone is wondering, here are my 2.5D pipeline notebooks,</p>\n<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-train\" target=\"_blank\">RSNA-ATD: 2.5D Series Image [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-infer\" target=\"_blank\">RSNA-ATD: 2.5D Series Image [Infer]</a></li>\n</ul>",
      "rawMarkdown": "In case anyone is wondering, here are my 2.5D pipeline notebooks,\n\n* Train: [RSNA-ATD: 2.5D Series Image [Train]](https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-train)\n* Infer: [RSNA-ATD: 2.5D Series Image [Infer]](https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-infer)",
      "votes": 3
    },
    {
      "id": 2389357,
      "postDate": "2023-08-14T02:16:51.267Z",
      "content": "<p>2.5D approach works for me and the cv result correlates well with public lb, I used the weighted mean method as a post processing method, which improves both cv and lb.</p>",
      "rawMarkdown": "2.5D approach works for me and the cv result correlates well with public lb, I used the weighted mean method as a post processing method, which improves both cv and lb.",
      "votes": 2,
      "replies": [
        {
          "id": 2389391,
          "postDate": "2023-08-14T03:10:35.620Z",
          "content": "<p>How do you select scans?</p>",
          "rawMarkdown": "How do you select scans?",
          "replies": [
            {
              "id": 2389408,
              "postDate": "2023-08-14T03:32:48Z",
              "content": "<p>similar approach to u, n scans with equal distance apart from one another</p>",
              "rawMarkdown": "similar approach to u, n scans with equal distance apart from one another",
              "votes": 1
            },
            {
              "id": 2403009,
              "postDate": "2023-08-22T13:06:06.223Z",
              "content": "<p>whats the difference between 3D and 2.5D? Selecting n scans with equal distance sounds like 3D approach.<br>\nDoes the difference lay in type of convolutions, i.e. 2D convolutions for 2.5D approach and 3D convolutions for 3D approach?</p>",
              "rawMarkdown": "whats the difference between 3D and 2.5D? Selecting n scans with equal distance sounds like 3D approach.\nDoes the difference lay in type of convolutions, i.e. 2D convolutions for 2.5D approach and 3D convolutions for 3D approach?"
            },
            {
              "id": 2405324,
              "postDate": "2023-08-23T20:21:22.723Z",
              "content": "<p>Using one slice/image is 2D; stacking consecutive slices (e.g. n slices together) is 2.5D; an entire volume of the image is 3D. 3D models converged generally converge must faster but obviously have much greater demand for compute and memory compared to 2.5D or 2D models.</p>",
              "rawMarkdown": "Using one slice/image is 2D; stacking consecutive slices (e.g. n slices together) is 2.5D; an entire volume of the image is 3D. 3D models converged generally converge must faster but obviously have much greater demand for compute and memory compared to 2.5D or 2D models."
            },
            {
              "id": 2405342,
              "postDate": "2023-08-23T20:27:52.753Z",
              "content": "<p><a href=\"https://www.kaggle.com/robotmf\" target=\"_blank\">@robotmf</a> you might be wrong with 2.5D. I did a literature review and I'm 99% sure that 2.5D is not n, but exactly 3 (to mimic RGB but from 3 different grayscale scans)</p>",
              "rawMarkdown": "@robotmf you might be wrong with 2.5D. I did a literature review and I'm 99% sure that 2.5D is not n, but exactly 3 (to mimic RGB but from 3 different grayscale scans)"
            },
            {
              "id": 2406135,
              "postDate": "2023-08-24T08:36:47.907Z",
              "content": "<p><a href=\"https://www.kaggle.com/janglinko2\" target=\"_blank\">@janglinko2</a> To simplify it a bit, 2.5D is used to refer using 2D models to handle 3D data. E.g.:</p>\n<ul>\n<li>Using a 3D CNN on scans is 3D</li>\n<li>Using a 2D CNN on all the slices of the scan, and aggregating the results along the z axis afterwards is 2.5D</li>\n</ul>",
              "rawMarkdown": "@janglinko2 To simplify it a bit, 2.5D is used to refer using 2D models to handle 3D data. E.g.:\n- Using a 3D CNN on scans is 3D\n- Using a 2D CNN on all the slices of the scan, and aggregating the results along the z axis afterwards is 2.5D",
              "votes": 1
            },
            {
              "id": 2406170,
              "postDate": "2023-08-24T08:53:53.777Z",
              "content": "<p>got it, thx</p>",
              "rawMarkdown": "got it, thx"
            }
          ]
        }
      ]
    },
    {
      "id": 2388920,
      "postDate": "2023-08-13T17:20:19.080Z",
      "content": "<p>I haven't tried the 2.5D approach but I was resizing the entire series sections to a 128x128x128 3D block and passing it to a 3D CNN. Unfortunately, there seems to only give an LB of between 0.96 to 1.78 on various runs and hyperparameters. The f1 on validation set especially fails for bowel and extravasation injury with 0 for threshold of 0.5 with other remaining injuries having about 0.7 to 0.8. But this doesn't seem to translate well to LB scores so overfitting is occurring during training.</p>",
      "rawMarkdown": "I haven't tried the 2.5D approach but I was resizing the entire series sections to a 128x128x128 3D block and passing it to a 3D CNN. Unfortunately, there seems to only give an LB of between 0.96 to 1.78 on various runs and hyperparameters. The f1 on validation set especially fails for bowel and extravasation injury with 0 for threshold of 0.5 with other remaining injuries having about 0.7 to 0.8. But this doesn't seem to translate well to LB scores so overfitting is occurring during training.",
      "replies": [
        {
          "id": 2392767,
          "postDate": "2023-08-15T21:15:15.893Z",
          "content": "<p>I was also thinking of using 3D scans too but then realized the test data is only slices. So I thought we can't use 3D scans. Is that right or am I missing something? </p>",
          "rawMarkdown": "I was also thinking of using 3D scans too but then realized the test data is only slices. So I thought we can't use 3D scans. Is that right or am I missing something? ",
          "replies": [
            {
              "id": 2400665,
              "postDate": "2023-08-21T07:23:38.237Z",
              "content": "<p>The data in test_images is only a placeholder. It will be replaced with ~1100 patients during LB evaluation.</p>",
              "rawMarkdown": "The data in test_images is only a placeholder. It will be replaced with ~1100 patients during LB evaluation.",
              "votes": 2
            },
            {
              "id": 2401170,
              "postDate": "2023-08-21T13:14:33.383Z",
              "content": "<p>Ah true.. although I think 2D input would still be a better start since it adds generalizability to the model since it doesn't necessitate the input to be 3D. Not sure if 3D model will have better performance though!</p>",
              "rawMarkdown": "Ah true.. although I think 2D input would still be a better start since it adds generalizability to the model since it doesn't necessitate the input to be 3D. Not sure if 3D model will have better performance though!"
            }
          ]
        }
      ]
    },
    {
      "id": 2388690,
      "postDate": "2023-08-13T15:20:36.507Z",
      "content": "<p>question, for training your model you used part of the training image and making a pipeline and then applying to whole dataset or your method is something else? </p>",
      "rawMarkdown": "question, for training your model you used part of the training image and making a pipeline and then applying to whole dataset or your method is something else? "
    },
    {
      "id": 2383598,
      "postDate": "2023-08-10T13:11:53.480Z",
      "content": "<p>I think labels are too weak for 2D or 2.5D models. How exactly are you validating your 2.5D models? Are predictions aggregated on scan and patient levels or are you calculating score for each slice?</p>",
      "rawMarkdown": "I think labels are too weak for 2D or 2.5D models. How exactly are you validating your 2.5D models? Are predictions aggregated on scan and patient levels or are you calculating score for each slice?",
      "replies": [
        {
          "id": 2383613,
          "postDate": "2023-08-10T13:20:23.230Z",
          "content": "<p>Labels for 2.5D should be stronger than 2D as we're considering each series as one input. But the result of 2.5D is worse than 2D.<br>\nI evaluated the 2.5D method <code>serieswise</code> whereas in 2D method I evaluated <code>scanwise</code> it could be the reason behind the discrepancy of scores between 2D and 2.5D but still 2.5D should have better results on <code>leaderbard</code> as its labels are stronger than 2D.</p>",
          "rawMarkdown": "Labels for 2.5D should be stronger than 2D as we're considering each series as one input. But the result of 2.5D is worse than 2D.\nI evaluated the 2.5D method `serieswise` whereas in 2D method I evaluated `scanwise` it could be the reason behind the discrepancy of scores between 2D and 2.5D but still 2.5D should have better results on `leaderbard` as its labels are stronger than 2D.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2390239,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-08-14T13:03:56.557000",
      "content": "<p>RSNA competitions are usually quite hard to approach. 3D data is tougher than 2D to tackle, but much more interesting.<br>\nAlso, labels may be weak and the log loss heavily penalizes mistakes. <br>\nSometimes a simple choice made when building the pipeline can be the difference between a model that will beat the baselines or not.</p>\n<p>Great to see that <a href=\"https://www.kaggle.com/fengqilong\" target=\"_blank\">@fengqilong</a> could make it work though. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 2390247,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2023-08-14T13:12:28.743000",
          "content": "<p>Weird Fact!!<br>\nI had a bug in inference code; when I fixed it lb dropped for <strong>2D</strong>  (<code>0.81</code> -&gt; <code>1.03</code>) but for <strong>2.5D</strong> it improved from (<code>1.21</code> -&gt; <code>0.84</code>) xD</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2391267,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2023-08-15T04:56:10.630000",
      "content": "<p>In case anyone is wondering, here are my 2.5D pipeline notebooks,</p>\n<ul>\n<li>Train: <a href=\"https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-train\" target=\"_blank\">RSNA-ATD: 2.5D Series Image [Train]</a></li>\n<li>Infer: <a href=\"https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-infer\" target=\"_blank\">RSNA-ATD: 2.5D Series Image [Infer]</a></li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2389357,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-08-14T02:16:51.267000",
      "content": "<p>2.5D approach works for me and the cv result correlates well with public lb, I used the weighted mean method as a post processing method, which improves both cv and lb.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2389391,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2023-08-14T03:10:35.620000",
          "content": "<p>How do you select scans?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2389408,
              "author_name": "Feng Qilong",
              "author_url": "",
              "post_date": "2023-08-14T03:32:48",
              "content": "<p>similar approach to u, n scans with equal distance apart from one another</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2403009,
              "author_name": "JanGlinko2",
              "author_url": "",
              "post_date": "2023-08-22T13:06:06.223000",
              "content": "<p>whats the difference between 3D and 2.5D? Selecting n scans with equal distance sounds like 3D approach.<br>\nDoes the difference lay in type of convolutions, i.e. 2D convolutions for 2.5D approach and 3D convolutions for 3D approach?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2405324,
              "author_name": "Mark",
              "author_url": "",
              "post_date": "2023-08-23T20:21:22.723000",
              "content": "<p>Using one slice/image is 2D; stacking consecutive slices (e.g. n slices together) is 2.5D; an entire volume of the image is 3D. 3D models converged generally converge must faster but obviously have much greater demand for compute and memory compared to 2.5D or 2D models.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2405342,
              "author_name": "JanGlinko2",
              "author_url": "",
              "post_date": "2023-08-23T20:27:52.753000",
              "content": "<p><a href=\"https://www.kaggle.com/robotmf\" target=\"_blank\">@robotmf</a> you might be wrong with 2.5D. I did a literature review and I'm 99% sure that 2.5D is not n, but exactly 3 (to mimic RGB but from 3 different grayscale scans)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2406135,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2023-08-24T08:36:47.907000",
              "content": "<p><a href=\"https://www.kaggle.com/janglinko2\" target=\"_blank\">@janglinko2</a> To simplify it a bit, 2.5D is used to refer using 2D models to handle 3D data. E.g.:</p>\n<ul>\n<li>Using a 3D CNN on scans is 3D</li>\n<li>Using a 2D CNN on all the slices of the scan, and aggregating the results along the z axis afterwards is 2.5D</li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2406170,
              "author_name": "JanGlinko2",
              "author_url": "",
              "post_date": "2023-08-24T08:53:53.777000",
              "content": "<p>got it, thx</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2388920,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2023-08-13T17:20:19.080000",
      "content": "<p>I haven't tried the 2.5D approach but I was resizing the entire series sections to a 128x128x128 3D block and passing it to a 3D CNN. Unfortunately, there seems to only give an LB of between 0.96 to 1.78 on various runs and hyperparameters. The f1 on validation set especially fails for bowel and extravasation injury with 0 for threshold of 0.5 with other remaining injuries having about 0.7 to 0.8. But this doesn't seem to translate well to LB scores so overfitting is occurring during training.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2392767,
          "author_name": "Parham Mostame",
          "author_url": "",
          "post_date": "2023-08-15T21:15:15.893000",
          "content": "<p>I was also thinking of using 3D scans too but then realized the test data is only slices. So I thought we can't use 3D scans. Is that right or am I missing something? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2400665,
              "author_name": "zacstewart",
              "author_url": "",
              "post_date": "2023-08-21T07:23:38.237000",
              "content": "<p>The data in test_images is only a placeholder. It will be replaced with ~1100 patients during LB evaluation.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2401170,
              "author_name": "Parham Mostame",
              "author_url": "",
              "post_date": "2023-08-21T13:14:33.383000",
              "content": "<p>Ah true.. although I think 2D input would still be a better start since it adds generalizability to the model since it doesn't necessitate the input to be 3D. Not sure if 3D model will have better performance though!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2388690,
      "author_name": "golnaz ahmadvand",
      "author_url": "",
      "post_date": "2023-08-13T15:20:36.507000",
      "content": "<p>question, for training your model you used part of the training image and making a pipeline and then applying to whole dataset or your method is something else? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2383598,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2023-08-10T13:11:53.480000",
      "content": "<p>I think labels are too weak for 2D or 2.5D models. How exactly are you validating your 2.5D models? Are predictions aggregated on scan and patient levels or are you calculating score for each slice?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2383613,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2023-08-10T13:20:23.230000",
          "content": "<p>Labels for 2.5D should be stronger than 2D as we're considering each series as one input. But the result of 2.5D is worse than 2D.<br>\nI evaluated the 2.5D method <code>serieswise</code> whereas in 2D method I evaluated <code>scanwise</code> it could be the reason behind the discrepancy of scores between 2D and 2.5D but still 2.5D should have better results on <code>leaderbard</code> as its labels are stronger than 2D.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2383578": "I tried the 2.5D method, where I took 4 slices from each `series`. This approach improved the validation accuracy from `0.73+` to `0.93+`, and the validation loss from `3.4+` to `1.7+`. However, it performed poorly on the leaderboard with a score of `1.12`, which is worse than the 2D method (loss = `0.81`).\n\nI attempted to identify errors but couldn't find any. I also experimented with different models, numbers of slices, augmentations, and more, yet I consistently obtained the same overfitted result. I even considered the possibility of data leakage since the 2.5D method improved the score by `20%`. To mitigate this, I reduced the training epoch, which should have minimized leakage and improved the score. Paradoxically, it actually lowered the score.\n\nFor the slices, I divided the series into 6 segments and selected the middle 4 slices for all cases. This leads me to wonder whether the test distribution differs significantly from the training distribution.\n\nBased on the leaderboard, it appears that the top scores are still achieved through the `weighted mean` submission. This suggests that the `3D` and `2.5D` methods have not proven successful for many participants. If anyone has encountered similar results, I kindly ask for confirmation.",
    "2390239": "RSNA competitions are usually quite hard to approach. 3D data is tougher than 2D to tackle, but much more interesting.\nAlso, labels may be weak and the log loss heavily penalizes mistakes. \nSometimes a simple choice made when building the pipeline can be the difference between a model that will beat the baselines or not.\n\nGreat to see that @fengqilong could make it work though. ",
    "2391267": "In case anyone is wondering, here are my 2.5D pipeline notebooks,\n\n* Train: [RSNA-ATD: 2.5D Series Image [Train]](https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-train)\n* Infer: [RSNA-ATD: 2.5D Series Image [Infer]](https://www.kaggle.com/awsaf49/rsna-atd-2-5d-series-image-infer)",
    "2389357": "2.5D approach works for me and the cv result correlates well with public lb, I used the weighted mean method as a post processing method, which improves both cv and lb.",
    "2388920": "I haven't tried the 2.5D approach but I was resizing the entire series sections to a 128x128x128 3D block and passing it to a 3D CNN. Unfortunately, there seems to only give an LB of between 0.96 to 1.78 on various runs and hyperparameters. The f1 on validation set especially fails for bowel and extravasation injury with 0 for threshold of 0.5 with other remaining injuries having about 0.7 to 0.8. But this doesn't seem to translate well to LB scores so overfitting is occurring during training.",
    "2388690": "question, for training your model you used part of the training image and making a pipeline and then applying to whole dataset or your method is something else? ",
    "2383598": "I think labels are too weak for 2D or 2.5D models. How exactly are you validating your 2.5D models? Are predictions aggregated on scan and patient levels or are you calculating score for each slice?"
  }
}