{
  "id": 499240,
  "title": "1st Place Solution - Tips on Pre- & Post-Processing with Single Model",
  "url": "/competitions/spr-head-ct-age-prediction-challenge/discussion/499240",
  "author_name": "YYama",
  "post_date": "2024-05-01T06:28:02.428000",
  "votes": 7,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, thank you for organizing this wonderful competition.<br>\nIt was very rewarding to work with such a large-scale dataset across multiple facilities.</p>\n<p>Here’s my solution:</p>\n<h1>Overview</h1>\n<p>I used a single fold and a single model approach with tf-efficientnetv2-s, training on the age of each study and aggregating predictions from multiple slices to obtain the final prediction for each study.</p>\n<h1>Data Split</h1>\n<p>I used the holdout method. Training and validation were performed using StratifiedGroupKFold. Age was used for stratification both as a continuous and as a categorical variable, with grouping by patient. For validation data: fold == 0, training data: fold != 0.</p>\n<h1>Preprocessing</h1>\n<h3>Windowing</h3>\n<p>Initially, I stacked images in the channel direction using brain parenchyma, bone, and soft tissue windows (3ch in total).<br>\nreference: <a href=\"https://radiopaedia.org/articles/windowing-ct\" target=\"_blank\">https://radiopaedia.org/articles/windowing-ct</a></p>\n<h3>Removing Unnecessary Images</h3>\n<p>Due to likely anonymization, there was added noise around the images. To remove this, I applied rule-based processing. It was crucial to use the very high CT values of bone. In channels conditioned on bone, the maximum value should represent bone. Observations showed that 5-10% of images included only noise or soft tissues at the top of the head. For my bone window condition, the maximum pixel value of these images was 77. Consequently, I removed all images where the maximum pixel value in the bone window did not exceed 100.</p>\n<h3>No Resizing in x, y, z Directions</h3>\n<p>Initially, I resized in x, y, and z directions to obtain 2D images, but this led to decreased performance. Ultimately, I performed resizing to 512x512 size without compressing the z-direction to keep sizes close to the original.</p>\n<h1>Modeling</h1>\n<p>I only used tf-efficientnetv2-s, employing basic augmentations such as horizontal flip and cutout. The model was trained for 5 epochs using a scheduler that started with a initial lr of 0 and maximum lr of 1e-3, which was progressively reduced to 1e-6 through a warmup process.</p>\n<h1>Prediction Aggregation (Post-Processing)</h1>\n<p>Predictions were obtained for varying numbers of slices per study. Simply taking the mean and median of these predictions, the median outperformed in terms of my validation score.<br>\nFurther, I aggregated predictions considering the percentile position of each slice within a study. This approach, using the mean of predictions where slices were positioned in the tail 40th percentile, improved performance, reducing CV from about 4.0 to 3.6 (these are approximate values as I’m compiling this solution remotely). The public leaderboard score showed a slight improvement from 3.165 to 3.123, and a more notable improvement from 2.774 to 2.596 in the private leaderboard.</p>\n<h1>What Didn’t Work</h1>\n<ul>\n<li>3D Architectures: I used 3D DenseNet121. Even the best models scored around 6.7 on the public leaderboard, and predictions were around 6.5 on the private leaderboard.</li>\n</ul>\n<h1>Untried but Potentially Effective Strategies</h1>\n<ul>\n<li>Second-stage Model: Although I used the aforementioned rule-based processing for prediction aggregation, developing a model that uses predictions from the first stage as features to output final predictions could potentially improve scores further.</li>\n<li>Ensemble: Using an ensemble of multiple models or multi-fold models could also likely enhance scores.</li>\n</ul>",
  "messages": [
    {
      "id": 2786080,
      "postDate": "2024-05-01T06:28:02.427Z",
      "content": "<p>First of all, thank you for organizing this wonderful competition.<br>\nIt was very rewarding to work with such a large-scale dataset across multiple facilities.</p>\n<p>Here’s my solution:</p>\n<h1>Overview</h1>\n<p>I used a single fold and a single model approach with tf-efficientnetv2-s, training on the age of each study and aggregating predictions from multiple slices to obtain the final prediction for each study.</p>\n<h1>Data Split</h1>\n<p>I used the holdout method. Training and validation were performed using StratifiedGroupKFold. Age was used for stratification both as a continuous and as a categorical variable, with grouping by patient. For validation data: fold == 0, training data: fold != 0.</p>\n<h1>Preprocessing</h1>\n<h3>Windowing</h3>\n<p>Initially, I stacked images in the channel direction using brain parenchyma, bone, and soft tissue windows (3ch in total).<br>\nreference: <a href=\"https://radiopaedia.org/articles/windowing-ct\" target=\"_blank\">https://radiopaedia.org/articles/windowing-ct</a></p>\n<h3>Removing Unnecessary Images</h3>\n<p>Due to likely anonymization, there was added noise around the images. To remove this, I applied rule-based processing. It was crucial to use the very high CT values of bone. In channels conditioned on bone, the maximum value should represent bone. Observations showed that 5-10% of images included only noise or soft tissues at the top of the head. For my bone window condition, the maximum pixel value of these images was 77. Consequently, I removed all images where the maximum pixel value in the bone window did not exceed 100.</p>\n<h3>No Resizing in x, y, z Directions</h3>\n<p>Initially, I resized in x, y, and z directions to obtain 2D images, but this led to decreased performance. Ultimately, I performed resizing to 512x512 size without compressing the z-direction to keep sizes close to the original.</p>\n<h1>Modeling</h1>\n<p>I only used tf-efficientnetv2-s, employing basic augmentations such as horizontal flip and cutout. The model was trained for 5 epochs using a scheduler that started with a initial lr of 0 and maximum lr of 1e-3, which was progressively reduced to 1e-6 through a warmup process.</p>\n<h1>Prediction Aggregation (Post-Processing)</h1>\n<p>Predictions were obtained for varying numbers of slices per study. Simply taking the mean and median of these predictions, the median outperformed in terms of my validation score.<br>\nFurther, I aggregated predictions considering the percentile position of each slice within a study. This approach, using the mean of predictions where slices were positioned in the tail 40th percentile, improved performance, reducing CV from about 4.0 to 3.6 (these are approximate values as I’m compiling this solution remotely). The public leaderboard score showed a slight improvement from 3.165 to 3.123, and a more notable improvement from 2.774 to 2.596 in the private leaderboard.</p>\n<h1>What Didn’t Work</h1>\n<ul>\n<li>3D Architectures: I used 3D DenseNet121. Even the best models scored around 6.7 on the public leaderboard, and predictions were around 6.5 on the private leaderboard.</li>\n</ul>\n<h1>Untried but Potentially Effective Strategies</h1>\n<ul>\n<li>Second-stage Model: Although I used the aforementioned rule-based processing for prediction aggregation, developing a model that uses predictions from the first stage as features to output final predictions could potentially improve scores further.</li>\n<li>Ensemble: Using an ensemble of multiple models or multi-fold models could also likely enhance scores.</li>\n</ul>",
      "rawMarkdown": "First of all, thank you for organizing this wonderful competition.\nIt was very rewarding to work with such a large-scale dataset across multiple facilities.\n\nHere’s my solution:\n\n# Overview\nI used a single fold and a single model approach with tf-efficientnetv2-s, training on the age of each study and aggregating predictions from multiple slices to obtain the final prediction for each study.\n\n# Data Split\nI used the holdout method. Training and validation were performed using StratifiedGroupKFold. Age was used for stratification both as a continuous and as a categorical variable, with grouping by patient. For validation data: fold == 0, training data: fold != 0.\n\n# Preprocessing\n### Windowing\nInitially, I stacked images in the channel direction using brain parenchyma, bone, and soft tissue windows (3ch in total).\nreference: https://radiopaedia.org/articles/windowing-ct\n\n### Removing Unnecessary Images\nDue to likely anonymization, there was added noise around the images. To remove this, I applied rule-based processing. It was crucial to use the very high CT values of bone. In channels conditioned on bone, the maximum value should represent bone. Observations showed that 5-10% of images included only noise or soft tissues at the top of the head. For my bone window condition, the maximum pixel value of these images was 77. Consequently, I removed all images where the maximum pixel value in the bone window did not exceed 100.\n\n### No Resizing in x, y, z Directions\nInitially, I resized in x, y, and z directions to obtain 2D images, but this led to decreased performance. Ultimately, I performed resizing to 512x512 size without compressing the z-direction to keep sizes close to the original.\n\n# Modeling\nI only used tf-efficientnetv2-s, employing basic augmentations such as horizontal flip and cutout. The model was trained for 5 epochs using a scheduler that started with a initial lr of 0 and maximum lr of 1e-3, which was progressively reduced to 1e-6 through a warmup process.\n\n# Prediction Aggregation (Post-Processing)\nPredictions were obtained for varying numbers of slices per study. Simply taking the mean and median of these predictions, the median outperformed in terms of my validation score.\nFurther, I aggregated predictions considering the percentile position of each slice within a study. This approach, using the mean of predictions where slices were positioned in the tail 40th percentile, improved performance, reducing CV from about 4.0 to 3.6 (these are approximate values as I’m compiling this solution remotely). The public leaderboard score showed a slight improvement from 3.165 to 3.123, and a more notable improvement from 2.774 to 2.596 in the private leaderboard.\n\n# What Didn’t Work\n- 3D Architectures: I used 3D DenseNet121. Even the best models scored around 6.7 on the public leaderboard, and predictions were around 6.5 on the private leaderboard.\n\n# Untried but Potentially Effective Strategies\n- Second-stage Model: Although I used the aforementioned rule-based processing for prediction aggregation, developing a model that uses predictions from the first stage as features to output final predictions could potentially improve scores further.\n- Ensemble: Using an ensemble of multiple models or multi-fold models could also likely enhance scores.\n",
      "votes": 6
    },
    {
      "id": 2827358,
      "postDate": "2024-05-21T12:25:26.053Z",
      "content": "<p>Congratulations for your solution! Are you uploading the preprocessing/training script you used in your model?</p>",
      "rawMarkdown": "Congratulations for your solution! Are you uploading the preprocessing/training script you used in your model?",
      "votes": 1,
      "replies": [
        {
          "id": 2834859,
          "postDate": "2024-05-25T04:31:39.270Z",
          "content": "<p>Sorry for the delay! I've been busy with work.</p>\n<p>I have shared the training and inference ipynb file of my best models.<br>\nI will upload the preprocess files within a few days.</p>\n<p><a href=\"https://github.com/yamagishi0824/spr_head_ct_age_1st\" target=\"_blank\">https://github.com/yamagishi0824/spr_head_ct_age_1st</a></p>",
          "rawMarkdown": "Sorry for the delay! I've been busy with work.\n\nI have shared the training and inference ipynb file of my best models.\nI will upload the preprocess files within a few days.\n\nhttps://github.com/yamagishi0824/spr_head_ct_age_1st",
          "votes": 1,
          "replies": [
            {
              "id": 2838676,
              "postDate": "2024-05-27T07:01:31.163Z",
              "content": "<p>thank you! incredibly helpful</p>",
              "rawMarkdown": "thank you! incredibly helpful",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2788266,
      "postDate": "2024-05-02T06:52:23.550Z",
      "content": "<p>Congratulations! A very clear approach, thanks for sharing! May I ask what was your strategy (if there was any) for sampling which slices to select from a series?</p>",
      "rawMarkdown": "Congratulations! A very clear approach, thanks for sharing! May I ask what was your strategy (if there was any) for sampling which slices to select from a series?",
      "votes": 1,
      "replies": [
        {
          "id": 2790022,
          "postDate": "2024-05-03T01:16:59.103Z",
          "content": "<p>Thank you for your question!</p>\n<p>During the training of the model, I used all images except for removing noise images as I wrote in the solution above. Furthermore, during inference, I used all slices in the 40th percentile of the caudal-end. The slices I adopted for inference were selected through a simple process without much scrutiny, so I think there might have been a better method.</p>",
          "rawMarkdown": "Thank you for your question!\n\nDuring the training of the model, I used all images except for removing noise images as I wrote in the solution above. Furthermore, during inference, I used all slices in the 40th percentile of the caudal-end. The slices I adopted for inference were selected through a simple process without much scrutiny, so I think there might have been a better method.",
          "votes": 1,
          "replies": [
            {
              "id": 2791011,
              "postDate": "2024-05-03T12:40:04.773Z",
              "content": "<p>thank you, will you be open-sourcing the code? I'd love to explore and learn more</p>",
              "rawMarkdown": "thank you, will you be open-sourcing the code? I'd love to explore and learn more",
              "votes": 1
            },
            {
              "id": 2791017,
              "postDate": "2024-05-03T12:43:32.757Z",
              "content": "<p>Yes!<br>\nOnce I've cleaned up the codes, I plan to make them public!</p>",
              "rawMarkdown": "Yes!\nOnce I've cleaned up the codes, I plan to make them public!",
              "votes": 1
            },
            {
              "id": 2791044,
              "postDate": "2024-05-03T12:55:17.943Z",
              "content": "<p>awesome thank you</p>",
              "rawMarkdown": "awesome thank you",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2786133,
      "postDate": "2024-05-01T07:06:41.183Z",
      "content": "<p>Thanks for sharing! Did you mean you used 2d instead of 2.5D? What loss function are you using?</p>",
      "rawMarkdown": "Thanks for sharing! Did you mean you used 2d instead of 2.5D? What loss function are you using?",
      "votes": 2,
      "replies": [
        {
          "id": 2786146,
          "postDate": "2024-05-01T07:18:49.033Z",
          "content": "<p>That's right! It’s simply a normal 2D model, with windows from 2D images stacked in the channel direction. <br>\nThe same \"Age\" label is assigned to all slices within the same study. For the loss function, I used MAE loss (nn.L1Loss).</p>",
          "rawMarkdown": "That's right! It’s simply a normal 2D model, with windows from 2D images stacked in the channel direction. \nThe same \"Age\" label is assigned to all slices within the same study. For the loss function, I used MAE loss (nn.L1Loss).",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2827358,
      "author_name": "Artur Paulo",
      "author_url": "",
      "post_date": "2024-05-21T12:25:26.053000",
      "content": "<p>Congratulations for your solution! Are you uploading the preprocessing/training script you used in your model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2834859,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2024-05-25T04:31:39.270000",
          "content": "<p>Sorry for the delay! I've been busy with work.</p>\n<p>I have shared the training and inference ipynb file of my best models.<br>\nI will upload the preprocess files within a few days.</p>\n<p><a href=\"https://github.com/yamagishi0824/spr_head_ct_age_1st\" target=\"_blank\">https://github.com/yamagishi0824/spr_head_ct_age_1st</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2838676,
              "author_name": "Shreyas Daniel Gaddam",
              "author_url": "",
              "post_date": "2024-05-27T07:01:31.163000",
              "content": "<p>thank you! incredibly helpful</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2788266,
      "author_name": "Shreyas Daniel Gaddam",
      "author_url": "",
      "post_date": "2024-05-02T06:52:23.550000",
      "content": "<p>Congratulations! A very clear approach, thanks for sharing! May I ask what was your strategy (if there was any) for sampling which slices to select from a series?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2790022,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2024-05-03T01:16:59.103000",
          "content": "<p>Thank you for your question!</p>\n<p>During the training of the model, I used all images except for removing noise images as I wrote in the solution above. Furthermore, during inference, I used all slices in the 40th percentile of the caudal-end. The slices I adopted for inference were selected through a simple process without much scrutiny, so I think there might have been a better method.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2791011,
              "author_name": "Shreyas Daniel Gaddam",
              "author_url": "",
              "post_date": "2024-05-03T12:40:04.773000",
              "content": "<p>thank you, will you be open-sourcing the code? I'd love to explore and learn more</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2791017,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2024-05-03T12:43:32.757000",
              "content": "<p>Yes!<br>\nOnce I've cleaned up the codes, I plan to make them public!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2791044,
              "author_name": "Shreyas Daniel Gaddam",
              "author_url": "",
              "post_date": "2024-05-03T12:55:17.943000",
              "content": "<p>awesome thank you</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2786133,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2024-05-01T07:06:41.183000",
      "content": "<p>Thanks for sharing! Did you mean you used 2d instead of 2.5D? What loss function are you using?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2786146,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2024-05-01T07:18:49.033000",
          "content": "<p>That's right! It’s simply a normal 2D model, with windows from 2D images stacked in the channel direction. <br>\nThe same \"Age\" label is assigned to all slices within the same study. For the loss function, I used MAE loss (nn.L1Loss).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2786080": "First of all, thank you for organizing this wonderful competition.\nIt was very rewarding to work with such a large-scale dataset across multiple facilities.\n\nHere’s my solution:\n\n# Overview\nI used a single fold and a single model approach with tf-efficientnetv2-s, training on the age of each study and aggregating predictions from multiple slices to obtain the final prediction for each study.\n\n# Data Split\nI used the holdout method. Training and validation were performed using StratifiedGroupKFold. Age was used for stratification both as a continuous and as a categorical variable, with grouping by patient. For validation data: fold == 0, training data: fold != 0.\n\n# Preprocessing\n### Windowing\nInitially, I stacked images in the channel direction using brain parenchyma, bone, and soft tissue windows (3ch in total).\nreference: https://radiopaedia.org/articles/windowing-ct\n\n### Removing Unnecessary Images\nDue to likely anonymization, there was added noise around the images. To remove this, I applied rule-based processing. It was crucial to use the very high CT values of bone. In channels conditioned on bone, the maximum value should represent bone. Observations showed that 5-10% of images included only noise or soft tissues at the top of the head. For my bone window condition, the maximum pixel value of these images was 77. Consequently, I removed all images where the maximum pixel value in the bone window did not exceed 100.\n\n### No Resizing in x, y, z Directions\nInitially, I resized in x, y, and z directions to obtain 2D images, but this led to decreased performance. Ultimately, I performed resizing to 512x512 size without compressing the z-direction to keep sizes close to the original.\n\n# Modeling\nI only used tf-efficientnetv2-s, employing basic augmentations such as horizontal flip and cutout. The model was trained for 5 epochs using a scheduler that started with a initial lr of 0 and maximum lr of 1e-3, which was progressively reduced to 1e-6 through a warmup process.\n\n# Prediction Aggregation (Post-Processing)\nPredictions were obtained for varying numbers of slices per study. Simply taking the mean and median of these predictions, the median outperformed in terms of my validation score.\nFurther, I aggregated predictions considering the percentile position of each slice within a study. This approach, using the mean of predictions where slices were positioned in the tail 40th percentile, improved performance, reducing CV from about 4.0 to 3.6 (these are approximate values as I’m compiling this solution remotely). The public leaderboard score showed a slight improvement from 3.165 to 3.123, and a more notable improvement from 2.774 to 2.596 in the private leaderboard.\n\n# What Didn’t Work\n- 3D Architectures: I used 3D DenseNet121. Even the best models scored around 6.7 on the public leaderboard, and predictions were around 6.5 on the private leaderboard.\n\n# Untried but Potentially Effective Strategies\n- Second-stage Model: Although I used the aforementioned rule-based processing for prediction aggregation, developing a model that uses predictions from the first stage as features to output final predictions could potentially improve scores further.\n- Ensemble: Using an ensemble of multiple models or multi-fold models could also likely enhance scores.\n",
    "2827358": "Congratulations for your solution! Are you uploading the preprocessing/training script you used in your model?",
    "2788266": "Congratulations! A very clear approach, thanks for sharing! May I ask what was your strategy (if there was any) for sampling which slices to select from a series?",
    "2786133": "Thanks for sharing! Did you mean you used 2d instead of 2.5D? What loss function are you using?"
  }
}