{
  "id": 611893,
  "title": "4th Place Solution",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/611893",
  "author_name": "Harshit Sheoran",
  "post_date": "2025-10-15T11:56:38.092000",
  "votes": 44,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Sincere Thank You to RSNA and Kaggle for hosting this competition </p>\n<p>I tackled the challenge as a speedrun as I had only 14 days left on the competition when I started, which is nowhere near enough to cover everything one can do with this dataset, the dataset is super cool and I have barely scratched the surface. I managed to get in the gold medal line in 8 days.</p>\n<p>Finally a GM, yay!</p>\n<h3>Segmentation Model Failure</h3>\n<p>I started with training a segmentation model… which was not really happy to learn beyond 0.6 Dice in the short time I tried which was nowhere near enough for my confidence in the approach.</p>\n<h3>ROI Crop coordinates Model</h3>\n<p>I shifted to a ViT (dinov3) model that takes in input of equidistant 48 slices for each patient, input shape: 48x1x128x128 volume and predicts x1, x2, y1, y2 values of where the crop boundaries of segmetation mask should be at patient-level, this model was far more reliable and efficient for ROI cropping</p>\n<p>Original:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F9468fad56fbe01d66d223f4a50c0304f%2Funcrop.png?generation=1760528693093368&amp;alt=media\" alt=\"\"></p>\n<p>ROI Crop:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F3cea5be1e8254117722f8b66541e9682%2Froi_crop.png?generation=1760528860832407&amp;alt=media\" alt=\"\"></p>\n<p>I kept the ROI relatively large as I did not want to miss out on aneurisms, I tested with the coordinate data and doing it in this configuration let me keep 95%+ of aneurism locations inside the cropped region</p>\n<h3>Classification Model</h3>\n<p>I took ~2.2k samples in the train_localizer csv file and all samples from negative patients which gives me a dataset of:</p>\n<ul>\n<li>545k samples with ~2.2K positive samples and rest negative samples (1 in 250 positive), </li>\n<li>There are 14 labels in the dataset, which keeps it that model needs to predict 1/1700 values to be 1, rest 0s</li>\n</ul>\n<p>I managed to find a combination of pipeline where model starts to learn even in this highly imbalance without weighted sampling (weighted sampling did not help).</p>\n<p>My entire solution is this classification model and here’s how I improve it’s score (very simple once baseline starts learning)</p>\n<p>My baseline hyperparameters:</p>\n<ul>\n<li>Model: CoaT-Lite-Medium</li>\n<li>Optimizer: AdamW | Learning-Rate: 1e-4 cosine annealing down to ~3e-5</li>\n<li>Augmentations: HFlip, VFlip, no TTA</li>\n<li>Image size: 384x384</li>\n<li>Global Batch-Size: 192</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>CV</th>\n<th>Public/Private Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CoaT-Lite-Medium on ROI Crops</td>\n<td>0.805</td>\n<td>0.78/0.78</td>\n</tr>\n<tr>\n<td>+ Rotate (-25, 25)</td>\n<td>0.83</td>\n<td>0.8/0.81</td>\n</tr>\n<tr>\n<td>+ 2.5D (-2, 0, +2)</td>\n<td>0.86</td>\n<td>0.86/0.83</td>\n</tr>\n<tr>\n<td>+ Soft Pseudo-Distill on remaining 500k potentially positive samples</td>\n<td>0.89</td>\n<td>0.86/0.84</td>\n</tr>\n<tr>\n<td>Ensembling with MaxViT + CoaT</td>\n<td>0.896</td>\n<td>0.87/0.84</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>On inference, the model predicts on every image for each patient, and final predictions are simply max aggregated</p>\n<p><strong>Thank you for reading!</strong></p>\n<p>Inference Code <a href=\"https://www.kaggle.com/code/harshitsheoran/rsna-aneurysm-detection-demo-submission\" target=\"_blank\">here</a><br>\nTraining Code <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code\" target=\"_blank\">here</a><br>\nModel Weights <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 3302248,
      "postDate": "2025-10-15T11:56:38.093Z",
      "content": "<p>Sincere Thank You to RSNA and Kaggle for hosting this competition </p>\n<p>I tackled the challenge as a speedrun as I had only 14 days left on the competition when I started, which is nowhere near enough to cover everything one can do with this dataset, the dataset is super cool and I have barely scratched the surface. I managed to get in the gold medal line in 8 days.</p>\n<p>Finally a GM, yay!</p>\n<h3>Segmentation Model Failure</h3>\n<p>I started with training a segmentation model… which was not really happy to learn beyond 0.6 Dice in the short time I tried which was nowhere near enough for my confidence in the approach.</p>\n<h3>ROI Crop coordinates Model</h3>\n<p>I shifted to a ViT (dinov3) model that takes in input of equidistant 48 slices for each patient, input shape: 48x1x128x128 volume and predicts x1, x2, y1, y2 values of where the crop boundaries of segmetation mask should be at patient-level, this model was far more reliable and efficient for ROI cropping</p>\n<p>Original:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F9468fad56fbe01d66d223f4a50c0304f%2Funcrop.png?generation=1760528693093368&amp;alt=media\" alt=\"\"></p>\n<p>ROI Crop:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F3cea5be1e8254117722f8b66541e9682%2Froi_crop.png?generation=1760528860832407&amp;alt=media\" alt=\"\"></p>\n<p>I kept the ROI relatively large as I did not want to miss out on aneurisms, I tested with the coordinate data and doing it in this configuration let me keep 95%+ of aneurism locations inside the cropped region</p>\n<h3>Classification Model</h3>\n<p>I took ~2.2k samples in the train_localizer csv file and all samples from negative patients which gives me a dataset of:</p>\n<ul>\n<li>545k samples with ~2.2K positive samples and rest negative samples (1 in 250 positive), </li>\n<li>There are 14 labels in the dataset, which keeps it that model needs to predict 1/1700 values to be 1, rest 0s</li>\n</ul>\n<p>I managed to find a combination of pipeline where model starts to learn even in this highly imbalance without weighted sampling (weighted sampling did not help).</p>\n<p>My entire solution is this classification model and here’s how I improve it’s score (very simple once baseline starts learning)</p>\n<p>My baseline hyperparameters:</p>\n<ul>\n<li>Model: CoaT-Lite-Medium</li>\n<li>Optimizer: AdamW | Learning-Rate: 1e-4 cosine annealing down to ~3e-5</li>\n<li>Augmentations: HFlip, VFlip, no TTA</li>\n<li>Image size: 384x384</li>\n<li>Global Batch-Size: 192</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Approach</th>\n<th>CV</th>\n<th>Public/Private Leaderboard</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CoaT-Lite-Medium on ROI Crops</td>\n<td>0.805</td>\n<td>0.78/0.78</td>\n</tr>\n<tr>\n<td>+ Rotate (-25, 25)</td>\n<td>0.83</td>\n<td>0.8/0.81</td>\n</tr>\n<tr>\n<td>+ 2.5D (-2, 0, +2)</td>\n<td>0.86</td>\n<td>0.86/0.83</td>\n</tr>\n<tr>\n<td>+ Soft Pseudo-Distill on remaining 500k potentially positive samples</td>\n<td>0.89</td>\n<td>0.86/0.84</td>\n</tr>\n<tr>\n<td>Ensembling with MaxViT + CoaT</td>\n<td>0.896</td>\n<td>0.87/0.84</td>\n</tr>\n</tbody>\n</table>\n<p><br></p>\n<p>On inference, the model predicts on every image for each patient, and final predictions are simply max aggregated</p>\n<p><strong>Thank you for reading!</strong></p>\n<p>Inference Code <a href=\"https://www.kaggle.com/code/harshitsheoran/rsna-aneurysm-detection-demo-submission\" target=\"_blank\">here</a><br>\nTraining Code <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code\" target=\"_blank\">here</a><br>\nModel Weights <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Sincere Thank You to RSNA and Kaggle for hosting this competition \n\nI tackled the challenge as a speedrun as I had only 14 days left on the competition when I started, which is nowhere near enough to cover everything one can do with this dataset, the dataset is super cool and I have barely scratched the surface. I managed to get in the gold medal line in 8 days.\n\nFinally a GM, yay!\n\n### Segmentation Model Failure\n\nI started with training a segmentation model… which was not really happy to learn beyond 0.6 Dice in the short time I tried which was nowhere near enough for my confidence in the approach.\n\n### ROI Crop coordinates Model\n\nI shifted to a ViT (dinov3) model that takes in input of equidistant 48 slices for each patient, input shape: 48x1x128x128 volume and predicts x1, x2, y1, y2 values of where the crop boundaries of segmetation mask should be at patient-level, this model was far more reliable and efficient for ROI cropping\n\nOriginal:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F9468fad56fbe01d66d223f4a50c0304f%2Funcrop.png?generation=1760528693093368&alt=media)\n\nROI Crop:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F3cea5be1e8254117722f8b66541e9682%2Froi_crop.png?generation=1760528860832407&alt=media)\n\nI kept the ROI relatively large as I did not want to miss out on aneurisms, I tested with the coordinate data and doing it in this configuration let me keep 95%+ of aneurism locations inside the cropped region\n\n### Classification Model\n\nI took ~2.2k samples in the train_localizer csv file and all samples from negative patients which gives me a dataset of:\n- 545k samples with ~2.2K positive samples and rest negative samples (1 in 250 positive), \n- There are 14 labels in the dataset, which keeps it that model needs to predict 1/1700 values to be 1, rest 0s\n\nI managed to find a combination of pipeline where model starts to learn even in this highly imbalance without weighted sampling (weighted sampling did not help).\n\nMy entire solution is this classification model and here’s how I improve it’s score (very simple once baseline starts learning)\n\nMy baseline hyperparameters:\n- Model: CoaT-Lite-Medium\n- Optimizer: AdamW | Learning-Rate: 1e-4 cosine annealing down to ~3e-5\n- Augmentations: HFlip, VFlip, no TTA\n- Image size: 384x384\n- Global Batch-Size: 192\n\n| Approach | CV | Public/Private Leaderboard |\n|---|---|---|\n| CoaT-Lite-Medium on ROI Crops | 0.805 | 0.78/0.78 |\n| + Rotate (-25, 25) | 0.83 | 0.8/0.81 |\n| + 2.5D (-2, 0, +2) | 0.86 | 0.86/0.83 |\n| + Soft Pseudo-Distill on remaining 500k potentially positive samples | 0.89 | 0.86/0.84 |\n| Ensembling with MaxViT + CoaT | 0.896 | 0.87/0.84 |\n\n<br>\n\nOn inference, the model predicts on every image for each patient, and final predictions are simply max aggregated\n\n**Thank you for reading!**\n\n\nInference Code [here](https://www.kaggle.com/code/harshitsheoran/rsna-aneurysm-detection-demo-submission)\nTraining Code [here](https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code)\nModel Weights [here](https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models)",
      "votes": 44
    },
    {
      "id": 3304973,
      "postDate": "2025-10-21T17:41:09.837Z",
      "content": "<p>Amazing work! The stepwise improvements and final ensemble strategy are very insightful. Thanks for sharing your process and code!</p>",
      "rawMarkdown": "Amazing work! The stepwise improvements and final ensemble strategy are very insightful. Thanks for sharing your process and code!",
      "votes": 1
    },
    {
      "id": 3304486,
      "postDate": "2025-10-20T16:08:57.830Z",
      "content": "<p>Congrats!!Thank you very much for providing such a simple solution!<br>\nI have a question about Training.<br>\nWhen I was training my model, I used BCELoss and FocalLoss, but the training loss was quite unstable no matter how much I adjusted the learning rate or scheduler.<br>\nDid you experience the same issue in your training pipeline?<br>\nOr did you confirm that the loss converged smoothly?</p>",
      "rawMarkdown": "Congrats!!Thank you very much for providing such a simple solution!\nI have a question about Training.\nWhen I was training my model, I used BCELoss and FocalLoss, but the training loss was quite unstable no matter how much I adjusted the learning rate or scheduler.\nDid you experience the same issue in your training pipeline?\nOr did you confirm that the loss converged smoothly?",
      "votes": 1,
      "replies": [
        {
          "id": 3304509,
          "postDate": "2025-10-20T17:08:44.100Z",
          "content": "<p>Hi, Glad you liked my solution</p>\n<p>Your score is near 0.7 which indicates that the model is not learning to predict the aneurism, it is likely that model has been picking up on other clues like weight imbalance between several modalities or other clues in the image to produce a score near 0.7, once your model really starts learning the competition target (0.78+ in my experiments), the training/validation loss/score metrics gets stable</p>\n<p>I did face heavy unstability issues when my score was lower than 0.7</p>\n<p>Yes, in my most recent work of scores 0.8+ the loss converges smoothly</p>",
          "rawMarkdown": "Hi, Glad you liked my solution\n\nYour score is near 0.7 which indicates that the model is not learning to predict the aneurism, it is likely that model has been picking up on other clues like weight imbalance between several modalities or other clues in the image to produce a score near 0.7, once your model really starts learning the competition target (0.78+ in my experiments), the training/validation loss/score metrics gets stable\n\nI did face heavy unstability issues when my score was lower than 0.7\n\nYes, in my most recent work of scores 0.8+ the loss converges smoothly",
          "votes": 1,
          "replies": [
            {
              "id": 3304628,
              "postDate": "2025-10-21T00:57:32.513Z",
              "content": "<p>Thank you for your reply! So, if the training is going well, the loss stabilizes and converges — that’s reassuring to hear.</p>\n<p>By the way, I have a question: when the loss became stable (from 0.7 to 0.78), did you do anything in particular to achieve that? For example, did you add class weights or something similar?</p>",
              "rawMarkdown": "Thank you for your reply! So, if the training is going well, the loss stabilizes and converges — that’s reassuring to hear.\n\nBy the way, I have a question: when the loss became stable (from 0.7 to 0.78), did you do anything in particular to achieve that? For example, did you add class weights or something similar?"
            },
            {
              "id": 3304662,
              "postDate": "2025-10-21T03:42:04.980Z",
              "content": "<p>For me, it was data cleaning, I take 2k positive samples and 500k negative samples, it was hard for the model to start learning immediately</p>\n<p>with ROI crop I think my score went from 0.7 to 0.8 (baseline)</p>",
              "rawMarkdown": "For me, it was data cleaning, I take 2k positive samples and 500k negative samples, it was hard for the model to start learning immediately\n\nwith ROI crop I think my score went from 0.7 to 0.8 (baseline)"
            },
            {
              "id": 3304934,
              "postDate": "2025-10-21T15:19:54.337Z",
              "content": "<p>It seems that increasing the resolution based on the ROI is the key point in this comp.<br>\nI understand it well now. Thank you very much!!</p>",
              "rawMarkdown": "It seems that increasing the resolution based on the ROI is the key point in this comp.\nI understand it well now. Thank you very much!!"
            }
          ]
        }
      ]
    },
    {
      "id": 3303608,
      "postDate": "2025-10-18T14:06:33.007Z",
      "content": "<p>congrats. This is the simplest and most clean solution I've ever seen</p>",
      "rawMarkdown": "congrats. This is the simplest and most clean solution I've ever seen",
      "votes": 1
    },
    {
      "id": 3302606,
      "postDate": "2025-10-16T06:31:00.357Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
      "rawMarkdown": "Congrats @harshitsheoran ",
      "votes": 1
    },
    {
      "id": 3302267,
      "postDate": "2025-10-15T12:35:33.097Z",
      "content": "<p>Congratulations! This is remarkably simple!<br>\nAny chance you can share your training code so that your solution can be reproduced ?</p>",
      "rawMarkdown": "Congratulations! This is remarkably simple!\nAny chance you can share your training code so that your solution can be reproduced ?",
      "votes": 1,
      "replies": [
        {
          "id": 3302275,
          "postDate": "2025-10-15T12:57:06.300Z",
          "content": "<p>Yes, Simple is what one gets in 8 days :D, </p>\n<p>I actually did not plan to use only a 2.5D model but wanted to use it as a pretraining for a 2.5D+1D pipeline that takes in multiple images  but that was not improving the score any further for me, and I managed to make a really good 2.5D model so I ran with it to pseudo-distillation</p>\n<p>I am sure that you are really excited to experiment with my code so I am making a priliminary open (sorry it's unclean):<br>\n<a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code</a></p>\n<p>Ask if you are missing any file that I am forgetting to share</p>",
          "rawMarkdown": "Yes, Simple is what one gets in 8 days :D, \n\nI actually did not plan to use only a 2.5D model but wanted to use it as a pretraining for a 2.5D+1D pipeline that takes in multiple images  but that was not improving the score any further for me, and I managed to make a really good 2.5D model so I ran with it to pseudo-distillation\n\nI am sure that you are really excited to experiment with my code so I am making a priliminary open (sorry it's unclean):\nhttps://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code\n\nAsk if you are missing any file that I am forgetting to share",
          "votes": 2,
          "replies": [
            {
              "id": 3302279,
              "postDate": "2025-10-15T13:04:16.543Z",
              "content": "<p>Unclean is an understatement here ! 😅</p>",
              "rawMarkdown": "Unclean is an understatement here ! 😅"
            }
          ]
        }
      ]
    },
    {
      "id": 3304692,
      "postDate": "2025-10-21T05:04:41.740Z",
      "content": "<p>HI <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>, congratulations on your 4th-place finish! Your approach is simple yet very effective. I’m really impressed that you managed to develop everything within just 14 days, it must have taken quite some effort to verify the effectiveness of each component.</p>\n<p>I have two questions: </p>\n<ul>\n<li><p>it seems the 2.5D contributed the most to your approach. Our team didn’t experiment with different sampling ranges like (-2, 0, 2) or wider. How did the sampling range affect your results? Have you tried other ranges?</p></li>\n<li><p>Also, I’m curious, what was your motivation for choosing ROI classification instead of pursuing a more refined object detection approach? Our team observes doing localization + classification task have better result then only doing classification.</p></li>\n</ul>\n<p>Thank you!</p>",
      "rawMarkdown": "HI @harshitsheoran, congratulations on your 4th-place finish! Your approach is simple yet very effective. I’m really impressed that you managed to develop everything within just 14 days, it must have taken quite some effort to verify the effectiveness of each component.\n\nI have two questions: \n* it seems the 2.5D contributed the most to your approach. Our team didn’t experiment with different sampling ranges like (-2, 0, 2) or wider. How did the sampling range affect your results? Have you tried other ranges?\n\n* Also, I’m curious, what was your motivation for choosing ROI classification instead of pursuing a more refined object detection approach? Our team observes doing localization + classification task have better result then only doing classification.\n\nThank you!",
      "votes": 2,
      "replies": [
        {
          "id": 3304716,
          "postDate": "2025-10-21T06:43:55.047Z",
          "content": "<p>Hi Tom,</p>\n<p>Thank you for appreciating my work</p>\n<p>14 days meant limited experiments on my part and I wish that I could have refined every method more much</p>\n<p>I had tried with [-1, 0, 1] earlier but at that time my sampling was positive-overweighted and it did not improve (neither did [-2,0,2]), I did not check on [-1,0,1] when the model reached 0.8 AUC</p>\n<p>ROI classification is simpler, yet effective, it's a tried and tested technique that I have developed many times, being short on time and thinking about complexity that an object detection approach would have both to train and infer, direct ROI coordinate prediction was just a couple hours of work from idea to finish.</p>",
          "rawMarkdown": "Hi Tom,\n\nThank you for appreciating my work\n\n14 days meant limited experiments on my part and I wish that I could have refined every method more much\n\nI had tried with [-1, 0, 1] earlier but at that time my sampling was positive-overweighted and it did not improve (neither did [-2,0,2]), I did not check on [-1,0,1] when the model reached 0.8 AUC\n\nROI classification is simpler, yet effective, it's a tried and tested technique that I have developed many times, being short on time and thinking about complexity that an object detection approach would have both to train and infer, direct ROI coordinate prediction was just a couple hours of work from idea to finish.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3304499,
      "postDate": "2025-10-20T16:25:32.570Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> congrats on GM!, Your solution is very simple and elegant. I am trying to replicate your solution by taking some insights from your <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code\" target=\"_blank\">prelims-code</a>. But it doesn't seems to have the training pipeline for cropping (maybe I just missed it). I am curious to what IoU score you got when doing such cropping using dinoV3? I am getting around ~0.65 on val</p>",
      "rawMarkdown": "Hey @harshitsheoran congrats on GM!, Your solution is very simple and elegant. I am trying to replicate your solution by taking some insights from your [prelims-code](https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code). But it doesn't seems to have the training pipeline for cropping (maybe I just missed it). I am curious to what IoU score you got when doing such cropping using dinoV3? I am getting around ~0.65 on val",
      "votes": 2,
      "replies": [
        {
          "id": 3304511,
          "postDate": "2025-10-20T17:20:05.293Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>,</p>\n<p>Congratulations to you too, Happy that you liked my solution and are excited to dive in</p>\n<p>For cropping, I am training a ViT-Small-Plus DinoV3 pretrained model, you can use any model really…, to take input a volume <code>(48, 1, 128, 128)</code> and predict <code>(x1, x2, y1, y2)</code> for where the volume should be cropped<br>\nx1, x2, y1, y2 are calculated by taking xmin, xmax, ymin, ymax around the 3D mask volume<br>\nThe metric I am using is MAE, I get an MAE of 0.027 (xy values normalized to 0 and 1)</p>\n<p>for cropping I take <code>image[0.9*ymin:1.1*ymax, 0.9*xmin:1.1*xmax]</code></p>\n<p>I am sharing just about all of my code to reproduce everything:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code</a></p>\n<p>And all the raw logs, OOF outputs, and model weights:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models/data\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models/data</a></p>",
          "rawMarkdown": "Hi @iamparadox,\n\nCongratulations to you too, Happy that you liked my solution and are excited to dive in\n\nFor cropping, I am training a ViT-Small-Plus DinoV3 pretrained model, you can use any model really..., to take input a volume `(48, 1, 128, 128)` and predict `(x1, x2, y1, y2)` for where the volume should be cropped\nx1, x2, y1, y2 are calculated by taking xmin, xmax, ymin, ymax around the 3D mask volume\nThe metric I am using is MAE, I get an MAE of 0.027 (xy values normalized to 0 and 1)\n\nfor cropping I take `image[0.9*ymin:1.1*ymax, 0.9*xmin:1.1*xmax]`\n\nI am sharing just about all of my code to reproduce everything:\n\nhttps://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code\n\nAnd all the raw logs, OOF outputs, and model weights:\n\nhttps://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models/data",
          "votes": 1
        }
      ]
    },
    {
      "id": 3303159,
      "postDate": "2025-10-17T09:00:15.577Z",
      "content": "<p>Congratz on GM !</p>\n<p>Looks like we went for quite similar approaches, yours is much simpler though, I'm impressed it scores that well.</p>",
      "rawMarkdown": "Congratz on GM !\n\nLooks like we went for quite similar approaches, yours is much simpler though, I'm impressed it scores that well.",
      "votes": 2,
      "replies": [
        {
          "id": 3303183,
          "postDate": "2025-10-17T10:10:54.810Z",
          "content": "<p>Thank you and congrats to you too on developing a very stable solution (public vs private)</p>\n<p>Yes, I think our approach aligns a lot, If I had time to work on each area, I may have developed to something as refined as yours</p>\n<p>The score difference is a bit confusing but it might just be coat-lite-medium being an exceptionally good model in my experiments (on CV, did not compare on LB)</p>",
          "rawMarkdown": "Thank you and congrats to you too on developing a very stable solution (public vs private)\n\nYes, I think our approach aligns a lot, If I had time to work on each area, I may have developed to something as refined as yours\n\nThe score difference is a bit confusing but it might just be coat-lite-medium being an exceptionally good model in my experiments (on CV, did not compare on LB)",
          "votes": 2,
          "replies": [
            {
              "id": 3304106,
              "postDate": "2025-10-19T20:11:26.277Z",
              "content": "<p>I tried coat-lite-medium as well, it's CV was not as good. Did not submit though, maybe it does well on LB.</p>",
              "rawMarkdown": "I tried coat-lite-medium as well, it's CV was not as good. Did not submit though, maybe it does well on LB.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3401660,
      "postDate": "2026-02-04T08:51:16.153Z",
      "content": "<p>Nice solution <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>  .. may I know what gpu you used for training the model and the training time in approx ?</p>",
      "rawMarkdown": "Nice solution @harshitsheoran  .. may I know what gpu you used for training the model and the training time in approx ?"
    },
    {
      "id": 3302401,
      "postDate": "2025-10-15T18:29:06.170Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3304973,
      "author_name": "Dimpi Mittal",
      "author_url": "",
      "post_date": "2025-10-21T17:41:09.837000",
      "content": "<p>Amazing work! The stepwise improvements and final ensemble strategy are very insightful. Thanks for sharing your process and code!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3304486,
      "author_name": "shiba-inu",
      "author_url": "",
      "post_date": "2025-10-20T16:08:57.830000",
      "content": "<p>Congrats!!Thank you very much for providing such a simple solution!<br>\nI have a question about Training.<br>\nWhen I was training my model, I used BCELoss and FocalLoss, but the training loss was quite unstable no matter how much I adjusted the learning rate or scheduler.<br>\nDid you experience the same issue in your training pipeline?<br>\nOr did you confirm that the loss converged smoothly?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3304509,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2025-10-20T17:08:44.100000",
          "content": "<p>Hi, Glad you liked my solution</p>\n<p>Your score is near 0.7 which indicates that the model is not learning to predict the aneurism, it is likely that model has been picking up on other clues like weight imbalance between several modalities or other clues in the image to produce a score near 0.7, once your model really starts learning the competition target (0.78+ in my experiments), the training/validation loss/score metrics gets stable</p>\n<p>I did face heavy unstability issues when my score was lower than 0.7</p>\n<p>Yes, in my most recent work of scores 0.8+ the loss converges smoothly</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3304628,
              "author_name": "shiba-inu",
              "author_url": "",
              "post_date": "2025-10-21T00:57:32.513000",
              "content": "<p>Thank you for your reply! So, if the training is going well, the loss stabilizes and converges — that’s reassuring to hear.</p>\n<p>By the way, I have a question: when the loss became stable (from 0.7 to 0.78), did you do anything in particular to achieve that? For example, did you add class weights or something similar?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3304662,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2025-10-21T03:42:04.980000",
              "content": "<p>For me, it was data cleaning, I take 2k positive samples and 500k negative samples, it was hard for the model to start learning immediately</p>\n<p>with ROI crop I think my score went from 0.7 to 0.8 (baseline)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3304934,
              "author_name": "shiba-inu",
              "author_url": "",
              "post_date": "2025-10-21T15:19:54.337000",
              "content": "<p>It seems that increasing the resolution based on the ROI is the key point in this comp.<br>\nI understand it well now. Thank you very much!!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3303608,
      "author_name": "ban #wave",
      "author_url": "",
      "post_date": "2025-10-18T14:06:33.007000",
      "content": "<p>congrats. This is the simplest and most clean solution I've ever seen</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3302606,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2025-10-16T06:31:00.357000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3302267,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2025-10-15T12:35:33.097000",
      "content": "<p>Congratulations! This is remarkably simple!<br>\nAny chance you can share your training code so that your solution can be reproduced ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3302275,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2025-10-15T12:57:06.300000",
          "content": "<p>Yes, Simple is what one gets in 8 days :D, </p>\n<p>I actually did not plan to use only a 2.5D model but wanted to use it as a pretraining for a 2.5D+1D pipeline that takes in multiple images  but that was not improving the score any further for me, and I managed to make a really good 2.5D model so I ran with it to pseudo-distillation</p>\n<p>I am sure that you are really excited to experiment with my code so I am making a priliminary open (sorry it's unclean):<br>\n<a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code</a></p>\n<p>Ask if you are missing any file that I am forgetting to share</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3302279,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2025-10-15T13:04:16.543000",
              "content": "<p>Unclean is an understatement here ! 😅</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3304692,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-10-21T05:04:41.740000",
      "content": "<p>HI <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>, congratulations on your 4th-place finish! Your approach is simple yet very effective. I’m really impressed that you managed to develop everything within just 14 days, it must have taken quite some effort to verify the effectiveness of each component.</p>\n<p>I have two questions: </p>\n<ul>\n<li><p>it seems the 2.5D contributed the most to your approach. Our team didn’t experiment with different sampling ranges like (-2, 0, 2) or wider. How did the sampling range affect your results? Have you tried other ranges?</p></li>\n<li><p>Also, I’m curious, what was your motivation for choosing ROI classification instead of pursuing a more refined object detection approach? Our team observes doing localization + classification task have better result then only doing classification.</p></li>\n</ul>\n<p>Thank you!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3304716,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2025-10-21T06:43:55.047000",
          "content": "<p>Hi Tom,</p>\n<p>Thank you for appreciating my work</p>\n<p>14 days meant limited experiments on my part and I wish that I could have refined every method more much</p>\n<p>I had tried with [-1, 0, 1] earlier but at that time my sampling was positive-overweighted and it did not improve (neither did [-2,0,2]), I did not check on [-1,0,1] when the model reached 0.8 AUC</p>\n<p>ROI classification is simpler, yet effective, it's a tried and tested technique that I have developed many times, being short on time and thinking about complexity that an object detection approach would have both to train and infer, direct ROI coordinate prediction was just a couple hours of work from idea to finish.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3304499,
      "author_name": "IAmParadox",
      "author_url": "",
      "post_date": "2025-10-20T16:25:32.570000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> congrats on GM!, Your solution is very simple and elegant. I am trying to replicate your solution by taking some insights from your <a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code\" target=\"_blank\">prelims-code</a>. But it doesn't seems to have the training pipeline for cropping (maybe I just missed it). I am curious to what IoU score you got when doing such cropping using dinoV3? I am getting around ~0.65 on val</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3304511,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2025-10-20T17:20:05.293000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a>,</p>\n<p>Congratulations to you too, Happy that you liked my solution and are excited to dive in</p>\n<p>For cropping, I am training a ViT-Small-Plus DinoV3 pretrained model, you can use any model really…, to take input a volume <code>(48, 1, 128, 128)</code> and predict <code>(x1, x2, y1, y2)</code> for where the volume should be cropped<br>\nx1, x2, y1, y2 are calculated by taking xmin, xmax, ymin, ymax around the 3D mask volume<br>\nThe metric I am using is MAE, I get an MAE of 0.027 (xy values normalized to 0 and 1)</p>\n<p>for cropping I take <code>image[0.9*ymin:1.1*ymax, 0.9*xmin:1.1*xmax]</code></p>\n<p>I am sharing just about all of my code to reproduce everything:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code</a></p>\n<p>And all the raw logs, OOF outputs, and model weights:</p>\n<p><a href=\"https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models/data\" target=\"_blank\">https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models/data</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3303159,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2025-10-17T09:00:15.577000",
      "content": "<p>Congratz on GM !</p>\n<p>Looks like we went for quite similar approaches, yours is much simpler though, I'm impressed it scores that well.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3303183,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2025-10-17T10:10:54.810000",
          "content": "<p>Thank you and congrats to you too on developing a very stable solution (public vs private)</p>\n<p>Yes, I think our approach aligns a lot, If I had time to work on each area, I may have developed to something as refined as yours</p>\n<p>The score difference is a bit confusing but it might just be coat-lite-medium being an exceptionally good model in my experiments (on CV, did not compare on LB)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3304106,
              "author_name": "Theo Viel",
              "author_url": "",
              "post_date": "2025-10-19T20:11:26.277000",
              "content": "<p>I tried coat-lite-medium as well, it's CV was not as good. Did not submit though, maybe it does well on LB.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3401660,
      "author_name": "AK",
      "author_url": "",
      "post_date": "2026-02-04T08:51:16.153000",
      "content": "<p>Nice solution <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a>  .. may I know what gpu you used for training the model and the training time in approx ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3302401,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-10-15T18:29:06.170000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3302248": "Sincere Thank You to RSNA and Kaggle for hosting this competition \n\nI tackled the challenge as a speedrun as I had only 14 days left on the competition when I started, which is nowhere near enough to cover everything one can do with this dataset, the dataset is super cool and I have barely scratched the surface. I managed to get in the gold medal line in 8 days.\n\nFinally a GM, yay!\n\n### Segmentation Model Failure\n\nI started with training a segmentation model… which was not really happy to learn beyond 0.6 Dice in the short time I tried which was nowhere near enough for my confidence in the approach.\n\n### ROI Crop coordinates Model\n\nI shifted to a ViT (dinov3) model that takes in input of equidistant 48 slices for each patient, input shape: 48x1x128x128 volume and predicts x1, x2, y1, y2 values of where the crop boundaries of segmetation mask should be at patient-level, this model was far more reliable and efficient for ROI cropping\n\nOriginal:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F9468fad56fbe01d66d223f4a50c0304f%2Funcrop.png?generation=1760528693093368&alt=media)\n\nROI Crop:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1794509%2F3cea5be1e8254117722f8b66541e9682%2Froi_crop.png?generation=1760528860832407&alt=media)\n\nI kept the ROI relatively large as I did not want to miss out on aneurisms, I tested with the coordinate data and doing it in this configuration let me keep 95%+ of aneurism locations inside the cropped region\n\n### Classification Model\n\nI took ~2.2k samples in the train_localizer csv file and all samples from negative patients which gives me a dataset of:\n- 545k samples with ~2.2K positive samples and rest negative samples (1 in 250 positive), \n- There are 14 labels in the dataset, which keeps it that model needs to predict 1/1700 values to be 1, rest 0s\n\nI managed to find a combination of pipeline where model starts to learn even in this highly imbalance without weighted sampling (weighted sampling did not help).\n\nMy entire solution is this classification model and here’s how I improve it’s score (very simple once baseline starts learning)\n\nMy baseline hyperparameters:\n- Model: CoaT-Lite-Medium\n- Optimizer: AdamW | Learning-Rate: 1e-4 cosine annealing down to ~3e-5\n- Augmentations: HFlip, VFlip, no TTA\n- Image size: 384x384\n- Global Batch-Size: 192\n\n| Approach | CV | Public/Private Leaderboard |\n|---|---|---|\n| CoaT-Lite-Medium on ROI Crops | 0.805 | 0.78/0.78 |\n| + Rotate (-25, 25) | 0.83 | 0.8/0.81 |\n| + 2.5D (-2, 0, +2) | 0.86 | 0.86/0.83 |\n| + Soft Pseudo-Distill on remaining 500k potentially positive samples | 0.89 | 0.86/0.84 |\n| Ensembling with MaxViT + CoaT | 0.896 | 0.87/0.84 |\n\n<br>\n\nOn inference, the model predicts on every image for each patient, and final predictions are simply max aggregated\n\n**Thank you for reading!**\n\n\nInference Code [here](https://www.kaggle.com/code/harshitsheoran/rsna-aneurysm-detection-demo-submission)\nTraining Code [here](https://www.kaggle.com/datasets/harshitsheoran/rsna2025-training-code)\nModel Weights [here](https://www.kaggle.com/datasets/harshitsheoran/rsna2025-raw-models)",
    "3304973": "Amazing work! The stepwise improvements and final ensemble strategy are very insightful. Thanks for sharing your process and code!",
    "3304486": "Congrats!!Thank you very much for providing such a simple solution!\nI have a question about Training.\nWhen I was training my model, I used BCELoss and FocalLoss, but the training loss was quite unstable no matter how much I adjusted the learning rate or scheduler.\nDid you experience the same issue in your training pipeline?\nOr did you confirm that the loss converged smoothly?",
    "3303608": "congrats. This is the simplest and most clean solution I've ever seen",
    "3302606": "Congrats @harshitsheoran ",
    "3302267": "Congratulations! This is remarkably simple!\nAny chance you can share your training code so that your solution can be reproduced ?",
    "3304692": "HI @harshitsheoran, congratulations on your 4th-place finish! Your approach is simple yet very effective. I’m really impressed that you managed to develop everything within just 14 days, it must have taken quite some effort to verify the effectiveness of each component.\n\nI have two questions: \n* it seems the 2.5D contributed the most to your approach. Our team didn’t experiment with different sampling ranges like (-2, 0, 2) or wider. How did the sampling range affect your results? Have you tried other ranges?\n\n* Also, I’m curious, what was your motivation for choosing ROI classification instead of pursuing a more refined object detection approach? Our team observes doing localization + classification task have better result then only doing classification.\n\nThank you!",
    "3304499": "Hey @harshitsheoran congrats on GM!, Your solution is very simple and elegant. I am trying to replicate your solution by taking some insights from your [prelims-code](https://www.kaggle.com/datasets/harshitsheoran/rsna-prilim-code). But it doesn't seems to have the training pipeline for cropping (maybe I just missed it). I am curious to what IoU score you got when doing such cropping using dinoV3? I am getting around ~0.65 on val",
    "3303159": "Congratz on GM !\n\nLooks like we went for quite similar approaches, yours is much simpler though, I'm impressed it scores that well.",
    "3401660": "Nice solution @harshitsheoran  .. may I know what gpu you used for training the model and the training time in approx ?",
    "3302401": ""
  }
}