{
  "id": 447779,
  "title": "17th Place Solution - How to Learn and Practice as a Beginner",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447779",
  "author_name": "m1dsolo",
  "post_date": "2023-10-17T07:18:58.855000",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to RSNA for hosting this interesting competition and congrats to all the winnners for their hard work.<br>\nI am a beginner who changed my major from physics to software engineering for just one year, I like to learn from practice, so I choose the kaggle competition platform which has many excellent learning resources.</p>\n<p>I am very happy that I can go from having no knowledge about 3D image processing to beating the baseline and finally getting the silver medal. I hope my competition experience can bring you some inspiration, especially the beginners like me who are struggling to beat the baseline at first.</p>\n<h1>Method</h1>\n<h3>Summary</h3>\n<p>For learning purposes, I plan to try all of the 2d classification, 3d classification, 2d segmentation, 3d segmentation in this competition. So my pipeline may be a little complicated.</p>\n<h3>Stage1: 2D + 3D segmentation</h3>\n<ol>\n<li><code>2D UNet</code> to segment liver, spleen, kidney_left, kidney_right and bowel.</li>\n<li><code>3D UNet</code> to further finely segment spleen.</li>\n</ol>\n<h3>Stage2: liver, spleen, kidney: 3D classification</h3>\n<ol>\n<li>Crop organs from segmentation and resize: liver(64, 312, 312), spleen(80, 224, 224), kidney_left(40, 128, 128), kidney_right(40, 128, 128).</li>\n<li>Use 3D classifier <code>X3D_l</code> to classify liver, spleen and kidney(For kidney, only backpropagate kidney with higher probability of being positive)</li>\n</ol>\n<h3>Stage3: bowel: 2.5D + 1D classification</h3>\n<ol>\n<li>Crop bowel by using bbox of bowel's masks(only choose pixels of the mask &gt; 1000 for each slice)</li>\n<li>Sample 64 slice uniformly from the cropped bowel and resize each slice to (512, 512), each slice has 4 channels(z-1, z, z+1, mask). So the input shape is (B, N, 4, 512, 512), B is batch size and N is num of slice.</li>\n<li>Use <code>convnext_tiny</code> as feature extractor. Input data(B, N, 4, 512, 512) after passing through the feature extractor will be converted into features(B, N, 768).<br>\nIf <code>N &lt; 64</code>, will use zero features to padding it. so the final output features are shape of (B, 64, 768).</li>\n<li>Use <code>lstm</code> + <code>attention pooling concat maxpooling</code> to fusion features.</li>\n<li>Use <code>nn.Linear</code> to classify.</li>\n</ol>\n<h3>Stage4: extravasation: 2.5D + 1D classification</h3>\n<ol>\n<li>Sample 64 slice uniformly, For each slice I use 5crop(top_left, top_right, bottom_left, bottom_right, center), and then resize each crop to (512, 512), then stack them. So the input shape is (B, N, 5, 3, 512, 512), B is batch size and N is num of slice.</li>\n<li>Use <code>convnext_tiny</code> as feature extractor. Input data(B, N, 5, 3, 512, 512) will be converted into features(B, N, 5, 768).<br>\nIf <code>N &lt; 64</code>, will use zero features to padding it. so the final output features are shape of (B, 64, 5, 768).</li>\n<li>Use <code>attention_pooling concat maxpooling</code> to fuse 5 features of each slice, so the output is shape of (B, 64, 768).</li>\n<li>Use <code>lstm</code> + <code>attention pooling concat maxpooling</code> to fusion features.</li>\n<li>Use <code>nn.Linear</code> to classify.</li>\n</ol>\n<h3>Some useful details:</h3>\n<ol>\n<li>When I started working on the extravasation classification, I saw <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">IAN PAN's extravasation bbox discussion</a>(Thanks!), I think it will help with the classification of extravasation.<br>\nFor positive sample, I use it with <code>albumentations.BBoxSafeRandomCropFixedSize(288, 288)</code>, it is a useful function of data augmentation that can random crop a part of the image around bbox. So it can help <code>convnext_tiny</code> to pay more attention to small area.<br>\nFor negative sample, I use <code>albumentations.RandomCrop(288, 288)</code> to random crop a part of input.</li>\n<li>I count every positive bowel slice, I find all of them <code>mask.sum() &gt;= 1000</code>(mask is generated by TotalSegmentator). So I think the slice which <code>mask.sum() &lt; 1000</code> can be ignored, thus making the network more concentrated in the ROI area.</li>\n<li>By <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/430242\" target=\"_blank\">NISCHAY DHANKHAR's 3rd solution</a>, I learned to use pseudo labeled dataset for initial training with a large learning rate and finally use a fine-labeled dataset for fine-tuning. So when I train 2D UNet, I use TotalSegmentator's prediction as pseudo label then use 206 fine segmentation to fine-tuning it.</li>\n<li>Since about 5% TotalSegmentator's zero-shot segmentations have serious problems, so I use trained 2D UNet segmentation to calculate the dice with it. If it is less than 0.75, it will be reconsidered.</li>\n<li>For Notebook Out of Memory:</li>\n</ol>\n<ul>\n<li>by <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/443256\" target=\"_blank\">Shai Ronen's discussion</a>, convert 2D dicom to uint8 as soon as loaded and delete + <code>gc.collect()</code> helped a lot.</li>\n<li>reduce <code>DataLoader.num_workers</code> can save RAM, set <code>DataLoader.pin_memory=False</code> can save GPU memory.</li>\n</ul>\n<h3>Things may work</h3>\n<ol>\n<li>For the task which has only small area of contrast, <code>maxpooling</code> may better than <code>avgpooling</code></li>\n<li>Training <code>convnext_tiny</code> + <code>lstm</code> + <code>attention pooling concat maxpooling</code> + <code>nn.Linear</code> end-to-end.<br>\n(I see local CV increased a lot, But I didn't have enough time to submit it before the end of competition)</li>\n</ol>\n<h3>Things may not work</h3>\n<ol>\n<li>Use <code>uniformer</code> instead of <code>X3D</code></li>\n<li>Use <code>efficientnetv2_s</code> instead of <code>convnext_tiny</code></li>\n</ol>\n<h1>For beginners like me</h1>\n<p>I know it will be a little hard when you first join a competition which unfamiliar to you.<br>\nI will list some personal experience to help beginners take the first step.</p>\n<h3>Before joining the competition</h3>\n<p>Before joining the competition, you must first clarify the competition task type(3D CT multi-label classification) and estimate the required calculation and required computing resources and hard disk capacity.<br>\nThis is used to decide whether you should join this competition, because you may be distressed when you don't have enough hard disk capacity to store the processed data.</p>\n<h3>Before coding</h3>\n<p>For an unfamiliar task(3D CT classification), the best way to get started is to read solultions from similar competitions that have ended.<br>\nIt just so happens that RSNA has held many similar competitions in the past few years.<br>\nI find <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection\" target=\"_blank\">RSNA 2022 Cervical Spine Fracture Detection (last year)</a> and <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/\" target=\"_blank\">RSNA STR Pulmonary Embolism Detection (3 years ago)</a> are highly relevant to this competition, their tasks are all 3D CT classification.<br>\nYou should read a lot about the winner's solution to decide the method of your own experiments.<br>\nMy finding are that 2d backbone to extract features(or 2.5d) + 1d rnn to fusion them tend to perform best.</p>\n<h3>About coding</h3>\n<p>After reading the top solutions, you have two routes to develop your own pipeline.<br>\nFirst is to copy and edit public code, but you should understand every line of code.<br>\nSecond is to refer to public code and use your own programming habits or code segment to develop pipeline.<br>\nI use the second method because it can greatly improve my coding ability.<br>\nIn my learning process of deep learning in the last year, I continuously accumulate and write a <a href=\"https://github.com/m1dsolo/yangDL\" target=\"_blank\">simple pytorch-based framework for multi-fold train, val, test, predict</a> framework.<br>\nIt can exercise my coding skills and greatly improves the speed of my coding, it feels really good to have any code in your own hands.<br>\nAt the same time, this accumulation of code will also be beneficial to similar tasks in the future.</p>\n<h3>Design method</h3>\n<p>Most people's methods can't beat the baseline propbably because they just throw 3D data to the network for training.<br>\nI did this too at the beginning, but the results were poor.<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">GUANSHUO XU's 1st solution</a> tells me that cropping ROI is really important.<br>\nMaybe because it is difficult for the network to learn ROI from just a few thousand of training data in 3d classification task.<br>\nSo I think segmenting each organ is an important first step.</p>\n<p>Since there are 2D classification labels for bowel and extravasation, so I decide to use 2d classification method for both of them and 3d classification method for liver, spleen and kidney.</p>\n<p>By <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/364837\" target=\"_blank\">Selim's 4th solution (CSN)</a> and <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362651\" target=\"_blank\">IAN PAN's 6th solution (X3D)</a>, it seems that using <code>transformer</code> as backbone is not good for small amounts of training data. Finally I choose <code>X3D_l</code> as my 3d classifier.</p>\n<p>Next is to design the 2d + 1d method for bowel and extravasation classification.<br>\nBy <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607\" target=\"_blank\">QISHEN HA's 1st solution</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/392449\" target=\"_blank\">Đăng Nguyễn Hồng's 1st solution</a>, <code>convnext</code> should be a good choice as 2D features extractor.<br>\nBy <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">Guanshuo Xu's 1st solution</a>, model of <code>lstm</code> + <code>attention pooling concat maxpooling</code> is selected by me to fuse features.</p>\n<p>For segmentation, because I already wrote 2D segmentation code during last competition <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature\" target=\"_blank\">hubmap-hacking-the-human-vasculature</a>, so I just decide to train a 2D UNet in early experiments.<br>\nLater I discovered that 2D UNet was not very good for segmenting spleen (maybe because my bad training skill), so I use 3D UNet to further refine segment spleen by inputting only data cropped by bbox of 2D UNet's coarse segmentation (spleen dice: 0.880 -&gt; 0.943)</p>\n<h1>Further</h1>\n<p>After I briefly readed other winning teams' solution, I list some tips that I might try in the future.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447449\" target=\"_blank\">NISCHAY DHANKHAR's 1st solution</a>: Auxiliary Segmentation Loss, 3D segmentation, generate image level label from series level label by organ visibility.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447453\" target=\"_blank\">THEO VIEL's 2nd solution</a>: infer which organs are present each of the slice, heavy augmentation, 3D <code>resnet18</code> to crop organs, use <code>RNN</code> to aggregate information from previous model and optimize the competition metric directly.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447464\" target=\"_blank\">YUJIARIYASU's 3rd solution</a>: enlarge mask before crop, input all organs into one model.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447549\" target=\"_blank\">LLREDA's 7th solution</a>: idea of <code>Mask2Former</code>, use image level label to assist feature learning, method of crop.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447706\" target=\"_blank\">IAN PAN's 8th solution</a>: use square root to scale probabilities.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447506\" target=\"_blank\">KAPENON's 9th solution</a>: post-process to improve the optimization of any_injury.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447450\" target=\"_blank\">YU4U's 10th solution</a>: stacking model to optimize any_injury, upsample positive smaples, train <code>max(logits, gt)</code> instead of <code>gt</code> because noise in image level label, <code>region_crop()</code> to remove outer black areas.</li>\n</ol>",
  "messages": [
    {
      "id": 2485416,
      "postDate": "2023-10-17T07:18:58.857Z",
      "content": "<p>Thanks to RSNA for hosting this interesting competition and congrats to all the winnners for their hard work.<br>\nI am a beginner who changed my major from physics to software engineering for just one year, I like to learn from practice, so I choose the kaggle competition platform which has many excellent learning resources.</p>\n<p>I am very happy that I can go from having no knowledge about 3D image processing to beating the baseline and finally getting the silver medal. I hope my competition experience can bring you some inspiration, especially the beginners like me who are struggling to beat the baseline at first.</p>\n<h1>Method</h1>\n<h3>Summary</h3>\n<p>For learning purposes, I plan to try all of the 2d classification, 3d classification, 2d segmentation, 3d segmentation in this competition. So my pipeline may be a little complicated.</p>\n<h3>Stage1: 2D + 3D segmentation</h3>\n<ol>\n<li><code>2D UNet</code> to segment liver, spleen, kidney_left, kidney_right and bowel.</li>\n<li><code>3D UNet</code> to further finely segment spleen.</li>\n</ol>\n<h3>Stage2: liver, spleen, kidney: 3D classification</h3>\n<ol>\n<li>Crop organs from segmentation and resize: liver(64, 312, 312), spleen(80, 224, 224), kidney_left(40, 128, 128), kidney_right(40, 128, 128).</li>\n<li>Use 3D classifier <code>X3D_l</code> to classify liver, spleen and kidney(For kidney, only backpropagate kidney with higher probability of being positive)</li>\n</ol>\n<h3>Stage3: bowel: 2.5D + 1D classification</h3>\n<ol>\n<li>Crop bowel by using bbox of bowel's masks(only choose pixels of the mask &gt; 1000 for each slice)</li>\n<li>Sample 64 slice uniformly from the cropped bowel and resize each slice to (512, 512), each slice has 4 channels(z-1, z, z+1, mask). So the input shape is (B, N, 4, 512, 512), B is batch size and N is num of slice.</li>\n<li>Use <code>convnext_tiny</code> as feature extractor. Input data(B, N, 4, 512, 512) after passing through the feature extractor will be converted into features(B, N, 768).<br>\nIf <code>N &lt; 64</code>, will use zero features to padding it. so the final output features are shape of (B, 64, 768).</li>\n<li>Use <code>lstm</code> + <code>attention pooling concat maxpooling</code> to fusion features.</li>\n<li>Use <code>nn.Linear</code> to classify.</li>\n</ol>\n<h3>Stage4: extravasation: 2.5D + 1D classification</h3>\n<ol>\n<li>Sample 64 slice uniformly, For each slice I use 5crop(top_left, top_right, bottom_left, bottom_right, center), and then resize each crop to (512, 512), then stack them. So the input shape is (B, N, 5, 3, 512, 512), B is batch size and N is num of slice.</li>\n<li>Use <code>convnext_tiny</code> as feature extractor. Input data(B, N, 5, 3, 512, 512) will be converted into features(B, N, 5, 768).<br>\nIf <code>N &lt; 64</code>, will use zero features to padding it. so the final output features are shape of (B, 64, 5, 768).</li>\n<li>Use <code>attention_pooling concat maxpooling</code> to fuse 5 features of each slice, so the output is shape of (B, 64, 768).</li>\n<li>Use <code>lstm</code> + <code>attention pooling concat maxpooling</code> to fusion features.</li>\n<li>Use <code>nn.Linear</code> to classify.</li>\n</ol>\n<h3>Some useful details:</h3>\n<ol>\n<li>When I started working on the extravasation classification, I saw <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">IAN PAN's extravasation bbox discussion</a>(Thanks!), I think it will help with the classification of extravasation.<br>\nFor positive sample, I use it with <code>albumentations.BBoxSafeRandomCropFixedSize(288, 288)</code>, it is a useful function of data augmentation that can random crop a part of the image around bbox. So it can help <code>convnext_tiny</code> to pay more attention to small area.<br>\nFor negative sample, I use <code>albumentations.RandomCrop(288, 288)</code> to random crop a part of input.</li>\n<li>I count every positive bowel slice, I find all of them <code>mask.sum() &gt;= 1000</code>(mask is generated by TotalSegmentator). So I think the slice which <code>mask.sum() &lt; 1000</code> can be ignored, thus making the network more concentrated in the ROI area.</li>\n<li>By <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/430242\" target=\"_blank\">NISCHAY DHANKHAR's 3rd solution</a>, I learned to use pseudo labeled dataset for initial training with a large learning rate and finally use a fine-labeled dataset for fine-tuning. So when I train 2D UNet, I use TotalSegmentator's prediction as pseudo label then use 206 fine segmentation to fine-tuning it.</li>\n<li>Since about 5% TotalSegmentator's zero-shot segmentations have serious problems, so I use trained 2D UNet segmentation to calculate the dice with it. If it is less than 0.75, it will be reconsidered.</li>\n<li>For Notebook Out of Memory:</li>\n</ol>\n<ul>\n<li>by <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/443256\" target=\"_blank\">Shai Ronen's discussion</a>, convert 2D dicom to uint8 as soon as loaded and delete + <code>gc.collect()</code> helped a lot.</li>\n<li>reduce <code>DataLoader.num_workers</code> can save RAM, set <code>DataLoader.pin_memory=False</code> can save GPU memory.</li>\n</ul>\n<h3>Things may work</h3>\n<ol>\n<li>For the task which has only small area of contrast, <code>maxpooling</code> may better than <code>avgpooling</code></li>\n<li>Training <code>convnext_tiny</code> + <code>lstm</code> + <code>attention pooling concat maxpooling</code> + <code>nn.Linear</code> end-to-end.<br>\n(I see local CV increased a lot, But I didn't have enough time to submit it before the end of competition)</li>\n</ol>\n<h3>Things may not work</h3>\n<ol>\n<li>Use <code>uniformer</code> instead of <code>X3D</code></li>\n<li>Use <code>efficientnetv2_s</code> instead of <code>convnext_tiny</code></li>\n</ol>\n<h1>For beginners like me</h1>\n<p>I know it will be a little hard when you first join a competition which unfamiliar to you.<br>\nI will list some personal experience to help beginners take the first step.</p>\n<h3>Before joining the competition</h3>\n<p>Before joining the competition, you must first clarify the competition task type(3D CT multi-label classification) and estimate the required calculation and required computing resources and hard disk capacity.<br>\nThis is used to decide whether you should join this competition, because you may be distressed when you don't have enough hard disk capacity to store the processed data.</p>\n<h3>Before coding</h3>\n<p>For an unfamiliar task(3D CT classification), the best way to get started is to read solultions from similar competitions that have ended.<br>\nIt just so happens that RSNA has held many similar competitions in the past few years.<br>\nI find <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection\" target=\"_blank\">RSNA 2022 Cervical Spine Fracture Detection (last year)</a> and <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/\" target=\"_blank\">RSNA STR Pulmonary Embolism Detection (3 years ago)</a> are highly relevant to this competition, their tasks are all 3D CT classification.<br>\nYou should read a lot about the winner's solution to decide the method of your own experiments.<br>\nMy finding are that 2d backbone to extract features(or 2.5d) + 1d rnn to fusion them tend to perform best.</p>\n<h3>About coding</h3>\n<p>After reading the top solutions, you have two routes to develop your own pipeline.<br>\nFirst is to copy and edit public code, but you should understand every line of code.<br>\nSecond is to refer to public code and use your own programming habits or code segment to develop pipeline.<br>\nI use the second method because it can greatly improve my coding ability.<br>\nIn my learning process of deep learning in the last year, I continuously accumulate and write a <a href=\"https://github.com/m1dsolo/yangDL\" target=\"_blank\">simple pytorch-based framework for multi-fold train, val, test, predict</a> framework.<br>\nIt can exercise my coding skills and greatly improves the speed of my coding, it feels really good to have any code in your own hands.<br>\nAt the same time, this accumulation of code will also be beneficial to similar tasks in the future.</p>\n<h3>Design method</h3>\n<p>Most people's methods can't beat the baseline propbably because they just throw 3D data to the network for training.<br>\nI did this too at the beginning, but the results were poor.<br>\n<a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">GUANSHUO XU's 1st solution</a> tells me that cropping ROI is really important.<br>\nMaybe because it is difficult for the network to learn ROI from just a few thousand of training data in 3d classification task.<br>\nSo I think segmenting each organ is an important first step.</p>\n<p>Since there are 2D classification labels for bowel and extravasation, so I decide to use 2d classification method for both of them and 3d classification method for liver, spleen and kidney.</p>\n<p>By <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/364837\" target=\"_blank\">Selim's 4th solution (CSN)</a> and <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362651\" target=\"_blank\">IAN PAN's 6th solution (X3D)</a>, it seems that using <code>transformer</code> as backbone is not good for small amounts of training data. Finally I choose <code>X3D_l</code> as my 3d classifier.</p>\n<p>Next is to design the 2d + 1d method for bowel and extravasation classification.<br>\nBy <a href=\"https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607\" target=\"_blank\">QISHEN HA's 1st solution</a> and <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/392449\" target=\"_blank\">Đăng Nguyễn Hồng's 1st solution</a>, <code>convnext</code> should be a good choice as 2D features extractor.<br>\nBy <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">Guanshuo Xu's 1st solution</a>, model of <code>lstm</code> + <code>attention pooling concat maxpooling</code> is selected by me to fuse features.</p>\n<p>For segmentation, because I already wrote 2D segmentation code during last competition <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature\" target=\"_blank\">hubmap-hacking-the-human-vasculature</a>, so I just decide to train a 2D UNet in early experiments.<br>\nLater I discovered that 2D UNet was not very good for segmenting spleen (maybe because my bad training skill), so I use 3D UNet to further refine segment spleen by inputting only data cropped by bbox of 2D UNet's coarse segmentation (spleen dice: 0.880 -&gt; 0.943)</p>\n<h1>Further</h1>\n<p>After I briefly readed other winning teams' solution, I list some tips that I might try in the future.</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447449\" target=\"_blank\">NISCHAY DHANKHAR's 1st solution</a>: Auxiliary Segmentation Loss, 3D segmentation, generate image level label from series level label by organ visibility.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447453\" target=\"_blank\">THEO VIEL's 2nd solution</a>: infer which organs are present each of the slice, heavy augmentation, 3D <code>resnet18</code> to crop organs, use <code>RNN</code> to aggregate information from previous model and optimize the competition metric directly.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447464\" target=\"_blank\">YUJIARIYASU's 3rd solution</a>: enlarge mask before crop, input all organs into one model.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447549\" target=\"_blank\">LLREDA's 7th solution</a>: idea of <code>Mask2Former</code>, use image level label to assist feature learning, method of crop.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447706\" target=\"_blank\">IAN PAN's 8th solution</a>: use square root to scale probabilities.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447506\" target=\"_blank\">KAPENON's 9th solution</a>: post-process to improve the optimization of any_injury.</li>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447450\" target=\"_blank\">YU4U's 10th solution</a>: stacking model to optimize any_injury, upsample positive smaples, train <code>max(logits, gt)</code> instead of <code>gt</code> because noise in image level label, <code>region_crop()</code> to remove outer black areas.</li>\n</ol>",
      "rawMarkdown": "Thanks to RSNA for hosting this interesting competition and congrats to all the winnners for their hard work.\nI am a beginner who changed my major from physics to software engineering for just one year, I like to learn from practice, so I choose the kaggle competition platform which has many excellent learning resources.\n\nI am very happy that I can go from having no knowledge about 3D image processing to beating the baseline and finally getting the silver medal. I hope my competition experience can bring you some inspiration, especially the beginners like me who are struggling to beat the baseline at first.\n\n# Method\n\n### Summary\n\nFor learning purposes, I plan to try all of the 2d classification, 3d classification, 2d segmentation, 3d segmentation in this competition. So my pipeline may be a little complicated.\n\n### Stage1: 2D + 3D segmentation\n1. `2D UNet` to segment liver, spleen, kidney_left, kidney_right and bowel.\n2. `3D UNet` to further finely segment spleen.\n\n### Stage2: liver, spleen, kidney: 3D classification\n\n1. Crop organs from segmentation and resize: liver(64, 312, 312), spleen(80, 224, 224), kidney_left(40, 128, 128), kidney_right(40, 128, 128).\n2. Use 3D classifier `X3D_l` to classify liver, spleen and kidney(For kidney, only backpropagate kidney with higher probability of being positive)\n\n### Stage3: bowel: 2.5D + 1D classification\n\n1. Crop bowel by using bbox of bowel's masks(only choose pixels of the mask > 1000 for each slice)\n2. Sample 64 slice uniformly from the cropped bowel and resize each slice to (512, 512), each slice has 4 channels(z-1, z, z+1, mask). So the input shape is (B, N, 4, 512, 512), B is batch size and N is num of slice.\n3. Use `convnext_tiny` as feature extractor. Input data(B, N, 4, 512, 512) after passing through the feature extractor will be converted into features(B, N, 768).\nIf `N < 64`, will use zero features to padding it. so the final output features are shape of (B, 64, 768).\n4. Use `lstm` + `attention pooling concat maxpooling` to fusion features.\n5. Use `nn.Linear` to classify.\n\n### Stage4: extravasation: 2.5D + 1D classification\n\n1. Sample 64 slice uniformly, For each slice I use 5crop(top_left, top_right, bottom_left, bottom_right, center), and then resize each crop to (512, 512), then stack them. So the input shape is (B, N, 5, 3, 512, 512), B is batch size and N is num of slice.\n2. Use `convnext_tiny` as feature extractor. Input data(B, N, 5, 3, 512, 512) will be converted into features(B, N, 5, 768).\nIf `N < 64`, will use zero features to padding it. so the final output features are shape of (B, 64, 5, 768).\n3. Use `attention_pooling concat maxpooling` to fuse 5 features of each slice, so the output is shape of (B, 64, 768).\n4. Use `lstm` + `attention pooling concat maxpooling` to fusion features.\n5. Use `nn.Linear` to classify.\n\n### Some useful details:\n\n1. When I started working on the extravasation classification, I saw [IAN PAN's extravasation bbox discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402)(Thanks!), I think it will help with the classification of extravasation.\nFor positive sample, I use it with `albumentations.BBoxSafeRandomCropFixedSize(288, 288)`, it is a useful function of data augmentation that can random crop a part of the image around bbox. So it can help `convnext_tiny` to pay more attention to small area.\nFor negative sample, I use `albumentations.RandomCrop(288, 288)` to random crop a part of input.\n2. I count every positive bowel slice, I find all of them `mask.sum() >= 1000`(mask is generated by TotalSegmentator). So I think the slice which `mask.sum() < 1000` can be ignored, thus making the network more concentrated in the ROI area.\n3. By [NISCHAY DHANKHAR's 3rd solution](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/430242), I learned to use pseudo labeled dataset for initial training with a large learning rate and finally use a fine-labeled dataset for fine-tuning. So when I train 2D UNet, I use TotalSegmentator's prediction as pseudo label then use 206 fine segmentation to fine-tuning it.\n4. Since about 5% TotalSegmentator's zero-shot segmentations have serious problems, so I use trained 2D UNet segmentation to calculate the dice with it. If it is less than 0.75, it will be reconsidered.\n5. For Notebook Out of Memory:\n- by [Shai Ronen's discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/443256), convert 2D dicom to uint8 as soon as loaded and delete + `gc.collect()` helped a lot.\n- reduce `DataLoader.num_workers` can save RAM, set `DataLoader.pin_memory=False` can save GPU memory.\n\n### Things may work\n\n1. For the task which has only small area of contrast, `maxpooling` may better than `avgpooling`\n2. Training `convnext_tiny` + `lstm` + `attention pooling concat maxpooling` + `nn.Linear` end-to-end.\n(I see local CV increased a lot, But I didn't have enough time to submit it before the end of competition)\n\n### Things may not work\n\n1. Use `uniformer` instead of `X3D`\n2. Use `efficientnetv2_s` instead of `convnext_tiny`\n\n# For beginners like me\n\nI know it will be a little hard when you first join a competition which unfamiliar to you.\nI will list some personal experience to help beginners take the first step.\n\n### Before joining the competition\n\nBefore joining the competition, you must first clarify the competition task type(3D CT multi-label classification) and estimate the required calculation and required computing resources and hard disk capacity.\nThis is used to decide whether you should join this competition, because you may be distressed when you don't have enough hard disk capacity to store the processed data.\n\n### Before coding\n\nFor an unfamiliar task(3D CT classification), the best way to get started is to read solultions from similar competitions that have ended.\nIt just so happens that RSNA has held many similar competitions in the past few years.\nI find [RSNA 2022 Cervical Spine Fracture Detection (last year)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection) and [RSNA STR Pulmonary Embolism Detection (3 years ago)](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/) are highly relevant to this competition, their tasks are all 3D CT classification.\nYou should read a lot about the winner's solution to decide the method of your own experiments.\nMy finding are that 2d backbone to extract features(or 2.5d) + 1d rnn to fusion them tend to perform best.\n\n### About coding\n\nAfter reading the top solutions, you have two routes to develop your own pipeline.\nFirst is to copy and edit public code, but you should understand every line of code.\nSecond is to refer to public code and use your own programming habits or code segment to develop pipeline.\nI use the second method because it can greatly improve my coding ability.\nIn my learning process of deep learning in the last year, I continuously accumulate and write a [simple pytorch-based framework for multi-fold train, val, test, predict](https://github.com/m1dsolo/yangDL) framework.\nIt can exercise my coding skills and greatly improves the speed of my coding, it feels really good to have any code in your own hands.\nAt the same time, this accumulation of code will also be beneficial to similar tasks in the future.\n\n### Design method\n\nMost people's methods can't beat the baseline propbably because they just throw 3D data to the network for training.\nI did this too at the beginning, but the results were poor.\n[GUANSHUO XU's 1st solution](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) tells me that cropping ROI is really important.\nMaybe because it is difficult for the network to learn ROI from just a few thousand of training data in 3d classification task.\nSo I think segmenting each organ is an important first step.\n\nSince there are 2D classification labels for bowel and extravasation, so I decide to use 2d classification method for both of them and 3d classification method for liver, spleen and kidney.\n\nBy [Selim's 4th solution (CSN)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/364837) and [IAN PAN's 6th solution (X3D)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362651), it seems that using `transformer` as backbone is not good for small amounts of training data. Finally I choose `X3D_l` as my 3d classifier.\n\nNext is to design the 2d + 1d method for bowel and extravasation classification.\nBy [QISHEN HA's 1st solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607) and [Đăng Nguyễn Hồng's 1st solution](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/392449), `convnext` should be a good choice as 2D features extractor.\nBy [Guanshuo Xu's 1st solution](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/194145), model of `lstm` + `attention pooling concat maxpooling` is selected by me to fuse features.\n\nFor segmentation, because I already wrote 2D segmentation code during last competition [hubmap-hacking-the-human-vasculature](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature), so I just decide to train a 2D UNet in early experiments.\nLater I discovered that 2D UNet was not very good for segmenting spleen (maybe because my bad training skill), so I use 3D UNet to further refine segment spleen by inputting only data cropped by bbox of 2D UNet's coarse segmentation (spleen dice: 0.880 -> 0.943)\n\n# Further\n\nAfter I briefly readed other winning teams' solution, I list some tips that I might try in the future.\n\n1. [NISCHAY DHANKHAR's 1st solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447449): Auxiliary Segmentation Loss, 3D segmentation, generate image level label from series level label by organ visibility.\n2. [THEO VIEL's 2nd solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447453): infer which organs are present each of the slice, heavy augmentation, 3D `resnet18` to crop organs, use `RNN` to aggregate information from previous model and optimize the competition metric directly.\n3. [YUJIARIYASU's 3rd solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447464): enlarge mask before crop, input all organs into one model.\n4. [LLREDA's 7th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447549): idea of `Mask2Former`, use image level label to assist feature learning, method of crop.\n5. [IAN PAN's 8th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447706): use square root to scale probabilities.\n6. [KAPENON's 9th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447506): post-process to improve the optimization of any_injury.\n7. [YU4U's 10th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447450): stacking model to optimize any_injury, upsample positive smaples, train `max(logits, gt)` instead of `gt` because noise in image level label, `region_crop()` to remove outer black areas.\n",
      "votes": 20
    },
    {
      "id": 2485666,
      "postDate": "2023-10-17T11:05:05.970Z",
      "content": "<p>Nice work. Thanks for the great write up!</p>",
      "rawMarkdown": "Nice work. Thanks for the great write up!",
      "votes": 1
    },
    {
      "id": 2486591,
      "postDate": "2023-10-18T03:56:58.697Z",
      "content": "<p>Thank you for sharing. Very helpful </p>",
      "rawMarkdown": "Thank you for sharing. Very helpful ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2485666,
      "author_name": "Cuog Nguyen",
      "author_url": "",
      "post_date": "2023-10-17T11:05:05.970000",
      "content": "<p>Nice work. Thanks for the great write up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2486591,
      "author_name": "crazy_rabbit_AI",
      "author_url": "",
      "post_date": "2023-10-18T03:56:58.697000",
      "content": "<p>Thank you for sharing. Very helpful </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2485416": "Thanks to RSNA for hosting this interesting competition and congrats to all the winnners for their hard work.\nI am a beginner who changed my major from physics to software engineering for just one year, I like to learn from practice, so I choose the kaggle competition platform which has many excellent learning resources.\n\nI am very happy that I can go from having no knowledge about 3D image processing to beating the baseline and finally getting the silver medal. I hope my competition experience can bring you some inspiration, especially the beginners like me who are struggling to beat the baseline at first.\n\n# Method\n\n### Summary\n\nFor learning purposes, I plan to try all of the 2d classification, 3d classification, 2d segmentation, 3d segmentation in this competition. So my pipeline may be a little complicated.\n\n### Stage1: 2D + 3D segmentation\n1. `2D UNet` to segment liver, spleen, kidney_left, kidney_right and bowel.\n2. `3D UNet` to further finely segment spleen.\n\n### Stage2: liver, spleen, kidney: 3D classification\n\n1. Crop organs from segmentation and resize: liver(64, 312, 312), spleen(80, 224, 224), kidney_left(40, 128, 128), kidney_right(40, 128, 128).\n2. Use 3D classifier `X3D_l` to classify liver, spleen and kidney(For kidney, only backpropagate kidney with higher probability of being positive)\n\n### Stage3: bowel: 2.5D + 1D classification\n\n1. Crop bowel by using bbox of bowel's masks(only choose pixels of the mask > 1000 for each slice)\n2. Sample 64 slice uniformly from the cropped bowel and resize each slice to (512, 512), each slice has 4 channels(z-1, z, z+1, mask). So the input shape is (B, N, 4, 512, 512), B is batch size and N is num of slice.\n3. Use `convnext_tiny` as feature extractor. Input data(B, N, 4, 512, 512) after passing through the feature extractor will be converted into features(B, N, 768).\nIf `N < 64`, will use zero features to padding it. so the final output features are shape of (B, 64, 768).\n4. Use `lstm` + `attention pooling concat maxpooling` to fusion features.\n5. Use `nn.Linear` to classify.\n\n### Stage4: extravasation: 2.5D + 1D classification\n\n1. Sample 64 slice uniformly, For each slice I use 5crop(top_left, top_right, bottom_left, bottom_right, center), and then resize each crop to (512, 512), then stack them. So the input shape is (B, N, 5, 3, 512, 512), B is batch size and N is num of slice.\n2. Use `convnext_tiny` as feature extractor. Input data(B, N, 5, 3, 512, 512) will be converted into features(B, N, 5, 768).\nIf `N < 64`, will use zero features to padding it. so the final output features are shape of (B, 64, 5, 768).\n3. Use `attention_pooling concat maxpooling` to fuse 5 features of each slice, so the output is shape of (B, 64, 768).\n4. Use `lstm` + `attention pooling concat maxpooling` to fusion features.\n5. Use `nn.Linear` to classify.\n\n### Some useful details:\n\n1. When I started working on the extravasation classification, I saw [IAN PAN's extravasation bbox discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402)(Thanks!), I think it will help with the classification of extravasation.\nFor positive sample, I use it with `albumentations.BBoxSafeRandomCropFixedSize(288, 288)`, it is a useful function of data augmentation that can random crop a part of the image around bbox. So it can help `convnext_tiny` to pay more attention to small area.\nFor negative sample, I use `albumentations.RandomCrop(288, 288)` to random crop a part of input.\n2. I count every positive bowel slice, I find all of them `mask.sum() >= 1000`(mask is generated by TotalSegmentator). So I think the slice which `mask.sum() < 1000` can be ignored, thus making the network more concentrated in the ROI area.\n3. By [NISCHAY DHANKHAR's 3rd solution](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/430242), I learned to use pseudo labeled dataset for initial training with a large learning rate and finally use a fine-labeled dataset for fine-tuning. So when I train 2D UNet, I use TotalSegmentator's prediction as pseudo label then use 206 fine segmentation to fine-tuning it.\n4. Since about 5% TotalSegmentator's zero-shot segmentations have serious problems, so I use trained 2D UNet segmentation to calculate the dice with it. If it is less than 0.75, it will be reconsidered.\n5. For Notebook Out of Memory:\n- by [Shai Ronen's discussion](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/443256), convert 2D dicom to uint8 as soon as loaded and delete + `gc.collect()` helped a lot.\n- reduce `DataLoader.num_workers` can save RAM, set `DataLoader.pin_memory=False` can save GPU memory.\n\n### Things may work\n\n1. For the task which has only small area of contrast, `maxpooling` may better than `avgpooling`\n2. Training `convnext_tiny` + `lstm` + `attention pooling concat maxpooling` + `nn.Linear` end-to-end.\n(I see local CV increased a lot, But I didn't have enough time to submit it before the end of competition)\n\n### Things may not work\n\n1. Use `uniformer` instead of `X3D`\n2. Use `efficientnetv2_s` instead of `convnext_tiny`\n\n# For beginners like me\n\nI know it will be a little hard when you first join a competition which unfamiliar to you.\nI will list some personal experience to help beginners take the first step.\n\n### Before joining the competition\n\nBefore joining the competition, you must first clarify the competition task type(3D CT multi-label classification) and estimate the required calculation and required computing resources and hard disk capacity.\nThis is used to decide whether you should join this competition, because you may be distressed when you don't have enough hard disk capacity to store the processed data.\n\n### Before coding\n\nFor an unfamiliar task(3D CT classification), the best way to get started is to read solultions from similar competitions that have ended.\nIt just so happens that RSNA has held many similar competitions in the past few years.\nI find [RSNA 2022 Cervical Spine Fracture Detection (last year)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection) and [RSNA STR Pulmonary Embolism Detection (3 years ago)](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/) are highly relevant to this competition, their tasks are all 3D CT classification.\nYou should read a lot about the winner's solution to decide the method of your own experiments.\nMy finding are that 2d backbone to extract features(or 2.5d) + 1d rnn to fusion them tend to perform best.\n\n### About coding\n\nAfter reading the top solutions, you have two routes to develop your own pipeline.\nFirst is to copy and edit public code, but you should understand every line of code.\nSecond is to refer to public code and use your own programming habits or code segment to develop pipeline.\nI use the second method because it can greatly improve my coding ability.\nIn my learning process of deep learning in the last year, I continuously accumulate and write a [simple pytorch-based framework for multi-fold train, val, test, predict](https://github.com/m1dsolo/yangDL) framework.\nIt can exercise my coding skills and greatly improves the speed of my coding, it feels really good to have any code in your own hands.\nAt the same time, this accumulation of code will also be beneficial to similar tasks in the future.\n\n### Design method\n\nMost people's methods can't beat the baseline propbably because they just throw 3D data to the network for training.\nI did this too at the beginning, but the results were poor.\n[GUANSHUO XU's 1st solution](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145) tells me that cropping ROI is really important.\nMaybe because it is difficult for the network to learn ROI from just a few thousand of training data in 3d classification task.\nSo I think segmenting each organ is an important first step.\n\nSince there are 2D classification labels for bowel and extravasation, so I decide to use 2d classification method for both of them and 3d classification method for liver, spleen and kidney.\n\nBy [Selim's 4th solution (CSN)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/364837) and [IAN PAN's 6th solution (X3D)](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362651), it seems that using `transformer` as backbone is not good for small amounts of training data. Finally I choose `X3D_l` as my 3d classifier.\n\nNext is to design the 2d + 1d method for bowel and extravasation classification.\nBy [QISHEN HA's 1st solution](https://www.kaggle.com/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/362607) and [Đăng Nguyễn Hồng's 1st solution](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/392449), `convnext` should be a good choice as 2D features extractor.\nBy [Guanshuo Xu's 1st solution](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/194145), model of `lstm` + `attention pooling concat maxpooling` is selected by me to fuse features.\n\nFor segmentation, because I already wrote 2D segmentation code during last competition [hubmap-hacking-the-human-vasculature](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature), so I just decide to train a 2D UNet in early experiments.\nLater I discovered that 2D UNet was not very good for segmenting spleen (maybe because my bad training skill), so I use 3D UNet to further refine segment spleen by inputting only data cropped by bbox of 2D UNet's coarse segmentation (spleen dice: 0.880 -> 0.943)\n\n# Further\n\nAfter I briefly readed other winning teams' solution, I list some tips that I might try in the future.\n\n1. [NISCHAY DHANKHAR's 1st solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447449): Auxiliary Segmentation Loss, 3D segmentation, generate image level label from series level label by organ visibility.\n2. [THEO VIEL's 2nd solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447453): infer which organs are present each of the slice, heavy augmentation, 3D `resnet18` to crop organs, use `RNN` to aggregate information from previous model and optimize the competition metric directly.\n3. [YUJIARIYASU's 3rd solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447464): enlarge mask before crop, input all organs into one model.\n4. [LLREDA's 7th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447549): idea of `Mask2Former`, use image level label to assist feature learning, method of crop.\n5. [IAN PAN's 8th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447706): use square root to scale probabilities.\n6. [KAPENON's 9th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447506): post-process to improve the optimization of any_injury.\n7. [YU4U's 10th solution](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/447450): stacking model to optimize any_injury, upsample positive smaples, train `max(logits, gt)` instead of `gt` because noise in image level label, `region_crop()` to remove outer black areas.\n",
    "2485666": "Nice work. Thanks for the great write up!",
    "2486591": "Thank you for sharing. Very helpful "
  }
}