{
  "id": 453827,
  "title": "439th-Solution-RSNA_ATD",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/453827",
  "author_name": "Alex Luna",
  "post_date": "2023-11-07T22:21:05.719000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Place 439th Solution for the RSNA Abdominal Trauma Detection competition and Insights.</h1>\n<h2><strong>Context</strong></h2>\n<p>Clinical Context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview</a></p>\n<p>Data Context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data</a></p>\n<h2>Intro</h2>\n<p>Firstly thanks to the RSNA for hosting the Abdominal Trauma Detection competition. It was both challenging and well-structured. A big shout-out to our community for the insightful discussions and for demonstrating what's possible, we learn a lot not also by making research on reading papers but also reading the solution notebooks of the community. Congratulations to the winning teams; we learned a lot from the solution write-ups, also congratulations to our team <a href=\"https://www.kaggle.com/diegoramirezmendoza\" target=\"_blank\">@diegoramirezmendoza</a>, <a href=\"https://www.kaggle.com/pedromartnezbarrn\" target=\"_blank\">@pedromartnezbarrn</a> , <a href=\"https://www.kaggle.com/arantzabazalda\" target=\"_blank\">@arantzabazalda</a>, <a href=\"https://www.kaggle.com/eliudlimon\" target=\"_blank\">@eliudlimon</a> for such an amazing collaboration and perfect teamwork!</p>\n<h2>Overview</h2>\n<p>We implement 2 different pipelines for this competition, a 3D-CNN and a 2D-CNN embedded with a LSTM arquitecture, those were 2 different approaches.</p>\n<h2>EDA</h2>\n<p>As usual, we couldn't start this competition without an exploratory data analysis. The aim of this stage was to understand the data, its particular distribution, getting acquainted with the labels to classify, and exploring metadata that could provide better quality data preprocessing. By doing this, we were able to build a robust strategy based on the characteristics of the data and also considering the available resources. This is something we will discuss later in this document, but for now, let's say this first step (EDA), as a data scientist's best practice, provides good insights about the data and possible strategies, but also about the computational resources required.</p>\n<h2>Data Preprocessing</h2>\n<p>The data preprocessing step was something laborious, the objective was to generate scripts to automate the generation and transformations of the competitions data that was stored locally. Even though the developed static functions and methods to fit something like image data generators we observed that the time taken to preprocess all this data during training wasn´t really efficient, instead when training the proposed models the time to train a single epoch was enormously high.</p>\n<p>Essentially the \"standard\" part of the data preprocessing was to use the raw dicom files transformed to arrays, then rescaled them from the original shape to 128 x 128, normalized and fixed the pixel value representations due to some dicom files storage characteristics. From here we start exploring different models and possible solutions that summarizes our 3 proposed models, the first model is a 3D CNN based from the work of [], the second model is a CNN + LSTM layers and the third model was designed to train from scratch some state-of-the-art CNN´s like VGG16 and ResNet50b (which don´t gave us good results).</p>\n<h3>The strength of segmentations data</h3>\n<p>From the EDA stage, we gained insights into the segmentation data provided for the competition. These NIfTI files could offer a better understanding of the data and help develop a more refined data preprocessing pipeline.</p>\n<p>Although there is a limited amount of segmentation data, it was deemed sufficient for our purposes. The proposed methodology was to train a U-Net model from scratch using this data and integrate its predictions into the main pipeline. But why? The reason is that some of the scans are not as informative as we would like, and this is particularly important given the competition's objective to focus on abdominal trauma. Consequently, some of the scans above and below the abdomen can be considered as \"noise\".</p>\n<p>Based on this proposition, we decided to use the segmentation data to train a U-Net for the task of segmenting the organs, as provided by the masks in the NIfTI files. We then developed a threshold function to clean the inferences, using the first appearance of the liver as the upper limit and the last segment of the bowel as the lower limit. This strategy for reducing the data was successfully implemented and integrated into the preprocessing pipeline.</p>\n<p>The experiment setup for training the U-Net model consisted of several steps. First, we implemented the model by extracting data as numpy files. Second, we defined the hyperparameters for the model. Third, we evaluated the performance of the model.</p>\n<p>The final model used for the data reduction preprocessing pipeline was trained using the Adam optimizer with a learning rate of 0.000001. We used categorical crossentropy as the loss function, MeanIoU as our tunable metric and set the batch size to 128. The model was trained for 128 epochs, although due to GPU resource limitations on Google Colab, our training crashed somewhere around epoch 75 and total amount of data used for training was 15,520 DICOM and NIfTI files, which included both images and segmentation masks. Despite the challenges, the trained U-Net model provided valuable information for the preprocessing pipeline and allowed us to effectively reduce the data to focus on the region of interest in abdominal trauma cases.</p>\n<h2>Training Results:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2Fa610792c5a9897038986087883df671c%2Fdescarga.png?generation=1699395350287888&amp;alt=media\" alt=\"\"></p>\n<h2>Inference Results:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F7b77812541848956d381fbbf1416ed2e%2Fdescarga%20(2).png?generation=1699395316405025&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F28d7f4e90e094f9dbd48760e27c4b037%2Fdescarga%20(3).png?generation=1699395391353207&amp;alt=media\" alt=\"\"></p>\n<h2>3D CNN - data preprocessing</h2>\n<p>For the 3D CNN model, after completing our \"standard\" preprocessing procedure and fixing the scale to 128 x 128, the question of determining the depth of the volume arose. To answer this, we referred to the work from [] and considered the GPU capacity. Ultimately, we set the volume depth to 64.</p>\n<p>To get the data into the shape of 128x128x64, we used the zoom [] function from Python's scipy module. To automate this preprocessing task and data generation, we developed a preprocessing script that maps the patient data, storing the DICOM paths in a DataFrame (which was later saved as a CSV file and also stored in an SQLite database). From these mapped paths of the original data, we then distributed the specific \"maps\" to correctly generate the data.</p>\n<p>The data was generated and saved as numpy files, normalized to values between 0 and 1, with the final shape of 128x128x64 for each series (patients' folders store series, with some patients having only one series and others having two series). As a reminder, the total number of training data 128x128x64 data volumes generated sums up the complete number of series present in the competition's dataset.</p>\n<h2>CNN-LSTM - data preprocessing</h2>\n<p>On this model we use the down-sampling block of the U-net previously trained for the Semantic Segmentation using the masks, to extract the Feature vector of the image, this part of the U-net is called Encoder, the pipeline was:</p>\n<p>Semantic Segmentation with the U-net -&gt; Reduce Volume Shape to 128x128x64 (using Zoom Function) -&gt; Feature Extraction (Pretrained Encoder) -&gt; Bidirectional-LSTM<br>\nNote: In the Feature Extraction we notice there were many 0's in the Feature Vectors, we applied a drop to that (it doesn't give us important information) and the final shape to be inputed in the model was (Batch_size, 64, 624)</p>\n<h2>Models</h2>\n<ul>\n<li>3D CNN</li>\n<li>CNN-LSTM (Pretrained Encoder from U-net -&gt; Bi-LSTM)</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>3D-CNN model has bad results, we think that this kind of arquitecture only works with a very good cropping of the volumes, the bad performance is because of the black pixels who tend to overfit the model.<br>\nCurrently working on different approaches…</li>\n</ul>\n<h2>NOTEBOOKS</h2>\n<p>Training: (The train of the model was on a local computer)</p>\n<p>Inference: <a href=\"https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv\" target=\"_blank\">https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv</a></p>",
  "messages": [
    {
      "id": 2516716,
      "postDate": "2023-11-07T22:21:05.720Z",
      "content": "<h1>Place 439th Solution for the RSNA Abdominal Trauma Detection competition and Insights.</h1>\n<h2><strong>Context</strong></h2>\n<p>Clinical Context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview</a></p>\n<p>Data Context: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data</a></p>\n<h2>Intro</h2>\n<p>Firstly thanks to the RSNA for hosting the Abdominal Trauma Detection competition. It was both challenging and well-structured. A big shout-out to our community for the insightful discussions and for demonstrating what's possible, we learn a lot not also by making research on reading papers but also reading the solution notebooks of the community. Congratulations to the winning teams; we learned a lot from the solution write-ups, also congratulations to our team <a href=\"https://www.kaggle.com/diegoramirezmendoza\" target=\"_blank\">@diegoramirezmendoza</a>, <a href=\"https://www.kaggle.com/pedromartnezbarrn\" target=\"_blank\">@pedromartnezbarrn</a> , <a href=\"https://www.kaggle.com/arantzabazalda\" target=\"_blank\">@arantzabazalda</a>, <a href=\"https://www.kaggle.com/eliudlimon\" target=\"_blank\">@eliudlimon</a> for such an amazing collaboration and perfect teamwork!</p>\n<h2>Overview</h2>\n<p>We implement 2 different pipelines for this competition, a 3D-CNN and a 2D-CNN embedded with a LSTM arquitecture, those were 2 different approaches.</p>\n<h2>EDA</h2>\n<p>As usual, we couldn't start this competition without an exploratory data analysis. The aim of this stage was to understand the data, its particular distribution, getting acquainted with the labels to classify, and exploring metadata that could provide better quality data preprocessing. By doing this, we were able to build a robust strategy based on the characteristics of the data and also considering the available resources. This is something we will discuss later in this document, but for now, let's say this first step (EDA), as a data scientist's best practice, provides good insights about the data and possible strategies, but also about the computational resources required.</p>\n<h2>Data Preprocessing</h2>\n<p>The data preprocessing step was something laborious, the objective was to generate scripts to automate the generation and transformations of the competitions data that was stored locally. Even though the developed static functions and methods to fit something like image data generators we observed that the time taken to preprocess all this data during training wasn´t really efficient, instead when training the proposed models the time to train a single epoch was enormously high.</p>\n<p>Essentially the \"standard\" part of the data preprocessing was to use the raw dicom files transformed to arrays, then rescaled them from the original shape to 128 x 128, normalized and fixed the pixel value representations due to some dicom files storage characteristics. From here we start exploring different models and possible solutions that summarizes our 3 proposed models, the first model is a 3D CNN based from the work of [], the second model is a CNN + LSTM layers and the third model was designed to train from scratch some state-of-the-art CNN´s like VGG16 and ResNet50b (which don´t gave us good results).</p>\n<h3>The strength of segmentations data</h3>\n<p>From the EDA stage, we gained insights into the segmentation data provided for the competition. These NIfTI files could offer a better understanding of the data and help develop a more refined data preprocessing pipeline.</p>\n<p>Although there is a limited amount of segmentation data, it was deemed sufficient for our purposes. The proposed methodology was to train a U-Net model from scratch using this data and integrate its predictions into the main pipeline. But why? The reason is that some of the scans are not as informative as we would like, and this is particularly important given the competition's objective to focus on abdominal trauma. Consequently, some of the scans above and below the abdomen can be considered as \"noise\".</p>\n<p>Based on this proposition, we decided to use the segmentation data to train a U-Net for the task of segmenting the organs, as provided by the masks in the NIfTI files. We then developed a threshold function to clean the inferences, using the first appearance of the liver as the upper limit and the last segment of the bowel as the lower limit. This strategy for reducing the data was successfully implemented and integrated into the preprocessing pipeline.</p>\n<p>The experiment setup for training the U-Net model consisted of several steps. First, we implemented the model by extracting data as numpy files. Second, we defined the hyperparameters for the model. Third, we evaluated the performance of the model.</p>\n<p>The final model used for the data reduction preprocessing pipeline was trained using the Adam optimizer with a learning rate of 0.000001. We used categorical crossentropy as the loss function, MeanIoU as our tunable metric and set the batch size to 128. The model was trained for 128 epochs, although due to GPU resource limitations on Google Colab, our training crashed somewhere around epoch 75 and total amount of data used for training was 15,520 DICOM and NIfTI files, which included both images and segmentation masks. Despite the challenges, the trained U-Net model provided valuable information for the preprocessing pipeline and allowed us to effectively reduce the data to focus on the region of interest in abdominal trauma cases.</p>\n<h2>Training Results:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2Fa610792c5a9897038986087883df671c%2Fdescarga.png?generation=1699395350287888&amp;alt=media\" alt=\"\"></p>\n<h2>Inference Results:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F7b77812541848956d381fbbf1416ed2e%2Fdescarga%20(2).png?generation=1699395316405025&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F28d7f4e90e094f9dbd48760e27c4b037%2Fdescarga%20(3).png?generation=1699395391353207&amp;alt=media\" alt=\"\"></p>\n<h2>3D CNN - data preprocessing</h2>\n<p>For the 3D CNN model, after completing our \"standard\" preprocessing procedure and fixing the scale to 128 x 128, the question of determining the depth of the volume arose. To answer this, we referred to the work from [] and considered the GPU capacity. Ultimately, we set the volume depth to 64.</p>\n<p>To get the data into the shape of 128x128x64, we used the zoom [] function from Python's scipy module. To automate this preprocessing task and data generation, we developed a preprocessing script that maps the patient data, storing the DICOM paths in a DataFrame (which was later saved as a CSV file and also stored in an SQLite database). From these mapped paths of the original data, we then distributed the specific \"maps\" to correctly generate the data.</p>\n<p>The data was generated and saved as numpy files, normalized to values between 0 and 1, with the final shape of 128x128x64 for each series (patients' folders store series, with some patients having only one series and others having two series). As a reminder, the total number of training data 128x128x64 data volumes generated sums up the complete number of series present in the competition's dataset.</p>\n<h2>CNN-LSTM - data preprocessing</h2>\n<p>On this model we use the down-sampling block of the U-net previously trained for the Semantic Segmentation using the masks, to extract the Feature vector of the image, this part of the U-net is called Encoder, the pipeline was:</p>\n<p>Semantic Segmentation with the U-net -&gt; Reduce Volume Shape to 128x128x64 (using Zoom Function) -&gt; Feature Extraction (Pretrained Encoder) -&gt; Bidirectional-LSTM<br>\nNote: In the Feature Extraction we notice there were many 0's in the Feature Vectors, we applied a drop to that (it doesn't give us important information) and the final shape to be inputed in the model was (Batch_size, 64, 624)</p>\n<h2>Models</h2>\n<ul>\n<li>3D CNN</li>\n<li>CNN-LSTM (Pretrained Encoder from U-net -&gt; Bi-LSTM)</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>3D-CNN model has bad results, we think that this kind of arquitecture only works with a very good cropping of the volumes, the bad performance is because of the black pixels who tend to overfit the model.<br>\nCurrently working on different approaches…</li>\n</ul>\n<h2>NOTEBOOKS</h2>\n<p>Training: (The train of the model was on a local computer)</p>\n<p>Inference: <a href=\"https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv\" target=\"_blank\">https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv</a></p>",
      "rawMarkdown": "# Place 439th Solution for the RSNA Abdominal Trauma Detection competition and Insights.\n\n## **Context**\nClinical Context: https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\n\nData Context: https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\n\n## Intro\nFirstly thanks to the RSNA for hosting the Abdominal Trauma Detection competition. It was both challenging and well-structured. A big shout-out to our community for the insightful discussions and for demonstrating what's possible, we learn a lot not also by making research on reading papers but also reading the solution notebooks of the community. Congratulations to the winning teams; we learned a lot from the solution write-ups, also congratulations to our team @diegoramirezmendoza, @pedromartnezbarrn , @arantzabazalda, @eliudlimon for such an amazing collaboration and perfect teamwork!\n\n## Overview\nWe implement 2 different pipelines for this competition, a 3D-CNN and a 2D-CNN embedded with a LSTM arquitecture, those were 2 different approaches.\n\n## EDA\nAs usual, we couldn't start this competition without an exploratory data analysis. The aim of this stage was to understand the data, its particular distribution, getting acquainted with the labels to classify, and exploring metadata that could provide better quality data preprocessing. By doing this, we were able to build a robust strategy based on the characteristics of the data and also considering the available resources. This is something we will discuss later in this document, but for now, let's say this first step (EDA), as a data scientist's best practice, provides good insights about the data and possible strategies, but also about the computational resources required.\n\n## Data Preprocessing\nThe data preprocessing step was something laborious, the objective was to generate scripts to automate the generation and transformations of the competitions data that was stored locally. Even though the developed static functions and methods to fit something like image data generators we observed that the time taken to preprocess all this data during training wasn´t really efficient, instead when training the proposed models the time to train a single epoch was enormously high.\n\nEssentially the \"standard\" part of the data preprocessing was to use the raw dicom files transformed to arrays, then rescaled them from the original shape to 128 x 128, normalized and fixed the pixel value representations due to some dicom files storage characteristics. From here we start exploring different models and possible solutions that summarizes our 3 proposed models, the first model is a 3D CNN based from the work of [], the second model is a CNN + LSTM layers and the third model was designed to train from scratch some state-of-the-art CNN´s like VGG16 and ResNet50b (which don´t gave us good results).\n\n### The strength of segmentations data\nFrom the EDA stage, we gained insights into the segmentation data provided for the competition. These NIfTI files could offer a better understanding of the data and help develop a more refined data preprocessing pipeline.\n\nAlthough there is a limited amount of segmentation data, it was deemed sufficient for our purposes. The proposed methodology was to train a U-Net model from scratch using this data and integrate its predictions into the main pipeline. But why? The reason is that some of the scans are not as informative as we would like, and this is particularly important given the competition's objective to focus on abdominal trauma. Consequently, some of the scans above and below the abdomen can be considered as \"noise\".\n\nBased on this proposition, we decided to use the segmentation data to train a U-Net for the task of segmenting the organs, as provided by the masks in the NIfTI files. We then developed a threshold function to clean the inferences, using the first appearance of the liver as the upper limit and the last segment of the bowel as the lower limit. This strategy for reducing the data was successfully implemented and integrated into the preprocessing pipeline.\n\nThe experiment setup for training the U-Net model consisted of several steps. First, we implemented the model by extracting data as numpy files. Second, we defined the hyperparameters for the model. Third, we evaluated the performance of the model.\n\nThe final model used for the data reduction preprocessing pipeline was trained using the Adam optimizer with a learning rate of 0.000001. We used categorical crossentropy as the loss function, MeanIoU as our tunable metric and set the batch size to 128. The model was trained for 128 epochs, although due to GPU resource limitations on Google Colab, our training crashed somewhere around epoch 75 and total amount of data used for training was 15,520 DICOM and NIfTI files, which included both images and segmentation masks. Despite the challenges, the trained U-Net model provided valuable information for the preprocessing pipeline and allowed us to effectively reduce the data to focus on the region of interest in abdominal trauma cases.\n\n## Training Results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2Fa610792c5a9897038986087883df671c%2Fdescarga.png?generation=1699395350287888&alt=media)\n\n## Inference Results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F7b77812541848956d381fbbf1416ed2e%2Fdescarga%20(2).png?generation=1699395316405025&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F28d7f4e90e094f9dbd48760e27c4b037%2Fdescarga%20(3).png?generation=1699395391353207&alt=media)\n\n\n## 3D CNN - data preprocessing\nFor the 3D CNN model, after completing our \"standard\" preprocessing procedure and fixing the scale to 128 x 128, the question of determining the depth of the volume arose. To answer this, we referred to the work from [] and considered the GPU capacity. Ultimately, we set the volume depth to 64.\n\nTo get the data into the shape of 128x128x64, we used the zoom [] function from Python's scipy module. To automate this preprocessing task and data generation, we developed a preprocessing script that maps the patient data, storing the DICOM paths in a DataFrame (which was later saved as a CSV file and also stored in an SQLite database). From these mapped paths of the original data, we then distributed the specific \"maps\" to correctly generate the data.\n\nThe data was generated and saved as numpy files, normalized to values between 0 and 1, with the final shape of 128x128x64 for each series (patients' folders store series, with some patients having only one series and others having two series). As a reminder, the total number of training data 128x128x64 data volumes generated sums up the complete number of series present in the competition's dataset.\n\n## CNN-LSTM - data preprocessing\nOn this model we use the down-sampling block of the U-net previously trained for the Semantic Segmentation using the masks, to extract the Feature vector of the image, this part of the U-net is called Encoder, the pipeline was:\n\nSemantic Segmentation with the U-net -> Reduce Volume Shape to 128x128x64 (using Zoom Function) -> Feature Extraction (Pretrained Encoder) -> Bidirectional-LSTM\nNote: In the Feature Extraction we notice there were many 0's in the Feature Vectors, we applied a drop to that (it doesn't give us important information) and the final shape to be inputed in the model was (Batch_size, 64, 624)\n\n## Models\n* 3D CNN\n* CNN-LSTM (Pretrained Encoder from U-net -> Bi-LSTM)\n\n## What did not work\n* 3D-CNN model has bad results, we think that this kind of arquitecture only works with a very good cropping of the volumes, the bad performance is because of the black pixels who tend to overfit the model.\nCurrently working on different approaches...\n## NOTEBOOKS\n\nTraining: (The train of the model was on a local computer)\n\nInference: https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2516716": "# Place 439th Solution for the RSNA Abdominal Trauma Detection competition and Insights.\n\n## **Context**\nClinical Context: https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\n\nData Context: https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\n\n## Intro\nFirstly thanks to the RSNA for hosting the Abdominal Trauma Detection competition. It was both challenging and well-structured. A big shout-out to our community for the insightful discussions and for demonstrating what's possible, we learn a lot not also by making research on reading papers but also reading the solution notebooks of the community. Congratulations to the winning teams; we learned a lot from the solution write-ups, also congratulations to our team @diegoramirezmendoza, @pedromartnezbarrn , @arantzabazalda, @eliudlimon for such an amazing collaboration and perfect teamwork!\n\n## Overview\nWe implement 2 different pipelines for this competition, a 3D-CNN and a 2D-CNN embedded with a LSTM arquitecture, those were 2 different approaches.\n\n## EDA\nAs usual, we couldn't start this competition without an exploratory data analysis. The aim of this stage was to understand the data, its particular distribution, getting acquainted with the labels to classify, and exploring metadata that could provide better quality data preprocessing. By doing this, we were able to build a robust strategy based on the characteristics of the data and also considering the available resources. This is something we will discuss later in this document, but for now, let's say this first step (EDA), as a data scientist's best practice, provides good insights about the data and possible strategies, but also about the computational resources required.\n\n## Data Preprocessing\nThe data preprocessing step was something laborious, the objective was to generate scripts to automate the generation and transformations of the competitions data that was stored locally. Even though the developed static functions and methods to fit something like image data generators we observed that the time taken to preprocess all this data during training wasn´t really efficient, instead when training the proposed models the time to train a single epoch was enormously high.\n\nEssentially the \"standard\" part of the data preprocessing was to use the raw dicom files transformed to arrays, then rescaled them from the original shape to 128 x 128, normalized and fixed the pixel value representations due to some dicom files storage characteristics. From here we start exploring different models and possible solutions that summarizes our 3 proposed models, the first model is a 3D CNN based from the work of [], the second model is a CNN + LSTM layers and the third model was designed to train from scratch some state-of-the-art CNN´s like VGG16 and ResNet50b (which don´t gave us good results).\n\n### The strength of segmentations data\nFrom the EDA stage, we gained insights into the segmentation data provided for the competition. These NIfTI files could offer a better understanding of the data and help develop a more refined data preprocessing pipeline.\n\nAlthough there is a limited amount of segmentation data, it was deemed sufficient for our purposes. The proposed methodology was to train a U-Net model from scratch using this data and integrate its predictions into the main pipeline. But why? The reason is that some of the scans are not as informative as we would like, and this is particularly important given the competition's objective to focus on abdominal trauma. Consequently, some of the scans above and below the abdomen can be considered as \"noise\".\n\nBased on this proposition, we decided to use the segmentation data to train a U-Net for the task of segmenting the organs, as provided by the masks in the NIfTI files. We then developed a threshold function to clean the inferences, using the first appearance of the liver as the upper limit and the last segment of the bowel as the lower limit. This strategy for reducing the data was successfully implemented and integrated into the preprocessing pipeline.\n\nThe experiment setup for training the U-Net model consisted of several steps. First, we implemented the model by extracting data as numpy files. Second, we defined the hyperparameters for the model. Third, we evaluated the performance of the model.\n\nThe final model used for the data reduction preprocessing pipeline was trained using the Adam optimizer with a learning rate of 0.000001. We used categorical crossentropy as the loss function, MeanIoU as our tunable metric and set the batch size to 128. The model was trained for 128 epochs, although due to GPU resource limitations on Google Colab, our training crashed somewhere around epoch 75 and total amount of data used for training was 15,520 DICOM and NIfTI files, which included both images and segmentation masks. Despite the challenges, the trained U-Net model provided valuable information for the preprocessing pipeline and allowed us to effectively reduce the data to focus on the region of interest in abdominal trauma cases.\n\n## Training Results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2Fa610792c5a9897038986087883df671c%2Fdescarga.png?generation=1699395350287888&alt=media)\n\n## Inference Results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F7b77812541848956d381fbbf1416ed2e%2Fdescarga%20(2).png?generation=1699395316405025&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11304439%2F28d7f4e90e094f9dbd48760e27c4b037%2Fdescarga%20(3).png?generation=1699395391353207&alt=media)\n\n\n## 3D CNN - data preprocessing\nFor the 3D CNN model, after completing our \"standard\" preprocessing procedure and fixing the scale to 128 x 128, the question of determining the depth of the volume arose. To answer this, we referred to the work from [] and considered the GPU capacity. Ultimately, we set the volume depth to 64.\n\nTo get the data into the shape of 128x128x64, we used the zoom [] function from Python's scipy module. To automate this preprocessing task and data generation, we developed a preprocessing script that maps the patient data, storing the DICOM paths in a DataFrame (which was later saved as a CSV file and also stored in an SQLite database). From these mapped paths of the original data, we then distributed the specific \"maps\" to correctly generate the data.\n\nThe data was generated and saved as numpy files, normalized to values between 0 and 1, with the final shape of 128x128x64 for each series (patients' folders store series, with some patients having only one series and others having two series). As a reminder, the total number of training data 128x128x64 data volumes generated sums up the complete number of series present in the competition's dataset.\n\n## CNN-LSTM - data preprocessing\nOn this model we use the down-sampling block of the U-net previously trained for the Semantic Segmentation using the masks, to extract the Feature vector of the image, this part of the U-net is called Encoder, the pipeline was:\n\nSemantic Segmentation with the U-net -> Reduce Volume Shape to 128x128x64 (using Zoom Function) -> Feature Extraction (Pretrained Encoder) -> Bidirectional-LSTM\nNote: In the Feature Extraction we notice there were many 0's in the Feature Vectors, we applied a drop to that (it doesn't give us important information) and the final shape to be inputed in the model was (Batch_size, 64, 624)\n\n## Models\n* 3D CNN\n* CNN-LSTM (Pretrained Encoder from U-net -> Bi-LSTM)\n\n## What did not work\n* 3D-CNN model has bad results, we think that this kind of arquitecture only works with a very good cropping of the volumes, the bad performance is because of the black pixels who tend to overfit the model.\nCurrently working on different approaches...\n## NOTEBOOKS\n\nTraining: (The train of the model was on a local computer)\n\nInference: https://www.kaggle.com/alejandrolunamtz/inference-rsna-atd-fv"
  }
}