{
  "id": 453651,
  "title": "MLDA-G12 - Ensembling: EfficientNet_b4 + Weighted Baseline - Part 1",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/453651",
  "author_name": "RohanAkode55",
  "post_date": "2023-11-07T07:23:20.728000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<h2>CONTEXT SECTION</h2>\n<ul>\n<li><strong>Business context</strong>: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">Contest Page</a></li>\n</ul>\n<h2>- <strong>Data context</strong>:  <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">Dataset</a></h2>\n<h2>OVERVIEW OF APPROACH</h2>\n<h3>DATASET PREPERATION:</h3>\n<p>Dataset provided for the contest consists of 3 main data sources:</p>\n<ol>\n<li>Metadata for each patient</li>\n<li>Dicom or CT-SCAN images for each patient</li>\n<li>NII files or MRI scan images for each patient<br>\nOur initial aim of this project was to leverage all 3 parts together, though given that this was a contest at the <br>\nend we had to drop the usage of some data sources due to added discrepancy in result and resort to a <br>\ncombination of just 2. The same is explained in the model bellow. Before going there, let us understand the <br>\ndata provided as images:<br>\nThe dataset provided consisted of 2 types of images (CT scans or CAT scans) - .dcm files and .ni files.</li>\n</ol>\n<h4>Exploring ‘.dcm’ Files:</h4>\n<p>.dcm stands as an extension for DICOM files, an abbreviation for Digital Imaging and Communications in <br>\nMedicine. It is a set or sequence of X-ray images that CT scan is comprised of providing details on organ <br>\nhealth.<br>\n<strong>examples of DICOM images:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F7ffb464d82e78e29bca3dafe437a1532%2FDICOM1.png?generation=1699299484698499&amp;alt=media\" alt=\"Patient 10004 - record 21057 - IMG 1000.dcm\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47a42e18f96eda654081f04d8eba4cc3%2FDICOM2.png?generation=1699299601100767&amp;alt=media\" alt=\"Patient 10004 – record 21057 – IMG 1029.dcm\"><br>\nPatient 10004 – record 21057 – IMG 1000.dcm,         Patient 10004 – record 21057 – IMG 1029.dcm</p>\n<h4>Exploring ‘.ni’ Files:</h4>\n<p>The other set of data present along with the training set was .ni file. During initial procedures we tried to <br>\nleverage this data to train model as well while actual model conditions and result deviations are discussed at <br>\na later part.<br>\nThe .ni files stand as an extension for Neuroimaging Informatics Technology Initiative, the organisation that <br>\ncreated .nii file format. It is commonly used to store magnetic resonance imaging (MRI) data.<br>\nThe MRI contains a 3D representation of full middle body section for analysis of bowel structure. These are <br>\nparticularly difficult to process to the ML model due to its nature. For this problem, we decided to take an <br>\nalternate route to the problem by converting the 3D lattice into 3 sets of lateral snapshots each having 2 axes <br>\nfixed at 0 and 1 available for lateral traversal. <br>\nTo explain it in simple terms. We changed value of z while keeping x and y at 0. This produced slices parallel <br>\nto x-y plane at regular intervals in z axis from z = 0 to z = max.<br>\nA visualisation of this lateral segment can be seen in the screenshot below. The screenshot is captured on a <br>\nweb-app available to public access via this <a href=\"https://socr.umich.edu/HTML5/BrainViewer/\" target=\"_blank\">link</a><br>\n<strong>examples of NII files:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fde92fe3838a9035bcb54aa3b1ec46d33%2FNII_1.jpg?generation=1699299895059718&amp;alt=media\" alt=\"Patient 10000 NII File\"></p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fa1b52b75588e1468665c0f2c5d578695%2FNII_2.jpg?generation=1699300536859629&amp;alt=media\" alt=\"\"></h2>\n<h3>EXPLORATORY DATA ANALYSIS:</h3>\n<h4>META DATA Cleanup</h4>\n<p>The data cleaning step included checking for any missing values and removing the rows altogether in the <br>\nmetadata of train dataset. Fortunately, we found no missing values. The only value present in the meta data <br>\nfor each patient was the aortic_hu that was absolute value. <br>\nFor the purpose of a uniform distribution for the conditional model, we normalised the value of aortic_hu to <br>\nlimited from 0-1 where 0 indicated lowest observed value and 1 represented highest value observed. <br>\nThis was done using the mathematical expression:<br>\n<strong>normalised-aortic-hu</strong> = (𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)/(𝐡𝐢𝐠𝐡𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)</p>\n<h4>NII FILE ANALYSIS</h4>\n<p><strong>Redundant Full Black image cleanup</strong><br>\nThe cleaning operation was primarily to remove the full black images from the nii file generated dataset. This <br>\nwas particularly difficult when we see that the number of such images in the dataset of nii file output was <br>\nnot the same. In some MRI images, first 12 images were fully black, in others first 16 were. This created an <br>\nirregular sized set of images if we proceeded with this tactic. This required an alternate route of first finding <br>\nthe minimum such count of fully black images and then reduce that size from both sides of an image thus <br>\nreducing the dataset, although keeping it uniform in all patient MRI images and in an optimised set size <br>\nreducing irregularity and difficulty for model to process each frame.<br>\nSo to talk numerically,<br>\nWe took 100 snapshots per axis of each MRI scan or each NII file.<br>\nSo, each NII file contributed along x = 100, y = 100, z = 100 =&gt; total 300 images.<br>\nWe found minimum non-full black image at 9th position. i.e. we decided to remove 8 images from each side <br>\non all axes. So, new dataset =&gt; x = 84 (100 - 8 - 8), y = 84, z = 84. =&gt; total = 252 images per NII file<br>\nSo effective dataset reduced by 16% after removal of redundant images.<br>\n<strong>Note:</strong> Feature Extraction could not be done well due to inaccurate slice ranges and image rotation in NII files unlike DICOM files.</p>\n<h4>DICOM FILES ANALYSIS</h4>\n<p><strong>Redundant Full Black image cleanup</strong><br>\nInitial Approach of the dicom file processing was like NII file about removing redundant images, but it turned <br>\nout that nearly all of the images were significant and unskipable. This was great because now we had an <br>\noption to utilise full dataset and each individual image training was meaningful.</p>\n<h4>Feature Extraction on DICOM Images</h4>\n<p>The second exploratory analysis process that we implemented was core feature extraction from images by <br>\nlocalisation of organs. How exactly? As seen from the image below, the organs are localised to certain <br>\npositions of CT scan or alternatively DICOM images. We tried to extract specific location of each organ with <br>\na 20% buffer border around each organ to adjust for any dislocation of organ due to natural causes like <br>\ngenetics or in body fat layers. This buffer was also to account for the organ movement due to diaphragm<br>\ncompression and relaxation during breathing.<br>\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the <br>\nDICOM based ML model into its subsections. Although this was successfully executed to separate specific <br>\norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.<br>\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the <br>\nDICOM based ML model into its subsections. Although this was successfully executed to separate specific <br>\norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.<br>\nIt so happened that the considered organs overlapped over each other’s specific sub-region images. So, an <br>\nimage for kidney health analysis contained a significant part of spleen as well, so if we went forward with this <br>\nimplementation, It was possible that damage on spleen could be reflected in damage on kidney by model <br>\npredictions due to spleen sharing a significant portion in images. <br>\nThis could be solved by removing the restriction of keeping the image sub-regions to be explicitly a rectangle <br>\nin shape and keep it as a free body line, though, this would be much difficult to proceed with given that free <br>\nbody line detection, isolation and processing for this model was quite a difficult task which we could not <br>\ncomplete in time. Thus, we dropped the idea and added it in the Future possibilities/improvements section </p>\n<h2>at the end of this report.</h2>\n<h3>Validation Stratergy</h3>\n<p>For testing of data on the dataset, we used 2 methods of testing excluding the public dataset-based testing. </p>\n<ol>\n<li>In-sample testing (dataset that was a part of model training)</li>\n<li>out-sample testing (dataset that model has never seen)<br>\n• For out-sample testing, we split the dataset into an 80:20 ratio of train : test. We reserved the 20% <br>\ndataset as out-sample testing dataset. <br>\n• For in-sample testing, we used randomised selection of 25% of the training dataset (training dataset = <br>\n80% of total dataset). The specific number 25% was a result of trying to match the out-sample testing to </li>\n</ol>\n<h2>create an effective 20% total dataset for in-sample testing as well.</h2>\n<h3>ML MODEL SECTION</h3>\n<h4>MODEL LOGIC</h4>\n<p><strong>Initial Approach:</strong><br>\nAs discussed earlier we planned to utilise all the 3 types of data together. But there was a problem with this <br>\napproach. The 3 datasets showed a lot of variations. If we were to treat MRI images and CT-SCAN images as <br>\na single input to the model, we were bound to face issues with training and model accuracy plunging down. <br>\nTo solve this issue, we decided to take an ensemble model like approach to the model where we would be <br>\ntreating each dataset separately with their own model and then we would combine the generated result <br>\nfrom each model with appropriate weights to decide on the best output to be returned as a result.<br>\nAn illustration of the same can be seen as:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47fae49205d7d037c21eadfb3f126f25%2FMLDA_Model.jpg?generation=1699305429235805&amp;alt=media\" alt=\"\"><br>\nThe data had disparity for each patient and not every patient had both CT scan (aka CAT scan in US) as well <br>\nas MRI done while diagnostics. In such cases, we simply changed the weights for such patients and distributed the other weights proportionally.<br>\nFor example:<br>\nIf the decided ideal weights were: α=0.4, β=0.4, γ=0.2<br>\nIf a patient A only performed CT-SCAN and not MRI, we could simply set value of β=0 and redistribute α and γ <br>\nproportionally as α = α/(α+ γ), γ = γ/(α+ γ)<br>\nAlthough this was our initial plan, we observed that the predicted values by the NII file model had a lot of <br>\ndiscrepancy. Due to this, the idealistic β value would have been near 0. Thus, to save processing time, we eliminated </p>\n<h2>the NII file processing segment and ML model entirely in final solution keeping just DICOM and Metadata.</h2>\n<h2>DETAILS OF THE SUBMISSION</h2>\n<h3>MODEL ALGORITHMN</h3>\n<h4>ML model trained on DICOM images</h4>\n<p>To build this model, we have taken help of EfficientNet_B4 prebuilt model. <br>\nEfficientNet is a convolutional neural network architecture and scaling method that uniformly scales all <br>\ndimensions of depth/width/resolution using a compound coefficient. Unlike conventional practice that <br>\narbitrary scales these factors, the EfficientNet scaling method uniformly scales network width, depth, and <br>\nresolution with a set of fixed scaling coefficients. <br>\nFor example, if we want to use 2^N times more computational resources, then we can simply increase the <br>\nnetwork depth by a^N, width by b^N, and image size by c^N, where a,b,c are constant coefficients <br>\ndetermined by a small grid search on the original small model. EfficientNet uses a compound <br>\ncoefficient phi to uniformly scales network width, depth, and resolution in a principled way.<br>\nThe compound scaling method is justified by the intuition that if the input image is bigger, then the network <br>\nneeds more layers to increase the receptive field and more channels to capture more fine-grained patterns <br>\non the bigger image. EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 <br>\n(91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer <br>\nparameters.</p>\n<h4>Pseudo Code for DICOM - EfficientNet_b4 implementation</h4>\n<pre><code>model = create_model(, =)\ntorch.save(model.state_dict(), )\nweights_path = \ndef build_model(num_classes):\n   model = create_model(, =)\n\n    os.path.exists(weights_path):\n       model.load_state_dict(torch.load(weights_path, =), =)  \n   :\n       (f)\n\n   model.classifier = nn.Linear(model.classifier.in_features, num_classes)\n\n   return model\ndevice = torch.device(  torch.cuda.is_available()  )\nmodel = build_model(len(train_df.columns) - 1).(device)\ncriterion = torch.nn.BCEWithLogitsLoss()\noptimizer = optim.Adam(model.parameters(), =0.001)\nnum_epochs = 3\n epoch  range(num_epochs):\n   model.train()\n   running_loss = 0.0\n    i, (inputs, labels)  enumerate(train_loader):\n       (i)\n       inputs, labels = inputs.(device), labels.(device)\n       optimizer.zero_grad()\n       outputs = model(inputs)\n       loss = criterion(outputs, labels)\n       loss.backward()\n       optimizer.()\n       running_loss += loss.item()\n   (f)\n</code></pre>\n<h4>Implementation of Model (Dicom Section)</h4>\n<ol>\n<li>We begin by preprocessing our DICOM images to prepare them for the required training process. These <br>\npreprocessing steps ensure that our input data is in the optimal format for subsequent model training.</li>\n<li>Next, we construct our deep learning model using the EfficientNet-B4 architecture. This model comes <br>\nwith pre-initialized weights, providing a foundation for the training process. We configure the model to <br>\naccept our specific dataset, consisting of nine distinct classes, as its input. This step is essential for the <br>\nmodel to understand the classification task it's meant to perform.</li>\n<li>To facilitate the training and testing phases, we select the Adam optimizer. The Adam optimizer is a <br>\nwidely used optimization algorithm for deep learning tasks. It helps fine-tune the model's parameters <br>\nduring the training process.</li>\n<li>Throughout the training phase, our model carefully analyzes each image, examining every pixel and <br>\ndynamically assigning weights to various regions of the images. This process allows the model to learn <br>\nand adapt to the specific characteristics of our dataset.</li>\n<li>As the training progresses, if the model recognizes patterns or features that it has encountered before, <br>\nit associates them with the appropriate class. This continuous learning and refinement of its internal <br>\nrepresentations contribute to the model's ability to make accurate predictions.</li>\n<li>In this particular scenario, our model is trained for a single epoch. An epoch represents a complete pass <br>\nthrough the training dataset. After this initial training phase, we evaluate the model's performance on a <br>\nseparate set of test images. The calculated loss on these test images provides valuable feedback on how <br>\nwell the model is generalizing to new, unseen data.</li>\n<li>Based on the test loss, we update the model's weights using a learning rate (LR) of 0.001. This step allows <br>\nthe model to make adjustments based on the errors it encountered during testing, further improving its </li>\n</ol>\n<h2>accuracy.</h2>\n<p>Part 2 contains Weighted Baseline Model implementation, Model Reasoning, HyperParameteric Tuning and Results Section. You can find a link to the article <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/453570\" target=\"_blank\">here</a>.<br>\nThank you for reading through.</p>",
  "messages": [
    {
      "id": 2515737,
      "postDate": "2023-11-07T07:23:20.730Z",
      "content": "<h2>CONTEXT SECTION</h2>\n<ul>\n<li><strong>Business context</strong>: <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview\" target=\"_blank\">Contest Page</a></li>\n</ul>\n<h2>- <strong>Data context</strong>:  <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data\" target=\"_blank\">Dataset</a></h2>\n<h2>OVERVIEW OF APPROACH</h2>\n<h3>DATASET PREPERATION:</h3>\n<p>Dataset provided for the contest consists of 3 main data sources:</p>\n<ol>\n<li>Metadata for each patient</li>\n<li>Dicom or CT-SCAN images for each patient</li>\n<li>NII files or MRI scan images for each patient<br>\nOur initial aim of this project was to leverage all 3 parts together, though given that this was a contest at the <br>\nend we had to drop the usage of some data sources due to added discrepancy in result and resort to a <br>\ncombination of just 2. The same is explained in the model bellow. Before going there, let us understand the <br>\ndata provided as images:<br>\nThe dataset provided consisted of 2 types of images (CT scans or CAT scans) - .dcm files and .ni files.</li>\n</ol>\n<h4>Exploring ‘.dcm’ Files:</h4>\n<p>.dcm stands as an extension for DICOM files, an abbreviation for Digital Imaging and Communications in <br>\nMedicine. It is a set or sequence of X-ray images that CT scan is comprised of providing details on organ <br>\nhealth.<br>\n<strong>examples of DICOM images:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F7ffb464d82e78e29bca3dafe437a1532%2FDICOM1.png?generation=1699299484698499&amp;alt=media\" alt=\"Patient 10004 - record 21057 - IMG 1000.dcm\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47a42e18f96eda654081f04d8eba4cc3%2FDICOM2.png?generation=1699299601100767&amp;alt=media\" alt=\"Patient 10004 – record 21057 – IMG 1029.dcm\"><br>\nPatient 10004 – record 21057 – IMG 1000.dcm,         Patient 10004 – record 21057 – IMG 1029.dcm</p>\n<h4>Exploring ‘.ni’ Files:</h4>\n<p>The other set of data present along with the training set was .ni file. During initial procedures we tried to <br>\nleverage this data to train model as well while actual model conditions and result deviations are discussed at <br>\na later part.<br>\nThe .ni files stand as an extension for Neuroimaging Informatics Technology Initiative, the organisation that <br>\ncreated .nii file format. It is commonly used to store magnetic resonance imaging (MRI) data.<br>\nThe MRI contains a 3D representation of full middle body section for analysis of bowel structure. These are <br>\nparticularly difficult to process to the ML model due to its nature. For this problem, we decided to take an <br>\nalternate route to the problem by converting the 3D lattice into 3 sets of lateral snapshots each having 2 axes <br>\nfixed at 0 and 1 available for lateral traversal. <br>\nTo explain it in simple terms. We changed value of z while keeping x and y at 0. This produced slices parallel <br>\nto x-y plane at regular intervals in z axis from z = 0 to z = max.<br>\nA visualisation of this lateral segment can be seen in the screenshot below. The screenshot is captured on a <br>\nweb-app available to public access via this <a href=\"https://socr.umich.edu/HTML5/BrainViewer/\" target=\"_blank\">link</a><br>\n<strong>examples of NII files:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fde92fe3838a9035bcb54aa3b1ec46d33%2FNII_1.jpg?generation=1699299895059718&amp;alt=media\" alt=\"Patient 10000 NII File\"></p>\n<h2><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fa1b52b75588e1468665c0f2c5d578695%2FNII_2.jpg?generation=1699300536859629&amp;alt=media\" alt=\"\"></h2>\n<h3>EXPLORATORY DATA ANALYSIS:</h3>\n<h4>META DATA Cleanup</h4>\n<p>The data cleaning step included checking for any missing values and removing the rows altogether in the <br>\nmetadata of train dataset. Fortunately, we found no missing values. The only value present in the meta data <br>\nfor each patient was the aortic_hu that was absolute value. <br>\nFor the purpose of a uniform distribution for the conditional model, we normalised the value of aortic_hu to <br>\nlimited from 0-1 where 0 indicated lowest observed value and 1 represented highest value observed. <br>\nThis was done using the mathematical expression:<br>\n<strong>normalised-aortic-hu</strong> = (𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)/(𝐡𝐢𝐠𝐡𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)</p>\n<h4>NII FILE ANALYSIS</h4>\n<p><strong>Redundant Full Black image cleanup</strong><br>\nThe cleaning operation was primarily to remove the full black images from the nii file generated dataset. This <br>\nwas particularly difficult when we see that the number of such images in the dataset of nii file output was <br>\nnot the same. In some MRI images, first 12 images were fully black, in others first 16 were. This created an <br>\nirregular sized set of images if we proceeded with this tactic. This required an alternate route of first finding <br>\nthe minimum such count of fully black images and then reduce that size from both sides of an image thus <br>\nreducing the dataset, although keeping it uniform in all patient MRI images and in an optimised set size <br>\nreducing irregularity and difficulty for model to process each frame.<br>\nSo to talk numerically,<br>\nWe took 100 snapshots per axis of each MRI scan or each NII file.<br>\nSo, each NII file contributed along x = 100, y = 100, z = 100 =&gt; total 300 images.<br>\nWe found minimum non-full black image at 9th position. i.e. we decided to remove 8 images from each side <br>\non all axes. So, new dataset =&gt; x = 84 (100 - 8 - 8), y = 84, z = 84. =&gt; total = 252 images per NII file<br>\nSo effective dataset reduced by 16% after removal of redundant images.<br>\n<strong>Note:</strong> Feature Extraction could not be done well due to inaccurate slice ranges and image rotation in NII files unlike DICOM files.</p>\n<h4>DICOM FILES ANALYSIS</h4>\n<p><strong>Redundant Full Black image cleanup</strong><br>\nInitial Approach of the dicom file processing was like NII file about removing redundant images, but it turned <br>\nout that nearly all of the images were significant and unskipable. This was great because now we had an <br>\noption to utilise full dataset and each individual image training was meaningful.</p>\n<h4>Feature Extraction on DICOM Images</h4>\n<p>The second exploratory analysis process that we implemented was core feature extraction from images by <br>\nlocalisation of organs. How exactly? As seen from the image below, the organs are localised to certain <br>\npositions of CT scan or alternatively DICOM images. We tried to extract specific location of each organ with <br>\na 20% buffer border around each organ to adjust for any dislocation of organ due to natural causes like <br>\ngenetics or in body fat layers. This buffer was also to account for the organ movement due to diaphragm<br>\ncompression and relaxation during breathing.<br>\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the <br>\nDICOM based ML model into its subsections. Although this was successfully executed to separate specific <br>\norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.<br>\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the <br>\nDICOM based ML model into its subsections. Although this was successfully executed to separate specific <br>\norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.<br>\nIt so happened that the considered organs overlapped over each other’s specific sub-region images. So, an <br>\nimage for kidney health analysis contained a significant part of spleen as well, so if we went forward with this <br>\nimplementation, It was possible that damage on spleen could be reflected in damage on kidney by model <br>\npredictions due to spleen sharing a significant portion in images. <br>\nThis could be solved by removing the restriction of keeping the image sub-regions to be explicitly a rectangle <br>\nin shape and keep it as a free body line, though, this would be much difficult to proceed with given that free <br>\nbody line detection, isolation and processing for this model was quite a difficult task which we could not <br>\ncomplete in time. Thus, we dropped the idea and added it in the Future possibilities/improvements section </p>\n<h2>at the end of this report.</h2>\n<h3>Validation Stratergy</h3>\n<p>For testing of data on the dataset, we used 2 methods of testing excluding the public dataset-based testing. </p>\n<ol>\n<li>In-sample testing (dataset that was a part of model training)</li>\n<li>out-sample testing (dataset that model has never seen)<br>\n• For out-sample testing, we split the dataset into an 80:20 ratio of train : test. We reserved the 20% <br>\ndataset as out-sample testing dataset. <br>\n• For in-sample testing, we used randomised selection of 25% of the training dataset (training dataset = <br>\n80% of total dataset). The specific number 25% was a result of trying to match the out-sample testing to </li>\n</ol>\n<h2>create an effective 20% total dataset for in-sample testing as well.</h2>\n<h3>ML MODEL SECTION</h3>\n<h4>MODEL LOGIC</h4>\n<p><strong>Initial Approach:</strong><br>\nAs discussed earlier we planned to utilise all the 3 types of data together. But there was a problem with this <br>\napproach. The 3 datasets showed a lot of variations. If we were to treat MRI images and CT-SCAN images as <br>\na single input to the model, we were bound to face issues with training and model accuracy plunging down. <br>\nTo solve this issue, we decided to take an ensemble model like approach to the model where we would be <br>\ntreating each dataset separately with their own model and then we would combine the generated result <br>\nfrom each model with appropriate weights to decide on the best output to be returned as a result.<br>\nAn illustration of the same can be seen as:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47fae49205d7d037c21eadfb3f126f25%2FMLDA_Model.jpg?generation=1699305429235805&amp;alt=media\" alt=\"\"><br>\nThe data had disparity for each patient and not every patient had both CT scan (aka CAT scan in US) as well <br>\nas MRI done while diagnostics. In such cases, we simply changed the weights for such patients and distributed the other weights proportionally.<br>\nFor example:<br>\nIf the decided ideal weights were: α=0.4, β=0.4, γ=0.2<br>\nIf a patient A only performed CT-SCAN and not MRI, we could simply set value of β=0 and redistribute α and γ <br>\nproportionally as α = α/(α+ γ), γ = γ/(α+ γ)<br>\nAlthough this was our initial plan, we observed that the predicted values by the NII file model had a lot of <br>\ndiscrepancy. Due to this, the idealistic β value would have been near 0. Thus, to save processing time, we eliminated </p>\n<h2>the NII file processing segment and ML model entirely in final solution keeping just DICOM and Metadata.</h2>\n<h2>DETAILS OF THE SUBMISSION</h2>\n<h3>MODEL ALGORITHMN</h3>\n<h4>ML model trained on DICOM images</h4>\n<p>To build this model, we have taken help of EfficientNet_B4 prebuilt model. <br>\nEfficientNet is a convolutional neural network architecture and scaling method that uniformly scales all <br>\ndimensions of depth/width/resolution using a compound coefficient. Unlike conventional practice that <br>\narbitrary scales these factors, the EfficientNet scaling method uniformly scales network width, depth, and <br>\nresolution with a set of fixed scaling coefficients. <br>\nFor example, if we want to use 2^N times more computational resources, then we can simply increase the <br>\nnetwork depth by a^N, width by b^N, and image size by c^N, where a,b,c are constant coefficients <br>\ndetermined by a small grid search on the original small model. EfficientNet uses a compound <br>\ncoefficient phi to uniformly scales network width, depth, and resolution in a principled way.<br>\nThe compound scaling method is justified by the intuition that if the input image is bigger, then the network <br>\nneeds more layers to increase the receptive field and more channels to capture more fine-grained patterns <br>\non the bigger image. EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 <br>\n(91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer <br>\nparameters.</p>\n<h4>Pseudo Code for DICOM - EfficientNet_b4 implementation</h4>\n<pre><code>model = create_model(, =)\ntorch.save(model.state_dict(), )\nweights_path = \ndef build_model(num_classes):\n   model = create_model(, =)\n\n    os.path.exists(weights_path):\n       model.load_state_dict(torch.load(weights_path, =), =)  \n   :\n       (f)\n\n   model.classifier = nn.Linear(model.classifier.in_features, num_classes)\n\n   return model\ndevice = torch.device(  torch.cuda.is_available()  )\nmodel = build_model(len(train_df.columns) - 1).(device)\ncriterion = torch.nn.BCEWithLogitsLoss()\noptimizer = optim.Adam(model.parameters(), =0.001)\nnum_epochs = 3\n epoch  range(num_epochs):\n   model.train()\n   running_loss = 0.0\n    i, (inputs, labels)  enumerate(train_loader):\n       (i)\n       inputs, labels = inputs.(device), labels.(device)\n       optimizer.zero_grad()\n       outputs = model(inputs)\n       loss = criterion(outputs, labels)\n       loss.backward()\n       optimizer.()\n       running_loss += loss.item()\n   (f)\n</code></pre>\n<h4>Implementation of Model (Dicom Section)</h4>\n<ol>\n<li>We begin by preprocessing our DICOM images to prepare them for the required training process. These <br>\npreprocessing steps ensure that our input data is in the optimal format for subsequent model training.</li>\n<li>Next, we construct our deep learning model using the EfficientNet-B4 architecture. This model comes <br>\nwith pre-initialized weights, providing a foundation for the training process. We configure the model to <br>\naccept our specific dataset, consisting of nine distinct classes, as its input. This step is essential for the <br>\nmodel to understand the classification task it's meant to perform.</li>\n<li>To facilitate the training and testing phases, we select the Adam optimizer. The Adam optimizer is a <br>\nwidely used optimization algorithm for deep learning tasks. It helps fine-tune the model's parameters <br>\nduring the training process.</li>\n<li>Throughout the training phase, our model carefully analyzes each image, examining every pixel and <br>\ndynamically assigning weights to various regions of the images. This process allows the model to learn <br>\nand adapt to the specific characteristics of our dataset.</li>\n<li>As the training progresses, if the model recognizes patterns or features that it has encountered before, <br>\nit associates them with the appropriate class. This continuous learning and refinement of its internal <br>\nrepresentations contribute to the model's ability to make accurate predictions.</li>\n<li>In this particular scenario, our model is trained for a single epoch. An epoch represents a complete pass <br>\nthrough the training dataset. After this initial training phase, we evaluate the model's performance on a <br>\nseparate set of test images. The calculated loss on these test images provides valuable feedback on how <br>\nwell the model is generalizing to new, unseen data.</li>\n<li>Based on the test loss, we update the model's weights using a learning rate (LR) of 0.001. This step allows <br>\nthe model to make adjustments based on the errors it encountered during testing, further improving its </li>\n</ol>\n<h2>accuracy.</h2>\n<p>Part 2 contains Weighted Baseline Model implementation, Model Reasoning, HyperParameteric Tuning and Results Section. You can find a link to the article <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/453570\" target=\"_blank\">here</a>.<br>\nThank you for reading through.</p>",
      "rawMarkdown": "\n## CONTEXT SECTION\n\n- **Business context**: [Contest Page](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview)\n- **Data context**:  [Dataset](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data)\n\n---\n\n## OVERVIEW OF APPROACH\n\n### DATASET PREPERATION:\n\nDataset provided for the contest consists of 3 main data sources:\n1. Metadata for each patient\n2. Dicom or CT-SCAN images for each patient\n3. NII files or MRI scan images for each patient\n\nOur initial aim of this project was to leverage all 3 parts together, though given that this was a contest at the \nend we had to drop the usage of some data sources due to added discrepancy in result and resort to a \ncombination of just 2. The same is explained in the model bellow. Before going there, let us understand the \ndata provided as images:\n\nThe dataset provided consisted of 2 types of images (CT scans or CAT scans) - .dcm files and .ni files.\n\n#### Exploring ‘.dcm’ Files:\n\n.dcm stands as an extension for DICOM files, an abbreviation for Digital Imaging and Communications in \nMedicine. It is a set or sequence of X-ray images that CT scan is comprised of providing details on organ \nhealth.\n\n**examples of DICOM images:**\n![Patient 10004 - record 21057 - IMG 1000.dcm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F7ffb464d82e78e29bca3dafe437a1532%2FDICOM1.png?generation=1699299484698499&alt=media) ![Patient 10004 – record 21057 – IMG 1029.dcm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47a42e18f96eda654081f04d8eba4cc3%2FDICOM2.png?generation=1699299601100767&alt=media)\nPatient 10004 – record 21057 – IMG 1000.dcm,         Patient 10004 – record 21057 – IMG 1029.dcm\n\n#### Exploring ‘.ni’ Files:\n\nThe other set of data present along with the training set was .ni file. During initial procedures we tried to \nleverage this data to train model as well while actual model conditions and result deviations are discussed at \na later part.\n\nThe .ni files stand as an extension for Neuroimaging Informatics Technology Initiative, the organisation that \ncreated .nii file format. It is commonly used to store magnetic resonance imaging (MRI) data.\nThe MRI contains a 3D representation of full middle body section for analysis of bowel structure. These are \nparticularly difficult to process to the ML model due to its nature. For this problem, we decided to take an \nalternate route to the problem by converting the 3D lattice into 3 sets of lateral snapshots each having 2 axes \nfixed at 0 and 1 available for lateral traversal. \n\nTo explain it in simple terms. We changed value of z while keeping x and y at 0. This produced slices parallel \nto x-y plane at regular intervals in z axis from z = 0 to z = max.\n\nA visualisation of this lateral segment can be seen in the screenshot below. The screenshot is captured on a \nweb-app available to public access via this [link](https://socr.umich.edu/HTML5/BrainViewer/)\n\n**examples of NII files:**\n![Patient 10000 NII File](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fde92fe3838a9035bcb54aa3b1ec46d33%2FNII_1.jpg?generation=1699299895059718&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fa1b52b75588e1468665c0f2c5d578695%2FNII_2.jpg?generation=1699300536859629&alt=media)\n\n---\n### EXPLORATORY DATA ANALYSIS:\n\n#### META DATA Cleanup\n\nThe data cleaning step included checking for any missing values and removing the rows altogether in the \nmetadata of train dataset. Fortunately, we found no missing values. The only value present in the meta data \nfor each patient was the aortic_hu that was absolute value. \nFor the purpose of a uniform distribution for the conditional model, we normalised the value of aortic_hu to \nlimited from 0-1 where 0 indicated lowest observed value and 1 represented highest value observed. \n\nThis was done using the mathematical expression:\n**normalised-aortic-hu** = (𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)/(𝐡𝐢𝐠𝐡𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)\n\n#### NII FILE ANALYSIS\n\n**Redundant Full Black image cleanup**\nThe cleaning operation was primarily to remove the full black images from the nii file generated dataset. This \nwas particularly difficult when we see that the number of such images in the dataset of nii file output was \nnot the same. In some MRI images, first 12 images were fully black, in others first 16 were. This created an \nirregular sized set of images if we proceeded with this tactic. This required an alternate route of first finding \nthe minimum such count of fully black images and then reduce that size from both sides of an image thus \nreducing the dataset, although keeping it uniform in all patient MRI images and in an optimised set size \nreducing irregularity and difficulty for model to process each frame.\n\nSo to talk numerically,\nWe took 100 snapshots per axis of each MRI scan or each NII file.\nSo, each NII file contributed along x = 100, y = 100, z = 100 => total 300 images.\nWe found minimum non-full black image at 9th position. i.e. we decided to remove 8 images from each side \non all axes. So, new dataset => x = 84 (100 - 8 - 8), y = 84, z = 84. => total = 252 images per NII file\nSo effective dataset reduced by 16% after removal of redundant images.\n\n**Note:** Feature Extraction could not be done well due to inaccurate slice ranges and image rotation in NII files unlike DICOM files.\n\n#### DICOM FILES ANALYSIS\n**Redundant Full Black image cleanup**\nInitial Approach of the dicom file processing was like NII file about removing redundant images, but it turned \nout that nearly all of the images were significant and unskipable. This was great because now we had an \noption to utilise full dataset and each individual image training was meaningful.\n\n#### Feature Extraction on DICOM Images\nThe second exploratory analysis process that we implemented was core feature extraction from images by \nlocalisation of organs. How exactly? As seen from the image below, the organs are localised to certain \npositions of CT scan or alternatively DICOM images. We tried to extract specific location of each organ with \na 20% buffer border around each organ to adjust for any dislocation of organ due to natural causes like \ngenetics or in body fat layers. This buffer was also to account for the organ movement due to diaphragm\ncompression and relaxation during breathing.\n\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the \nDICOM based ML model into its subsections. Although this was successfully executed to separate specific \norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.\n\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the \nDICOM based ML model into its subsections. Although this was successfully executed to separate specific \norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.\n\nIt so happened that the considered organs overlapped over each other’s specific sub-region images. So, an \nimage for kidney health analysis contained a significant part of spleen as well, so if we went forward with this \nimplementation, It was possible that damage on spleen could be reflected in damage on kidney by model \npredictions due to spleen sharing a significant portion in images. \n\nThis could be solved by removing the restriction of keeping the image sub-regions to be explicitly a rectangle \nin shape and keep it as a free body line, though, this would be much difficult to proceed with given that free \nbody line detection, isolation and processing for this model was quite a difficult task which we could not \ncomplete in time. Thus, we dropped the idea and added it in the Future possibilities/improvements section \nat the end of this report.\n\n---\n\n### Validation Stratergy\n\nFor testing of data on the dataset, we used 2 methods of testing excluding the public dataset-based testing. \n1. In-sample testing (dataset that was a part of model training)\n2. out-sample testing (dataset that model has never seen)\n• For out-sample testing, we split the dataset into an 80:20 ratio of train : test. We reserved the 20% \ndataset as out-sample testing dataset. \n• For in-sample testing, we used randomised selection of 25% of the training dataset (training dataset = \n80% of total dataset). The specific number 25% was a result of trying to match the out-sample testing to \ncreate an effective 20% total dataset for in-sample testing as well.\n\n---\n\n### ML MODEL SECTION\n\n#### MODEL LOGIC\n\n**Initial Approach:**\nAs discussed earlier we planned to utilise all the 3 types of data together. But there was a problem with this \napproach. The 3 datasets showed a lot of variations. If we were to treat MRI images and CT-SCAN images as \na single input to the model, we were bound to face issues with training and model accuracy plunging down. \nTo solve this issue, we decided to take an ensemble model like approach to the model where we would be \ntreating each dataset separately with their own model and then we would combine the generated result \nfrom each model with appropriate weights to decide on the best output to be returned as a result.\n\nAn illustration of the same can be seen as:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47fae49205d7d037c21eadfb3f126f25%2FMLDA_Model.jpg?generation=1699305429235805&alt=media)\n\nThe data had disparity for each patient and not every patient had both CT scan (aka CAT scan in US) as well \nas MRI done while diagnostics. In such cases, we simply changed the weights for such patients and distributed the other weights proportionally.\n\nFor example:\nIf the decided ideal weights were: α=0.4, β=0.4, γ=0.2\nIf a patient A only performed CT-SCAN and not MRI, we could simply set value of β=0 and redistribute α and γ \nproportionally as α = α/(α+ γ), γ = γ/(α+ γ)\nAlthough this was our initial plan, we observed that the predicted values by the NII file model had a lot of \ndiscrepancy. Due to this, the idealistic β value would have been near 0. Thus, to save processing time, we eliminated \nthe NII file processing segment and ML model entirely in final solution keeping just DICOM and Metadata.\n\n---\n\n## DETAILS OF THE SUBMISSION\n\n### MODEL ALGORITHMN\n\n#### ML model trained on DICOM images\nTo build this model, we have taken help of EfficientNet_B4 prebuilt model. \n\nEfficientNet is a convolutional neural network architecture and scaling method that uniformly scales all \ndimensions of depth/width/resolution using a compound coefficient. Unlike conventional practice that \narbitrary scales these factors, the EfficientNet scaling method uniformly scales network width, depth, and \nresolution with a set of fixed scaling coefficients. \n\nFor example, if we want to use 2^N times more computational resources, then we can simply increase the \nnetwork depth by a^N, width by b^N, and image size by c^N, where a,b,c are constant coefficients \ndetermined by a small grid search on the original small model. EfficientNet uses a compound \ncoefficient phi to uniformly scales network width, depth, and resolution in a principled way.\n\nThe compound scaling method is justified by the intuition that if the input image is bigger, then the network \nneeds more layers to increase the receptive field and more channels to capture more fine-grained patterns \non the bigger image. EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 \n(91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer \nparameters.\n\n#### Pseudo Code for DICOM - EfficientNet_b4 implementation\n```\nmodel = create_model('efficientnet_b4', pretrained=False)\ntorch.save(model.state_dict(), 'efficientnet_b4_weights.pth')\n\nweights_path = 'efficientnet_b4_weights.pth'\n\ndef build_model(num_classes):\n    model = create_model('efficientnet_b4', pretrained=False)\n    \n    if os.path.exists(weights_path):\n        model.load_state_dict(torch.load(weights_path, map_location='cpu'), strict=False)  \n    else:\n        print(f'Warning: Weights file not found in path {weights_path}, training from scratch.')\n    \n    model.classifier = nn.Linear(model.classifier.in_features, num_classes)\n    \n    return model\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel = build_model(len(train_df.columns) - 1).to(device)\n\ncriterion = torch.nn.BCEWithLogitsLoss()\noptimizer = optim.Adam(model.parameters(), lr=0.001)\n\nnum_epochs = 3\n\nfor epoch in range(num_epochs):\n    model.train()\n    running_loss = 0.0\n    for i, (inputs, labels) in enumerate(train_loader):\n        print(i)\n        inputs, labels = inputs.to(device), labels.to(device)\n        optimizer.zero_grad()\n        outputs = model(inputs)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n\n        running_loss += loss.item()\n\n    print(f\"Epoch {epoch+1}, Loss: {running_loss/len(train_loader)}\")\n```\n\n#### Implementation of Model (Dicom Section)\n\n1. We begin by preprocessing our DICOM images to prepare them for the required training process. These \npreprocessing steps ensure that our input data is in the optimal format for subsequent model training.\n2. Next, we construct our deep learning model using the EfficientNet-B4 architecture. This model comes \nwith pre-initialized weights, providing a foundation for the training process. We configure the model to \naccept our specific dataset, consisting of nine distinct classes, as its input. This step is essential for the \nmodel to understand the classification task it's meant to perform.\n3. To facilitate the training and testing phases, we select the Adam optimizer. The Adam optimizer is a \nwidely used optimization algorithm for deep learning tasks. It helps fine-tune the model's parameters \nduring the training process.\n4. Throughout the training phase, our model carefully analyzes each image, examining every pixel and \ndynamically assigning weights to various regions of the images. This process allows the model to learn \nand adapt to the specific characteristics of our dataset.\n5. As the training progresses, if the model recognizes patterns or features that it has encountered before, \nit associates them with the appropriate class. This continuous learning and refinement of its internal \nrepresentations contribute to the model's ability to make accurate predictions.\n6. In this particular scenario, our model is trained for a single epoch. An epoch represents a complete pass \nthrough the training dataset. After this initial training phase, we evaluate the model's performance on a \nseparate set of test images. The calculated loss on these test images provides valuable feedback on how \nwell the model is generalizing to new, unseen data.\n7. Based on the test loss, we update the model's weights using a learning rate (LR) of 0.001. This step allows \nthe model to make adjustments based on the errors it encountered during testing, further improving its \naccuracy.\n\n---\n\nPart 2 contains Weighted Baseline Model implementation, Model Reasoning, HyperParameteric Tuning and Results Section. You can find a link to the article [here](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/453570).\n\nThank you for reading through.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2515737": "\n## CONTEXT SECTION\n\n- **Business context**: [Contest Page](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/overview)\n- **Data context**:  [Dataset](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/data)\n\n---\n\n## OVERVIEW OF APPROACH\n\n### DATASET PREPERATION:\n\nDataset provided for the contest consists of 3 main data sources:\n1. Metadata for each patient\n2. Dicom or CT-SCAN images for each patient\n3. NII files or MRI scan images for each patient\n\nOur initial aim of this project was to leverage all 3 parts together, though given that this was a contest at the \nend we had to drop the usage of some data sources due to added discrepancy in result and resort to a \ncombination of just 2. The same is explained in the model bellow. Before going there, let us understand the \ndata provided as images:\n\nThe dataset provided consisted of 2 types of images (CT scans or CAT scans) - .dcm files and .ni files.\n\n#### Exploring ‘.dcm’ Files:\n\n.dcm stands as an extension for DICOM files, an abbreviation for Digital Imaging and Communications in \nMedicine. It is a set or sequence of X-ray images that CT scan is comprised of providing details on organ \nhealth.\n\n**examples of DICOM images:**\n![Patient 10004 - record 21057 - IMG 1000.dcm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F7ffb464d82e78e29bca3dafe437a1532%2FDICOM1.png?generation=1699299484698499&alt=media) ![Patient 10004 – record 21057 – IMG 1029.dcm](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47a42e18f96eda654081f04d8eba4cc3%2FDICOM2.png?generation=1699299601100767&alt=media)\nPatient 10004 – record 21057 – IMG 1000.dcm,         Patient 10004 – record 21057 – IMG 1029.dcm\n\n#### Exploring ‘.ni’ Files:\n\nThe other set of data present along with the training set was .ni file. During initial procedures we tried to \nleverage this data to train model as well while actual model conditions and result deviations are discussed at \na later part.\n\nThe .ni files stand as an extension for Neuroimaging Informatics Technology Initiative, the organisation that \ncreated .nii file format. It is commonly used to store magnetic resonance imaging (MRI) data.\nThe MRI contains a 3D representation of full middle body section for analysis of bowel structure. These are \nparticularly difficult to process to the ML model due to its nature. For this problem, we decided to take an \nalternate route to the problem by converting the 3D lattice into 3 sets of lateral snapshots each having 2 axes \nfixed at 0 and 1 available for lateral traversal. \n\nTo explain it in simple terms. We changed value of z while keeping x and y at 0. This produced slices parallel \nto x-y plane at regular intervals in z axis from z = 0 to z = max.\n\nA visualisation of this lateral segment can be seen in the screenshot below. The screenshot is captured on a \nweb-app available to public access via this [link](https://socr.umich.edu/HTML5/BrainViewer/)\n\n**examples of NII files:**\n![Patient 10000 NII File](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fde92fe3838a9035bcb54aa3b1ec46d33%2FNII_1.jpg?generation=1699299895059718&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2Fa1b52b75588e1468665c0f2c5d578695%2FNII_2.jpg?generation=1699300536859629&alt=media)\n\n---\n### EXPLORATORY DATA ANALYSIS:\n\n#### META DATA Cleanup\n\nThe data cleaning step included checking for any missing values and removing the rows altogether in the \nmetadata of train dataset. Fortunately, we found no missing values. The only value present in the meta data \nfor each patient was the aortic_hu that was absolute value. \nFor the purpose of a uniform distribution for the conditional model, we normalised the value of aortic_hu to \nlimited from 0-1 where 0 indicated lowest observed value and 1 represented highest value observed. \n\nThis was done using the mathematical expression:\n**normalised-aortic-hu** = (𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)/(𝐡𝐢𝐠𝐡𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞 − 𝐥𝐨𝐰𝐞𝐬𝐭-𝐚𝐨𝐫𝐭𝐢𝐜-𝐡𝐮𝐞)\n\n#### NII FILE ANALYSIS\n\n**Redundant Full Black image cleanup**\nThe cleaning operation was primarily to remove the full black images from the nii file generated dataset. This \nwas particularly difficult when we see that the number of such images in the dataset of nii file output was \nnot the same. In some MRI images, first 12 images were fully black, in others first 16 were. This created an \nirregular sized set of images if we proceeded with this tactic. This required an alternate route of first finding \nthe minimum such count of fully black images and then reduce that size from both sides of an image thus \nreducing the dataset, although keeping it uniform in all patient MRI images and in an optimised set size \nreducing irregularity and difficulty for model to process each frame.\n\nSo to talk numerically,\nWe took 100 snapshots per axis of each MRI scan or each NII file.\nSo, each NII file contributed along x = 100, y = 100, z = 100 => total 300 images.\nWe found minimum non-full black image at 9th position. i.e. we decided to remove 8 images from each side \non all axes. So, new dataset => x = 84 (100 - 8 - 8), y = 84, z = 84. => total = 252 images per NII file\nSo effective dataset reduced by 16% after removal of redundant images.\n\n**Note:** Feature Extraction could not be done well due to inaccurate slice ranges and image rotation in NII files unlike DICOM files.\n\n#### DICOM FILES ANALYSIS\n**Redundant Full Black image cleanup**\nInitial Approach of the dicom file processing was like NII file about removing redundant images, but it turned \nout that nearly all of the images were significant and unskipable. This was great because now we had an \noption to utilise full dataset and each individual image training was meaningful.\n\n#### Feature Extraction on DICOM Images\nThe second exploratory analysis process that we implemented was core feature extraction from images by \nlocalisation of organs. How exactly? As seen from the image below, the organs are localised to certain \npositions of CT scan or alternatively DICOM images. We tried to extract specific location of each organ with \na 20% buffer border around each organ to adjust for any dislocation of organ due to natural causes like \ngenetics or in body fat layers. This buffer was also to account for the organ movement due to diaphragm\ncompression and relaxation during breathing.\n\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the \nDICOM based ML model into its subsections. Although this was successfully executed to separate specific \norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.\n\nWe tried to localise this subsection of each organ and train it individually for each organ health splitting the \nDICOM based ML model into its subsections. Although this was successfully executed to separate specific \norgans from images separately with 20 buffers in each axis (10% on each border), there was still overlap.\n\nIt so happened that the considered organs overlapped over each other’s specific sub-region images. So, an \nimage for kidney health analysis contained a significant part of spleen as well, so if we went forward with this \nimplementation, It was possible that damage on spleen could be reflected in damage on kidney by model \npredictions due to spleen sharing a significant portion in images. \n\nThis could be solved by removing the restriction of keeping the image sub-regions to be explicitly a rectangle \nin shape and keep it as a free body line, though, this would be much difficult to proceed with given that free \nbody line detection, isolation and processing for this model was quite a difficult task which we could not \ncomplete in time. Thus, we dropped the idea and added it in the Future possibilities/improvements section \nat the end of this report.\n\n---\n\n### Validation Stratergy\n\nFor testing of data on the dataset, we used 2 methods of testing excluding the public dataset-based testing. \n1. In-sample testing (dataset that was a part of model training)\n2. out-sample testing (dataset that model has never seen)\n• For out-sample testing, we split the dataset into an 80:20 ratio of train : test. We reserved the 20% \ndataset as out-sample testing dataset. \n• For in-sample testing, we used randomised selection of 25% of the training dataset (training dataset = \n80% of total dataset). The specific number 25% was a result of trying to match the out-sample testing to \ncreate an effective 20% total dataset for in-sample testing as well.\n\n---\n\n### ML MODEL SECTION\n\n#### MODEL LOGIC\n\n**Initial Approach:**\nAs discussed earlier we planned to utilise all the 3 types of data together. But there was a problem with this \napproach. The 3 datasets showed a lot of variations. If we were to treat MRI images and CT-SCAN images as \na single input to the model, we were bound to face issues with training and model accuracy plunging down. \nTo solve this issue, we decided to take an ensemble model like approach to the model where we would be \ntreating each dataset separately with their own model and then we would combine the generated result \nfrom each model with appropriate weights to decide on the best output to be returned as a result.\n\nAn illustration of the same can be seen as:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16972845%2F47fae49205d7d037c21eadfb3f126f25%2FMLDA_Model.jpg?generation=1699305429235805&alt=media)\n\nThe data had disparity for each patient and not every patient had both CT scan (aka CAT scan in US) as well \nas MRI done while diagnostics. In such cases, we simply changed the weights for such patients and distributed the other weights proportionally.\n\nFor example:\nIf the decided ideal weights were: α=0.4, β=0.4, γ=0.2\nIf a patient A only performed CT-SCAN and not MRI, we could simply set value of β=0 and redistribute α and γ \nproportionally as α = α/(α+ γ), γ = γ/(α+ γ)\nAlthough this was our initial plan, we observed that the predicted values by the NII file model had a lot of \ndiscrepancy. Due to this, the idealistic β value would have been near 0. Thus, to save processing time, we eliminated \nthe NII file processing segment and ML model entirely in final solution keeping just DICOM and Metadata.\n\n---\n\n## DETAILS OF THE SUBMISSION\n\n### MODEL ALGORITHMN\n\n#### ML model trained on DICOM images\nTo build this model, we have taken help of EfficientNet_B4 prebuilt model. \n\nEfficientNet is a convolutional neural network architecture and scaling method that uniformly scales all \ndimensions of depth/width/resolution using a compound coefficient. Unlike conventional practice that \narbitrary scales these factors, the EfficientNet scaling method uniformly scales network width, depth, and \nresolution with a set of fixed scaling coefficients. \n\nFor example, if we want to use 2^N times more computational resources, then we can simply increase the \nnetwork depth by a^N, width by b^N, and image size by c^N, where a,b,c are constant coefficients \ndetermined by a small grid search on the original small model. EfficientNet uses a compound \ncoefficient phi to uniformly scales network width, depth, and resolution in a principled way.\n\nThe compound scaling method is justified by the intuition that if the input image is bigger, then the network \nneeds more layers to increase the receptive field and more channels to capture more fine-grained patterns \non the bigger image. EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 \n(91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer \nparameters.\n\n#### Pseudo Code for DICOM - EfficientNet_b4 implementation\n```\nmodel = create_model('efficientnet_b4', pretrained=False)\ntorch.save(model.state_dict(), 'efficientnet_b4_weights.pth')\n\nweights_path = 'efficientnet_b4_weights.pth'\n\ndef build_model(num_classes):\n    model = create_model('efficientnet_b4', pretrained=False)\n    \n    if os.path.exists(weights_path):\n        model.load_state_dict(torch.load(weights_path, map_location='cpu'), strict=False)  \n    else:\n        print(f'Warning: Weights file not found in path {weights_path}, training from scratch.')\n    \n    model.classifier = nn.Linear(model.classifier.in_features, num_classes)\n    \n    return model\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel = build_model(len(train_df.columns) - 1).to(device)\n\ncriterion = torch.nn.BCEWithLogitsLoss()\noptimizer = optim.Adam(model.parameters(), lr=0.001)\n\nnum_epochs = 3\n\nfor epoch in range(num_epochs):\n    model.train()\n    running_loss = 0.0\n    for i, (inputs, labels) in enumerate(train_loader):\n        print(i)\n        inputs, labels = inputs.to(device), labels.to(device)\n        optimizer.zero_grad()\n        outputs = model(inputs)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n\n        running_loss += loss.item()\n\n    print(f\"Epoch {epoch+1}, Loss: {running_loss/len(train_loader)}\")\n```\n\n#### Implementation of Model (Dicom Section)\n\n1. We begin by preprocessing our DICOM images to prepare them for the required training process. These \npreprocessing steps ensure that our input data is in the optimal format for subsequent model training.\n2. Next, we construct our deep learning model using the EfficientNet-B4 architecture. This model comes \nwith pre-initialized weights, providing a foundation for the training process. We configure the model to \naccept our specific dataset, consisting of nine distinct classes, as its input. This step is essential for the \nmodel to understand the classification task it's meant to perform.\n3. To facilitate the training and testing phases, we select the Adam optimizer. The Adam optimizer is a \nwidely used optimization algorithm for deep learning tasks. It helps fine-tune the model's parameters \nduring the training process.\n4. Throughout the training phase, our model carefully analyzes each image, examining every pixel and \ndynamically assigning weights to various regions of the images. This process allows the model to learn \nand adapt to the specific characteristics of our dataset.\n5. As the training progresses, if the model recognizes patterns or features that it has encountered before, \nit associates them with the appropriate class. This continuous learning and refinement of its internal \nrepresentations contribute to the model's ability to make accurate predictions.\n6. In this particular scenario, our model is trained for a single epoch. An epoch represents a complete pass \nthrough the training dataset. After this initial training phase, we evaluate the model's performance on a \nseparate set of test images. The calculated loss on these test images provides valuable feedback on how \nwell the model is generalizing to new, unseen data.\n7. Based on the test loss, we update the model's weights using a learning rate (LR) of 0.001. This step allows \nthe model to make adjustments based on the errors it encountered during testing, further improving its \naccuracy.\n\n---\n\nPart 2 contains Weighted Baseline Model implementation, Model Reasoning, HyperParameteric Tuning and Results Section. You can find a link to the article [here](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/453570).\n\nThank you for reading through."
  }
}