{
  "id": 358773,
  "title": "The power of pretraining on your own task-oriented images set",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/358773",
  "author_name": "José Miguel Máiz",
  "post_date": "2022-10-09T11:53:48.840000",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The purpose of this post is to show the power of pretraining a simple custom classifier on an image set linked to your current task instead of using a much more sophisticated model (such as ResNet or similar) pretrained on a set of images nothing to do with your specific task (being the task here to classify the blood clot origins in ischemic stroke as defined in the Mayo-Clinic - STRIP AI competition description).</p>\n<p>The training of the learner is based on multilabel stratified K-Folds cross-validation and makes heavy use of very simple randomization functions to augment the image set in every cross-validation loop (you don't really need albumentations for this).</p>\n<p>Using the aforementioned techniques I got an accuracy of 0.99 on a random test set of 79 images coming from the \"train\" set (I put them apart before training so that they weren't used for training purposes but only to test the model performance).</p>\n<p>I must stress that deleting blanks in the original images before downsampling them didn't improve the results of the classifier, showing that the only you need is to downsample to 512x512 size as is doesn't make you incur significant data loss, so that you can simplify your code and save a lot of preprocessing time.</p>\n<p>Summarizing, there are 2 stages in the training process:<br>\n<strong>1. Multi class classification</strong><br>\nThe model pretrains on an augmented set of images from both \"train\" and \"other\" sets provided in the competition.<br>\n<strong>2. Binary classification</strong><br>\nThe model turns in a binary classifier and fine-tunes the weights of the \"classifier\" (full connected modules) on the top of the convolutional part of the previously trained learner.</p>\n<p>This is the link to the notebook: <a href=\"https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706\" target=\"_blank\">https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706</a>. I hope you find it useful.</p>\n<p>Please feel free to give me your insights about the techniques shown here and your comments about your own experience using them.</p>\n<p>Thanks for taking your time to read this post.</p>",
  "messages": [
    {
      "id": 1979393,
      "postDate": "2022-10-09T11:53:48.840Z",
      "content": "<p>The purpose of this post is to show the power of pretraining a simple custom classifier on an image set linked to your current task instead of using a much more sophisticated model (such as ResNet or similar) pretrained on a set of images nothing to do with your specific task (being the task here to classify the blood clot origins in ischemic stroke as defined in the Mayo-Clinic - STRIP AI competition description).</p>\n<p>The training of the learner is based on multilabel stratified K-Folds cross-validation and makes heavy use of very simple randomization functions to augment the image set in every cross-validation loop (you don't really need albumentations for this).</p>\n<p>Using the aforementioned techniques I got an accuracy of 0.99 on a random test set of 79 images coming from the \"train\" set (I put them apart before training so that they weren't used for training purposes but only to test the model performance).</p>\n<p>I must stress that deleting blanks in the original images before downsampling them didn't improve the results of the classifier, showing that the only you need is to downsample to 512x512 size as is doesn't make you incur significant data loss, so that you can simplify your code and save a lot of preprocessing time.</p>\n<p>Summarizing, there are 2 stages in the training process:<br>\n<strong>1. Multi class classification</strong><br>\nThe model pretrains on an augmented set of images from both \"train\" and \"other\" sets provided in the competition.<br>\n<strong>2. Binary classification</strong><br>\nThe model turns in a binary classifier and fine-tunes the weights of the \"classifier\" (full connected modules) on the top of the convolutional part of the previously trained learner.</p>\n<p>This is the link to the notebook: <a href=\"https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706\" target=\"_blank\">https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706</a>. I hope you find it useful.</p>\n<p>Please feel free to give me your insights about the techniques shown here and your comments about your own experience using them.</p>\n<p>Thanks for taking your time to read this post.</p>",
      "rawMarkdown": "The purpose of this post is to show the power of pretraining a simple custom classifier on an image set linked to your current task instead of using a much more sophisticated model (such as ResNet or similar) pretrained on a set of images nothing to do with your specific task (being the task here to classify the blood clot origins in ischemic stroke as defined in the Mayo-Clinic - STRIP AI competition description).\n\nThe training of the learner is based on multilabel stratified K-Folds cross-validation and makes heavy use of very simple randomization functions to augment the image set in every cross-validation loop (you don't really need albumentations for this).\n\nUsing the aforementioned techniques I got an accuracy of 0.99 on a random test set of 79 images coming from the \"train\" set (I put them apart before training so that they weren't used for training purposes but only to test the model performance).\n\nI must stress that deleting blanks in the original images before downsampling them didn't improve the results of the classifier, showing that the only you need is to downsample to 512x512 size as is doesn't make you incur significant data loss, so that you can simplify your code and save a lot of preprocessing time.\n\nSummarizing, there are 2 stages in the training process:\n**1. Multi class classification**\nThe model pretrains on an augmented set of images from both \"train\" and \"other\" sets provided in the competition.\n**2. Binary classification**\nThe model turns in a binary classifier and fine-tunes the weights of the \"classifier\" (full connected modules) on the top of the convolutional part of the previously trained learner.\n\nThis is the link to the notebook: https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706. I hope you find it useful.\n\nPlease feel free to give me your insights about the techniques shown here and your comments about your own experience using them.\n\nThanks for taking your time to read this post.",
      "votes": 3
    },
    {
      "id": 1979836,
      "postDate": "2022-10-09T18:40:43.330Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/td-iceman\" target=\"_blank\">@td-iceman</a> for your comment about my notebook. The only metric I used to check CV performance was the average loss as the goal was to reach the minimal loss. I didn't keep track of the variance in order to keep the model as simple as possible even in terms of metrics, but I noticed that the increase of the number of CV folders up to 16 made the validation average loss to lower smoother and smoother in every epoch. I got the best results with this number of folders but the execution time made it difficult to use in a kaggle notebook due to the time restrictions, so I finally tried with only 5 folders.</p>\n<p>I tried several models using different activations and so on. The model I've shown here wasn't the best one in terms of validation loss but the convergence was quite fast in terms of execution time. The main issue using 5 folders is the sudden unstability in the validation mean loss every certain number of epochs (as you will see in the graphic I attach here), but you can avoid any side effects from this just implementing a callback to always get the best model (the one with the lowest validation loss) as I did.</p>\n<p>The model is ready to use differente activation functions for base and fine tuned models respectively but in this case I used GELU in both. LeakyReLU (with negative slope=0.02) is a great choice too. If you use 16 CV folders you get a very smooth curve using LeakyReLU and can see how both the base and fine tuned models decrease the validation loss really fast and get very low values in the minimum number of epochs in comparison with the model using GELU activations.</p>\n<p>This is the curve I got from Tensorboard when trained this model. I hope this graph, in addition with the figures from the training you can see in the notebook, will answer your question:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055197%2F65b4d2bac3da7660a622a0623bea9539%2FValidation_loss.PNG?generation=1665338795584061&amp;alt=media\" alt=\"\"></p>\n<p>This model wasn't ranked for the purpose of the competition due to an scoring error, but you can check its performance taking a look on the standard metrics I used for classifiers: ROC-AUC and accuracy scoring.</p>\n<p>Thanks for your attention and congrats for your excelent results in your first Kaggle competition.</p>",
      "rawMarkdown": "Thank you @td-iceman for your comment about my notebook. The only metric I used to check CV performance was the average loss as the goal was to reach the minimal loss. I didn't keep track of the variance in order to keep the model as simple as possible even in terms of metrics, but I noticed that the increase of the number of CV folders up to 16 made the validation average loss to lower smoother and smoother in every epoch. I got the best results with this number of folders but the execution time made it difficult to use in a kaggle notebook due to the time restrictions, so I finally tried with only 5 folders.\n\nI tried several models using different activations and so on. The model I've shown here wasn't the best one in terms of validation loss but the convergence was quite fast in terms of execution time. The main issue using 5 folders is the sudden unstability in the validation mean loss every certain number of epochs (as you will see in the graphic I attach here), but you can avoid any side effects from this just implementing a callback to always get the best model (the one with the lowest validation loss) as I did.\n\nThe model is ready to use differente activation functions for base and fine tuned models respectively but in this case I used GELU in both. LeakyReLU (with negative slope=0.02) is a great choice too. If you use 16 CV folders you get a very smooth curve using LeakyReLU and can see how both the base and fine tuned models decrease the validation loss really fast and get very low values in the minimum number of epochs in comparison with the model using GELU activations.\n\nThis is the curve I got from Tensorboard when trained this model. I hope this graph, in addition with the figures from the training you can see in the notebook, will answer your question:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055197%2F65b4d2bac3da7660a622a0623bea9539%2FValidation_loss.PNG?generation=1665338795584061&alt=media)\n\nThis model wasn't ranked for the purpose of the competition due to an scoring error, but you can check its performance taking a look on the standard metrics I used for classifiers: ROC-AUC and accuracy scoring.\n\nThanks for your attention and congrats for your excelent results in your first Kaggle competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1979853,
          "postDate": "2022-10-09T19:04:52.980Z",
          "content": "<p><a href=\"https://www.kaggle.com/josmiguelmiz\" target=\"_blank\">@josmiguelmiz</a> thanks so much for your detailed explanation! I get the idea clearly now, its great to see so many good implementation like yours in this competition. I also found it tough to track end to end CV for my training, and I understand your point of lack of GPU hours to try so many things and track metrics. I was only wondering because your training seemed to produce very impressive validation metrics, which I did not know was possible with this data - really nice work.</p>\n<p>And thanks for your final comment, and congrats for your result as well :)</p>",
          "rawMarkdown": "@josmiguelmiz thanks so much for your detailed explanation! I get the idea clearly now, its great to see so many good implementation like yours in this competition. I also found it tough to track end to end CV for my training, and I understand your point of lack of GPU hours to try so many things and track metrics. I was only wondering because your training seemed to produce very impressive validation metrics, which I did not know was possible with this data - really nice work.\n\nAnd thanks for your final comment, and congrats for your result as well :)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1979700,
      "postDate": "2022-10-09T16:28:51.040Z",
      "content": "<p><a href=\"https://www.kaggle.com/josmiguelmiz\" target=\"_blank\">@josmiguelmiz</a> its a nice idea to pretrain your network on all relevant data before training it to focus on the actual task, thanks for the notebook. The 2nd place solution did something similar. Did you have any CV metrics for your methodology, just to get a bigger picture of the training and fitting?</p>",
      "rawMarkdown": "@josmiguelmiz its a nice idea to pretrain your network on all relevant data before training it to focus on the actual task, thanks for the notebook. The 2nd place solution did something similar. Did you have any CV metrics for your methodology, just to get a bigger picture of the training and fitting?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1979836,
      "author_name": "José Miguel Máiz",
      "author_url": "",
      "post_date": "2022-10-09T18:40:43.330000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/td-iceman\" target=\"_blank\">@td-iceman</a> for your comment about my notebook. The only metric I used to check CV performance was the average loss as the goal was to reach the minimal loss. I didn't keep track of the variance in order to keep the model as simple as possible even in terms of metrics, but I noticed that the increase of the number of CV folders up to 16 made the validation average loss to lower smoother and smoother in every epoch. I got the best results with this number of folders but the execution time made it difficult to use in a kaggle notebook due to the time restrictions, so I finally tried with only 5 folders.</p>\n<p>I tried several models using different activations and so on. The model I've shown here wasn't the best one in terms of validation loss but the convergence was quite fast in terms of execution time. The main issue using 5 folders is the sudden unstability in the validation mean loss every certain number of epochs (as you will see in the graphic I attach here), but you can avoid any side effects from this just implementing a callback to always get the best model (the one with the lowest validation loss) as I did.</p>\n<p>The model is ready to use differente activation functions for base and fine tuned models respectively but in this case I used GELU in both. LeakyReLU (with negative slope=0.02) is a great choice too. If you use 16 CV folders you get a very smooth curve using LeakyReLU and can see how both the base and fine tuned models decrease the validation loss really fast and get very low values in the minimum number of epochs in comparison with the model using GELU activations.</p>\n<p>This is the curve I got from Tensorboard when trained this model. I hope this graph, in addition with the figures from the training you can see in the notebook, will answer your question:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055197%2F65b4d2bac3da7660a622a0623bea9539%2FValidation_loss.PNG?generation=1665338795584061&amp;alt=media\" alt=\"\"></p>\n<p>This model wasn't ranked for the purpose of the competition due to an scoring error, but you can check its performance taking a look on the standard metrics I used for classifiers: ROC-AUC and accuracy scoring.</p>\n<p>Thanks for your attention and congrats for your excelent results in your first Kaggle competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1979853,
          "author_name": "tdiceman",
          "author_url": "",
          "post_date": "2022-10-09T19:04:52.980000",
          "content": "<p><a href=\"https://www.kaggle.com/josmiguelmiz\" target=\"_blank\">@josmiguelmiz</a> thanks so much for your detailed explanation! I get the idea clearly now, its great to see so many good implementation like yours in this competition. I also found it tough to track end to end CV for my training, and I understand your point of lack of GPU hours to try so many things and track metrics. I was only wondering because your training seemed to produce very impressive validation metrics, which I did not know was possible with this data - really nice work.</p>\n<p>And thanks for your final comment, and congrats for your result as well :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1979700,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-10-09T16:28:51.040000",
      "content": "<p><a href=\"https://www.kaggle.com/josmiguelmiz\" target=\"_blank\">@josmiguelmiz</a> its a nice idea to pretrain your network on all relevant data before training it to focus on the actual task, thanks for the notebook. The 2nd place solution did something similar. Did you have any CV metrics for your methodology, just to get a bigger picture of the training and fitting?</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1979393": "The purpose of this post is to show the power of pretraining a simple custom classifier on an image set linked to your current task instead of using a much more sophisticated model (such as ResNet or similar) pretrained on a set of images nothing to do with your specific task (being the task here to classify the blood clot origins in ischemic stroke as defined in the Mayo-Clinic - STRIP AI competition description).\n\nThe training of the learner is based on multilabel stratified K-Folds cross-validation and makes heavy use of very simple randomization functions to augment the image set in every cross-validation loop (you don't really need albumentations for this).\n\nUsing the aforementioned techniques I got an accuracy of 0.99 on a random test set of 79 images coming from the \"train\" set (I put them apart before training so that they weren't used for training purposes but only to test the model performance).\n\nI must stress that deleting blanks in the original images before downsampling them didn't improve the results of the classifier, showing that the only you need is to downsample to 512x512 size as is doesn't make you incur significant data loss, so that you can simplify your code and save a lot of preprocessing time.\n\nSummarizing, there are 2 stages in the training process:\n**1. Multi class classification**\nThe model pretrains on an augmented set of images from both \"train\" and \"other\" sets provided in the competition.\n**2. Binary classification**\nThe model turns in a binary classifier and fine-tunes the weights of the \"classifier\" (full connected modules) on the top of the convolutional part of the previously trained learner.\n\nThis is the link to the notebook: https://www.kaggle.com/code/josmiguelmiz/mayo-clinic-image/notebook?scriptVersionId=107281706. I hope you find it useful.\n\nPlease feel free to give me your insights about the techniques shown here and your comments about your own experience using them.\n\nThanks for taking your time to read this post.",
    "1979836": "Thank you @td-iceman for your comment about my notebook. The only metric I used to check CV performance was the average loss as the goal was to reach the minimal loss. I didn't keep track of the variance in order to keep the model as simple as possible even in terms of metrics, but I noticed that the increase of the number of CV folders up to 16 made the validation average loss to lower smoother and smoother in every epoch. I got the best results with this number of folders but the execution time made it difficult to use in a kaggle notebook due to the time restrictions, so I finally tried with only 5 folders.\n\nI tried several models using different activations and so on. The model I've shown here wasn't the best one in terms of validation loss but the convergence was quite fast in terms of execution time. The main issue using 5 folders is the sudden unstability in the validation mean loss every certain number of epochs (as you will see in the graphic I attach here), but you can avoid any side effects from this just implementing a callback to always get the best model (the one with the lowest validation loss) as I did.\n\nThe model is ready to use differente activation functions for base and fine tuned models respectively but in this case I used GELU in both. LeakyReLU (with negative slope=0.02) is a great choice too. If you use 16 CV folders you get a very smooth curve using LeakyReLU and can see how both the base and fine tuned models decrease the validation loss really fast and get very low values in the minimum number of epochs in comparison with the model using GELU activations.\n\nThis is the curve I got from Tensorboard when trained this model. I hope this graph, in addition with the figures from the training you can see in the notebook, will answer your question:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5055197%2F65b4d2bac3da7660a622a0623bea9539%2FValidation_loss.PNG?generation=1665338795584061&alt=media)\n\nThis model wasn't ranked for the purpose of the competition due to an scoring error, but you can check its performance taking a look on the standard metrics I used for classifiers: ROC-AUC and accuracy scoring.\n\nThanks for your attention and congrats for your excelent results in your first Kaggle competition.",
    "1979700": "@josmiguelmiz its a nice idea to pretrain your network on all relevant data before training it to focus on the actual task, thanks for the notebook. The 2nd place solution did something similar. Did you have any CV metrics for your methodology, just to get a bigger picture of the training and fitting?"
  }
}