{
  "id": 169178,
  "title": "18th place solution: DenseNet + RNN based",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169178",
  "author_name": "Dmitry A. Grechka",
  "post_date": "2020-07-23T06:16:06.870000",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers, Kaggle team and all of the participating kagglers!</p>\n\n<p>That's the first time when I get so high place! I'm really glad that all of my work done in these 3 months is rewarded.\nAnd as usual I learned a lot during this challenge!</p>\n\n<p>My solution is quite different from the concat pooling (by Iafoss) based mainstream approach.</p>\n\n<p>It is:</p>\n\n<p>0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\n1) rotation of a whole (middle resolution) image to arbitrary angle with crop preventions\n2) extraction of tissue tiles (256x256)\n3) Global Contrast Normalization (across all of the extracted tiles from single image)\n4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n5) Multiple generations of discarding the \"hard or wrong labelled\" images by MAE&gt;2.5 threshold\n6) 5-Fold CV during training, keeping the gleason_score frequencies balanced while splitting the train and validation.\n7) 3 stage training:\n     - backbone frozen, long tile sequence 64 tiles\n     - all unfrozen, shorter tile sequence of 16 tiles\n     - backbone frozen, long tile sequence of 64 tiles again\n8) short train batch size of 2 (for regularizing effect)</p>",
  "messages": [
    {
      "id": 941064,
      "postDate": "2020-07-23T06:16:06.870Z",
      "content": "<p>First of all, thanks to the organizers, Kaggle team and all of the participating kagglers!</p>\n\n<p>That's the first time when I get so high place! I'm really glad that all of my work done in these 3 months is rewarded.\nAnd as usual I learned a lot during this challenge!</p>\n\n<p>My solution is quite different from the concat pooling (by Iafoss) based mainstream approach.</p>\n\n<p>It is:</p>\n\n<p>0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\n1) rotation of a whole (middle resolution) image to arbitrary angle with crop preventions\n2) extraction of tissue tiles (256x256)\n3) Global Contrast Normalization (across all of the extracted tiles from single image)\n4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n5) Multiple generations of discarding the \"hard or wrong labelled\" images by MAE&gt;2.5 threshold\n6) 5-Fold CV during training, keeping the gleason_score frequencies balanced while splitting the train and validation.\n7) 3 stage training:\n     - backbone frozen, long tile sequence 64 tiles\n     - all unfrozen, shorter tile sequence of 16 tiles\n     - backbone frozen, long tile sequence of 64 tiles again\n8) short train batch size of 2 (for regularizing effect)</p>",
      "rawMarkdown": "First of all, thanks to the organizers, Kaggle team and all of the participating kagglers!\n\nThat's the first time when I get so high place! I'm really glad that all of my work done in these 3 months is rewarded.\nAnd as usual I learned a lot during this challenge!\n\n\nMy solution is quite different from the concat pooling (by Iafoss) based mainstream approach.\n\nIt is:\n\n0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\n1) rotation of a whole (middle resolution) image to arbitrary angle with crop preventions\n2) extraction of tissue tiles (256x256)\n3) Global Contrast Normalization (across all of the extracted tiles from single image)\n4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n5) Multiple generations of discarding the \"hard or wrong labelled\" images by MAE&gt;2.5 threshold\n6) 5-Fold CV during training, keeping the gleason_score frequencies balanced while splitting the train and validation.\n7) 3 stage training:\n     - backbone frozen, long tile sequence 64 tiles\n     - all unfrozen, shorter tile sequence of 16 tiles\n     - backbone frozen, long tile sequence of 64 tiles again\n8) short train batch size of 2 (for regularizing effect)",
      "votes": 12
    },
    {
      "id": 941551,
      "postDate": "2020-07-23T09:42:13.343Z",
      "content": "<p>\"3 stage training:\n- backbone frozen, long tile sequence 64 tiles\n- all unfrozen, shorter tile sequence of 16 tiles\n- backbone frozen, long tile sequence of 64 tiles again\"\nWhat insight does this training approach come from? What's the difference between it and end-to-end training?</p>",
      "rawMarkdown": "\"3 stage training:\n- backbone frozen, long tile sequence 64 tiles\n- all unfrozen, shorter tile sequence of 16 tiles\n- backbone frozen, long tile sequence of 64 tiles again\"\nWhat insight does this training approach come from? What's the difference between it and end-to-end training?",
      "votes": 2,
      "replies": [
        {
          "id": 941577,
          "postDate": "2020-07-23T10:00:23.617Z",
          "content": "<p>First stage (frozen DenseNet backbone with imagenet weights) is for preventing imagenet features from being completely wiped by large error gradient originating from random initialized later layers. Thus transfer learning is utilized.</p>\n\n<p>2nd and 3rd are split due to GPU resources limitation.\nI would have used end-to-end training (with all the network unfrozen and long sequence for GRU units) but it did not fit into the GPU memory.</p>\n\n<p>Thus I split it.\n2nd phase (all network is unfrozen, shorter sequence) is to tune the visual features extractor (densenet weights) for the particular application.</p>\n\n<p>3rd phase is aimed to tune the GRU units with long enough sequences. Plus it can be seen as a variation of <a href=\"https://arxiv.org/pdf/1706.04983.pdf\">FreezeOut</a>. </p>",
          "rawMarkdown": "First stage (frozen DenseNet backbone with imagenet weights) is for preventing imagenet features from being completely wiped by large error gradient originating from random initialized later layers. Thus transfer learning is utilized.\n\n2nd and 3rd are split due to GPU resources limitation.\nI would have used end-to-end training (with all the network unfrozen and long sequence for GRU units) but it did not fit into the GPU memory.\n\nThus I split it.\n2nd phase (all network is unfrozen, shorter sequence) is to tune the visual features extractor (densenet weights) for the particular application.\n\n3rd phase is aimed to tune the GRU units with long enough sequences. Plus it can be seen as a variation of [FreezeOut](https://arxiv.org/pdf/1706.04983.pdf). ",
          "votes": 1
        },
        {
          "id": 947494,
          "postDate": "2020-07-27T09:56:07.937Z",
          "content": "<p>Got it, Thanks!\nBut I have another question about \"0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\", Could you explain it to me?</p>",
          "rawMarkdown": "Got it, Thanks!\nBut I have another question about \"0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\", Could you explain it to me?"
        },
        {
          "id": 948827,
          "postDate": "2020-07-28T08:45:53.877Z",
          "content": "<p><a href=\"/amshoreline\">@amshoreline</a> Sure.</p>\n\n<p>The idea is to prevent similar images from being in train and validation set. If such images appear, there will be information leak from training set to validation, and thus validation metrics will be biased (also <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155954\">discussed here</a>).</p>\n\n<p>For instance consider these two images:\n| 6226ebfc1f9b743a8b02db4eb7145738 | 3c659b2837afab3af6b952fcbaa6a515|\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2Fdde60e6fe94f5fa52ab84020f32dac4d%2F6226ebfc1f9b743a8b02db4eb7145738.png?generation=1595925176175386&amp;alt=media\" alt=\"\">  |  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2F9a2d686c4a616542bdadc7f810c31030%2F3c659b2837afab3af6b952fcbaa6a515.png?generation=1595925212356794&amp;alt=media\" alt=\"\"> |</p>\n\n<p>My decision was to put similar images like above either in training set or in validation, but prevent images of similar series from getting into the both.</p>\n\n<p>I calculated image hashes using <a href=\"https://pypi.org/project/ImageHash/\">ImageHash</a> to be able to evaluate how similar the train images are.\nThen I calculated distance (hamming distance) between each image pair. And merged the images that are close into the clusters (recursive approach: if any of the images of two clusters have distance smaller than a threshold, these two clusters merge into single one)</p>\n\n<p>Then I took \"representative\" from each of the cluster and did a training/validation k-fold split using these \"representatives\"</p>",
          "rawMarkdown": "@amshoreline Sure.\n\nThe idea is to prevent similar images from being in train and validation set. If such images appear, there will be information leak from training set to validation, and thus validation metrics will be biased (also [discussed here](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155954)).\n\nFor instance consider these two images:\n| 6226ebfc1f9b743a8b02db4eb7145738 | 3c659b2837afab3af6b952fcbaa6a515|\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2Fdde60e6fe94f5fa52ab84020f32dac4d%2F6226ebfc1f9b743a8b02db4eb7145738.png?generation=1595925176175386&amp;alt=media)  |  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2F9a2d686c4a616542bdadc7f810c31030%2F3c659b2837afab3af6b952fcbaa6a515.png?generation=1595925212356794&amp;alt=media) |\n\nMy decision was to put similar images like above either in training set or in validation, but prevent images of similar series from getting into the both.\n\n\n\nI calculated image hashes using [ImageHash](https://pypi.org/project/ImageHash/) to be able to evaluate how similar the train images are.\nThen I calculated distance (hamming distance) between each image pair. And merged the images that are close into the clusters (recursive approach: if any of the images of two clusters have distance smaller than a threshold, these two clusters merge into single one)\n\nThen I took \"representative\" from each of the cluster and did a training/validation k-fold split using these \"representatives\""
        }
      ]
    },
    {
      "id": 948202,
      "postDate": "2020-07-27T18:13:15.150Z",
      "content": "<p><a href=\"/dgrechka\">@dgrechka</a> \nCongrats high place &amp; thanks for sharing cool approach!</p>\n\n<blockquote>\n  <p>4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)</p>\n</blockquote>\n\n<p>Your approach with RNN is original and very interesting! <br>\nIn general, I think the order of input is often important when using RNNs. <br>\nHow did you do with the order in which you input tiles to model? (Or is it random?)</p>",
      "rawMarkdown": "@dgrechka \nCongrats high place &amp; thanks for sharing cool approach!\n\n&gt; 4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n\nYour approach with RNN is original and very interesting!  \nIn general, I think the order of input is often important when using RNNs.  \nHow did you do with the order in which you input tiles to model? (Or is it random?)",
      "replies": [
        {
          "id": 948826,
          "postDate": "2020-07-28T08:45:10.893Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 948835,
          "postDate": "2020-07-28T08:59:44.883Z",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> , Thank you!</p>\n\n<p>I sorted the tiles by the brightness descending (I worked with negative image: 255 - original Image). The brightest one goes first.</p>\n\n<p>But there was an important step before the ordering. I left only those tiles, which had more green than red. This is to filter out pen or marker marks which often look like large white straps in the negative image. If I did not do it, the white pen mark tiles went first in the brightness sorted sequence. And that was an issue.</p>\n\n<p>I coerced the training sequence to the needed length.\nIf there were too few tiles originally, I repeated (cycled) the sequence to match the needed length.\nIt there were too many tiles, I trimmed the sequence (discarded the \"tail\").</p>\n\n<p>The training was very sensitive to the input order, when I did order shuffling the model did not train at all.\nI guess that is because the tiles with relative information could be simply trimmed out for the cases when original sequence was too long, and the training signal was completely missing.</p>",
          "rawMarkdown": "@yukkyo , Thank you!\n\nI sorted the tiles by the brightness descending (I worked with negative image: 255 - original Image). The brightest one goes first.\n\nBut there was an important step before the ordering. I left only those tiles, which had more green than red. This is to filter out pen or marker marks which often look like large white straps in the negative image. If I did not do it, the white pen mark tiles went first in the brightness sorted sequence. And that was an issue.\n\nI coerced the training sequence to the needed length.\nIf there were too few tiles originally, I repeated (cycled) the sequence to match the needed length.\nIt there were too many tiles, I trimmed the sequence (discarded the \"tail\").\n\nThe training was very sensitive to the input order, when I did order shuffling the model did not train at all.\nI guess that is because the tiles with relative information could be simply trimmed out for the cases when original sequence was too long, and the training signal was completely missing.",
          "votes": 1
        },
        {
          "id": 949303,
          "postDate": "2020-07-28T14:44:26.727Z",
          "content": "<p><a href=\"/dgrechka\">@dgrechka</a> \nI see!\nYour preprocessing showed a deep understanding of the data, and this sort method made a lot of sense to me.</p>\n\n<p>Keep up the great work!</p>",
          "rawMarkdown": "@dgrechka \nI see!\nYour preprocessing showed a deep understanding of the data, and this sort method made a lot of sense to me.\n\nKeep up the great work!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 941551,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T09:42:13.343000",
      "content": "<p>\"3 stage training:\n- backbone frozen, long tile sequence 64 tiles\n- all unfrozen, shorter tile sequence of 16 tiles\n- backbone frozen, long tile sequence of 64 tiles again\"\nWhat insight does this training approach come from? What's the difference between it and end-to-end training?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 941577,
          "author_name": "Dmitry A. Grechka",
          "author_url": "",
          "post_date": "2020-07-23T10:00:23.617000",
          "content": "<p>First stage (frozen DenseNet backbone with imagenet weights) is for preventing imagenet features from being completely wiped by large error gradient originating from random initialized later layers. Thus transfer learning is utilized.</p>\n\n<p>2nd and 3rd are split due to GPU resources limitation.\nI would have used end-to-end training (with all the network unfrozen and long sequence for GRU units) but it did not fit into the GPU memory.</p>\n\n<p>Thus I split it.\n2nd phase (all network is unfrozen, shorter sequence) is to tune the visual features extractor (densenet weights) for the particular application.</p>\n\n<p>3rd phase is aimed to tune the GRU units with long enough sequences. Plus it can be seen as a variation of <a href=\"https://arxiv.org/pdf/1706.04983.pdf\">FreezeOut</a>. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 947494,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-27T09:56:07.937000",
          "content": "<p>Got it, Thanks!\nBut I have another question about \"0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\", Could you explain it to me?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 948827,
          "author_name": "Dmitry A. Grechka",
          "author_url": "",
          "post_date": "2020-07-28T08:45:53.877000",
          "content": "<p><a href=\"/amshoreline\">@amshoreline</a> Sure.</p>\n\n<p>The idea is to prevent similar images from being in train and validation set. If such images appear, there will be information leak from training set to validation, and thus validation metrics will be biased (also <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/155954\">discussed here</a>).</p>\n\n<p>For instance consider these two images:\n| 6226ebfc1f9b743a8b02db4eb7145738 | 3c659b2837afab3af6b952fcbaa6a515|\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2Fdde60e6fe94f5fa52ab84020f32dac4d%2F6226ebfc1f9b743a8b02db4eb7145738.png?generation=1595925176175386&amp;alt=media\" alt=\"\">  |  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F615797%2F9a2d686c4a616542bdadc7f810c31030%2F3c659b2837afab3af6b952fcbaa6a515.png?generation=1595925212356794&amp;alt=media\" alt=\"\"> |</p>\n\n<p>My decision was to put similar images like above either in training set or in validation, but prevent images of similar series from getting into the both.</p>\n\n<p>I calculated image hashes using <a href=\"https://pypi.org/project/ImageHash/\">ImageHash</a> to be able to evaluate how similar the train images are.\nThen I calculated distance (hamming distance) between each image pair. And merged the images that are close into the clusters (recursive approach: if any of the images of two clusters have distance smaller than a threshold, these two clusters merge into single one)</p>\n\n<p>Then I took \"representative\" from each of the cluster and did a training/validation k-fold split using these \"representatives\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948202,
      "author_name": "fam_taro",
      "author_url": "",
      "post_date": "2020-07-27T18:13:15.150000",
      "content": "<p><a href=\"/dgrechka\">@dgrechka</a> \nCongrats high place &amp; thanks for sharing cool approach!</p>\n\n<blockquote>\n  <p>4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)</p>\n</blockquote>\n\n<p>Your approach with RNN is original and very interesting! <br>\nIn general, I think the order of input is often important when using RNNs. <br>\nHow did you do with the order in which you input tiles to model? (Or is it random?)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 948826,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-28T08:45:10.893000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 948835,
          "author_name": "Dmitry A. Grechka",
          "author_url": "",
          "post_date": "2020-07-28T08:59:44.883000",
          "content": "<p><a href=\"/yukkyo\">@yukkyo</a> , Thank you!</p>\n\n<p>I sorted the tiles by the brightness descending (I worked with negative image: 255 - original Image). The brightest one goes first.</p>\n\n<p>But there was an important step before the ordering. I left only those tiles, which had more green than red. This is to filter out pen or marker marks which often look like large white straps in the negative image. If I did not do it, the white pen mark tiles went first in the brightness sorted sequence. And that was an issue.</p>\n\n<p>I coerced the training sequence to the needed length.\nIf there were too few tiles originally, I repeated (cycled) the sequence to match the needed length.\nIt there were too many tiles, I trimmed the sequence (discarded the \"tail\").</p>\n\n<p>The training was very sensitive to the input order, when I did order shuffling the model did not train at all.\nI guess that is because the tiles with relative information could be simply trimmed out for the cases when original sequence was too long, and the training signal was completely missing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 949303,
          "author_name": "fam_taro",
          "author_url": "",
          "post_date": "2020-07-28T14:44:26.727000",
          "content": "<p><a href=\"/dgrechka\">@dgrechka</a> \nI see!\nYour preprocessing showed a deep understanding of the data, and this sort method made a lot of sense to me.</p>\n\n<p>Keep up the great work!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "941064": "First of all, thanks to the organizers, Kaggle team and all of the participating kagglers!\n\nThat's the first time when I get so high place! I'm really glad that all of my work done in these 3 months is rewarded.\nAnd as usual I learned a lot during this challenge!\n\n\nMy solution is quite different from the concat pooling (by Iafoss) based mainstream approach.\n\nIt is:\n\n0) image similarity clustering via image hashing (splitting clusters into tr/val sets, not images)\n1) rotation of a whole (middle resolution) image to arbitrary angle with crop preventions\n2) extraction of tissue tiles (256x256)\n3) Global Contrast Normalization (across all of the extracted tiles from single image)\n4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n5) Multiple generations of discarding the \"hard or wrong labelled\" images by MAE&gt;2.5 threshold\n6) 5-Fold CV during training, keeping the gleason_score frequencies balanced while splitting the train and validation.\n7) 3 stage training:\n     - backbone frozen, long tile sequence 64 tiles\n     - all unfrozen, shorter tile sequence of 16 tiles\n     - backbone frozen, long tile sequence of 64 tiles again\n8) short train batch size of 2 (for regularizing effect)",
    "941551": "\"3 stage training:\n- backbone frozen, long tile sequence 64 tiles\n- all unfrozen, shorter tile sequence of 16 tiles\n- backbone frozen, long tile sequence of 64 tiles again\"\nWhat insight does this training approach come from? What's the difference between it and end-to-end training?",
    "948202": "@dgrechka \nCongrats high place &amp; thanks for sharing cool approach!\n\n&gt; 4) DenseNet121 backbone (imagenet pretrained) -&gt; Dense feature extractor -&gt; 2 GRU layers -&gt; single head ISUP grade regression (logcosh loss)\n\nYour approach with RNN is original and very interesting!  \nIn general, I think the order of input is often important when using RNNs.  \nHow did you do with the order in which you input tiles to model? (Or is it random?)"
  }
}