{
  "id": 152265,
  "title": "using @lafoss minimum model and how to improve",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/152265",
  "author_name": "hengck23",
  "post_date": "2020-05-19T04:43:27.681000",
  "votes": 69,
  "comment_count": 46,
  "views": 0,
  "content": "<p>as attached</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd5ca75ad58e41520efe7b3ecb587cd1%2FSlide1.png?generation=1589863350980222&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe81ae8bb26f98f33140c85c615d93%2FSlide2.png?generation=1589863399965273&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc894777e08d0a57a4e7bcf52a3ab0fe0%2FSlide3.png?generation=1589863398695926&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fee0d3d1eb9cf0e4828dd69d53878a021%2FSlide4.png?generation=1589863402368753&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 853305,
      "postDate": "2020-05-19T04:43:27.683Z",
      "content": "<p>as attached</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd5ca75ad58e41520efe7b3ecb587cd1%2FSlide1.png?generation=1589863350980222&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe81ae8bb26f98f33140c85c615d93%2FSlide2.png?generation=1589863399965273&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc894777e08d0a57a4e7bcf52a3ab0fe0%2FSlide3.png?generation=1589863398695926&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fee0d3d1eb9cf0e4828dd69d53878a021%2FSlide4.png?generation=1589863402368753&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "as attached\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd5ca75ad58e41520efe7b3ecb587cd1%2FSlide1.png?generation=1589863350980222&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe81ae8bb26f98f33140c85c615d93%2FSlide2.png?generation=1589863399965273&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc894777e08d0a57a4e7bcf52a3ab0fe0%2FSlide3.png?generation=1589863398695926&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fee0d3d1eb9cf0e4828dd69d53878a021%2FSlide4.png?generation=1589863402368753&amp;alt=media)\n",
      "votes": 67
    },
    {
      "id": 853715,
      "postDate": "2020-05-19T12:31:20.687Z",
      "content": "<p>\"Processing Megapixel Images with Deep Attention-Sampling Models\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1905.03711.pdf\">https://arxiv.org/pdf/1905.03711.pdf</a>\n<a href=\"https://www.youtube.com/watch?v=H6Qiegq_36c\">https://www.youtube.com/watch?v=H6Qiegq_36c</a>\n<a href=\"https://github.com/idiap/attention-sampling\">https://github.com/idiap/attention-sampling</a></p>\n\n<p>abstract:\nExisting deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a*<em>fully differentiable end-to-end trainable</em>* model that samples and processes only a fraction of the full resolution input image.</p>\n\n<p>The <strong>locations to process are sampled from an attention distribution computed from a low resolution view of the input.</strong> We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup.</p>",
      "rawMarkdown": "\"Processing Megapixel Images with Deep Attention-Sampling Models\"\n\nhttps://arxiv.org/pdf/1905.03711.pdf\nhttps://www.youtube.com/watch?v=H6Qiegq_36c\nhttps://github.com/idiap/attention-sampling\n\n\nabstract:\nExisting deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a**fully differentiable end-to-end trainable** model that samples and processes only a fraction of the full resolution input image.\n\nThe **locations to process are sampled from an attention distribution computed from a low resolution view of the input.** We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup.",
      "votes": 5,
      "replies": [
        {
          "id": 853717,
          "postDate": "2020-05-19T12:33:26.017Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1ad9553af04fb8b2718ed6d960aaa8c5%2F694eed8d9e67cb1aa2d8ee410a5c97d0a1f36714.png?generation=1589956576501217&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": " ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1ad9553af04fb8b2718ed6d960aaa8c5%2F694eed8d9e67cb1aa2d8ee410a5c97d0a1f36714.png?generation=1589956576501217&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 853720,
          "postDate": "2020-05-19T12:37:39.603Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F00a8df7fbc3a8dbbbaa42c271edf8f0e%2FSelection_156.png?generation=1589891957077940&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F00a8df7fbc3a8dbbbaa42c271edf8f0e%2FSelection_156.png?generation=1589891957077940&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 853738,
          "postDate": "2020-05-19T12:54:13.103Z",
          "content": "<p>related:\nStreaming convolutional neural networks for end-to-end learning with multi-megapixel images\n- Hans Pinckaers (Computational Pathology Group, Radboud University Medical Center)  </p>\n\n<p><a href=\"https://github.com/DIAGNijmegen/StreamingCNN\">https://github.com/DIAGNijmegen/StreamingCNN</a>\n<a href=\"https://www.computationalpathologygroup.eu/software/automated-gleason-grading/\">https://www.computationalpathologygroup.eu/software/automated-gleason-grading/</a></p>",
          "rawMarkdown": "related:\nStreaming convolutional neural networks for end-to-end learning with multi-megapixel images\n- Hans Pinckaers (Computational Pathology Group, Radboud University Medical Center)  \n\n\nhttps://github.com/DIAGNijmegen/StreamingCNN\nhttps://www.computationalpathologygroup.eu/software/automated-gleason-grading/",
          "votes": 3
        },
        {
          "id": 853845,
          "postDate": "2020-05-19T14:36:00.993Z",
          "content": "<p>Maybe you can use the masks to pretrain your attention network.</p>",
          "rawMarkdown": "Maybe you can use the masks to pretrain your attention network.",
          "votes": 2
        },
        {
          "id": 854944,
          "postDate": "2020-05-20T13:15:07.820Z",
          "content": "<p>pytorch version! <a href=\"https://github.com/idiap/attention-sampling/issues/1\">https://github.com/idiap/attention-sampling/issues/1</a></p>",
          "rawMarkdown": "pytorch version! https://github.com/idiap/attention-sampling/issues/1",
          "votes": 1
        }
      ]
    },
    {
      "id": 854734,
      "postDate": "2020-05-20T09:17:50.743Z",
      "content": "<p>Thanks a lot for the references. \nI've also thought about improving iafoss tiling method and created a very basic kernel which uses a hybrid CNN-LSTM in order to learn long-term dependencies between features <a href=\"https://www.kaggle.com/alexj21/hybrid-cnn-lstm-starter\">here</a>\nHope it can provide some ideas too</p>",
      "rawMarkdown": "Thanks a lot for the references. \nI've also thought about improving iafoss tiling method and created a very basic kernel which uses a hybrid CNN-LSTM in order to learn long-term dependencies between features [here](https://www.kaggle.com/alexj21/hybrid-cnn-lstm-starter)\nHope it can provide some ideas too",
      "votes": 3
    },
    {
      "id": 853339,
      "postDate": "2020-05-19T05:29:57.310Z",
      "content": "<p>Links to reference papers:\n<a href=\"https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982\">https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982</a>\n<a href=\"https://arxiv.org/abs/1905.13208\">https://arxiv.org/abs/1905.13208</a></p>",
      "rawMarkdown": "Links to reference papers:\nhttps://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982\nhttps://arxiv.org/abs/1905.13208",
      "votes": 4
    },
    {
      "id": 855472,
      "postDate": "2020-05-21T00:41:51.673Z",
      "content": "<p>Those approaches look exiting, but there may be an issue related to the nature of the labels: they are biased on the relative area. It would be interesting to think how to measure the area not having nearly all tissue tiles as an input. </p>",
      "rawMarkdown": "Those approaches look exiting, but there may be an issue related to the nature of the labels: they are biased on the relative area. It would be interesting to think how to measure the area not having nearly all tissue tiles as an input. ",
      "votes": 1
    },
    {
      "id": 857925,
      "postDate": "2020-05-23T04:07:23.953Z",
      "content": "<p>I tried <code>attention based pooling</code> mentioned in <a href=\"https://jmtomczak.github.io/pdf/Tooploox_2018_06_27.pdf\">this paper</a>. Compared to my previous <code>concat pooling</code> approach (almost identical to @lafoss 's notebook), it didn't give any boost in both CV and LB. Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc. Anyone managed to make it work?</p>",
      "rawMarkdown": "I tried `attention based pooling` mentioned in [this paper](https://jmtomczak.github.io/pdf/Tooploox_2018_06_27.pdf). Compared to my previous `concat pooling` approach (almost identical to @lafoss 's notebook), it didn't give any boost in both CV and LB. Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc. Anyone managed to make it work?",
      "votes": 2,
      "replies": [
        {
          "id": 858912,
          "postDate": "2020-05-23T23:59:20.787Z",
          "content": "<p>I just realized in <a href=\"/iafoss\">@iafoss</a> notebook what AdaptiveConcatPool2d does... I have just been using torch.mean on the feature vectors from the cnn backbone. Have you tried just using max or mean pooling and how much difference it makes?</p>",
          "rawMarkdown": "I just realized in @iafoss notebook what AdaptiveConcatPool2d does... I have just been using torch.mean on the feature vectors from the cnn backbone. Have you tried just using max or mean pooling and how much difference it makes?",
          "votes": 1
        },
        {
          "id": 858918,
          "postDate": "2020-05-24T00:21:46.913Z",
          "content": "<p>Based on my experiment with different pooling layers so far <code>GeM</code> pooling layers works the best. </p>\n\n<p>What is GeM?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>)</p>\n\n<p>code below:</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></p>",
          "rawMarkdown": "Based on my experiment with different pooling layers so far `GeM` pooling layers works the best. \n\nWhat is GeM?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065)\n\ncode below:\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n```",
          "votes": 13
        },
        {
          "id": 858962,
          "postDate": "2020-05-24T01:57:10.480Z",
          "content": "<p>Thanks for the tip. I will be trying this out</p>",
          "rawMarkdown": "Thanks for the tip. I will be trying this out",
          "votes": 1
        },
        {
          "id": 858982,
          "postDate": "2020-05-24T02:50:18.297Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you! I'll give it a try.</p>\n\n<p><a href=\"/shujun717\">@shujun717</a> I tried avg pool and concat pool. They both performed almost the same .</p>",
          "rawMarkdown": "@drhabib Thank you! I'll give it a try.\n\n@shujun717 I tried avg pool and concat pool. They both performed almost the same .",
          "votes": 1
        },
        {
          "id": 859922,
          "postDate": "2020-05-24T21:14:39.157Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> from what I understand AdaptiveConcatPool2d is no different from concating the mean and max of the set of feature vectors produced by the cnn in your notebook correct?</p>",
          "rawMarkdown": "@iafoss from what I understand AdaptiveConcatPool2d is no different from concating the mean and max of the set of feature vectors produced by the cnn in your notebook correct?"
        },
        {
          "id": 859937,
          "postDate": "2020-05-24T21:31:35.180Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a>  Yes. Precisely. It's taking both results of mean and max and concatenating them together. I think in this competition it wouldn't make that much of a difference because we are likely to be more interested in the results of the max pool (was there cancer or not? or what cancer was it?) as opposed to the average cancerness of the whole image. But in cases where we care about both, for example, what was the maximum haze in an image AND what was the overall haziness, then I guess concatenating both might give us more benefit.</p>",
          "rawMarkdown": "@shujun717  Yes. Precisely. It's taking both results of mean and max and concatenating them together. I think in this competition it wouldn't make that much of a difference because we are likely to be more interested in the results of the max pool (was there cancer or not? or what cancer was it?) as opposed to the average cancerness of the whole image. But in cases where we care about both, for example, what was the maximum haze in an image AND what was the overall haziness, then I guess concatenating both might give us more benefit.",
          "votes": 1
        },
        {
          "id": 859941,
          "postDate": "2020-05-24T21:36:20.600Z",
          "content": "<p><a href=\"/yousof9\">@yousof9</a> I would actually expect average to work better than max here considering grades are determined by the relative presence of gleason pattern in the sample. The average for that purpose is a better representation of the bag content.</p>",
          "rawMarkdown": "@yousof9 I would actually expect average to work better than max here considering grades are determined by the relative presence of gleason pattern in the sample. The average for that purpose is a better representation of the bag content."
        },
        {
          "id": 859969,
          "postDate": "2020-05-24T22:45:11.923Z",
          "content": "<p>Actually I would expect that for this particular labels both max and mean are important: mean gives the relative area of different patterns while max gives the worst cancer pattern. If you check the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/additional-resources\">additional clarification</a> provided for this competition by organizers, you will see that the first Gleason score is based on the pattern having the maximum area, while the second corresponds to the worst pattern, if its score is higher than one for the most common pattern, or the second most common pattern having more than 5% of the tissue area. If the second pattern is not present, the second Gleason score is set to be equal to the first one, if I understand it correctly.</p>\n\n<p>For this particular labels in my tests generalized GeM pooling (the value of p is selected individually for each channel) worked worse than concatenate pooling. It's a little bit surprising since in last two competitions I participated in GeM worked better (though the difference was small in comparison with concat), but I'd think it is resulted by the way how labels are selected: based on mean and max values.</p>\n\n<p>And regarding attention pooling (CBAM block added right before pooling), it actually never worked for me. I tried it in 3-4 competitions I participated in during the last year without much success. Probably, the model needs to be trained in some special manner, but somehow I couldn't make it work better than just simple concat or GeM pooling. From the first glance I wd think that attention applied right before pooling could really boost the results, but the reality is a little bit counterintuitive. Probably, some aux added to the attention could guide the model into the right direction, I don't know.</p>",
          "rawMarkdown": "Actually I would expect that for this particular labels both max and mean are important: mean gives the relative area of different patterns while max gives the worst cancer pattern. If you check the [additional clarification](https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/additional-resources) provided for this competition by organizers, you will see that the first Gleason score is based on the pattern having the maximum area, while the second corresponds to the worst pattern, if its score is higher than one for the most common pattern, or the second most common pattern having more than 5% of the tissue area. If the second pattern is not present, the second Gleason score is set to be equal to the first one, if I understand it correctly.\n\nFor this particular labels in my tests generalized GeM pooling (the value of p is selected individually for each channel) worked worse than concatenate pooling. It's a little bit surprising since in last two competitions I participated in GeM worked better (though the difference was small in comparison with concat), but I'd think it is resulted by the way how labels are selected: based on mean and max values.\n\nAnd regarding attention pooling (CBAM block added right before pooling), it actually never worked for me. I tried it in 3-4 competitions I participated in during the last year without much success. Probably, the model needs to be trained in some special manner, but somehow I couldn't make it work better than just simple concat or GeM pooling. From the first glance I wd think that attention applied right before pooling could really boost the results, but the reality is a little bit counterintuitive. Probably, some aux added to the attention could guide the model into the right direction, I don't know.",
          "votes": 5
        },
        {
          "id": 860001,
          "postDate": "2020-05-24T23:57:54.957Z",
          "content": "<p>True. Actually I use both anyway.</p>\n\n<p>Rergarding attention pooling I just developpe my own based on a paper above. I get similar performance as with concat pooling. Main advantage is that it gives some interpretability as it shows which tiles it is focusing on. But I don't really understand what cancer really looks like so I'm not sure it's doing the right thing actually :D</p>\n\n<p>Expecting a BN, C tensor here is my attention based on Deep MIL paper.\n```\nclass AttentionPool(nn.Module):</p>\n\n<pre><code>def __init__(self, c_in, d, n_tiles):\n    super().__init__()\n    self.maxpool = AdaptiveConcatPool2d() # Edit: Please remove this\n    self.lin_key = nn.Linear(c_in, d)\n    self.lin_w = nn.Linear(d, 1)\n    self.n_tiles = n_tiles\n\ndef forward(self, x):\n    bn, c = x.shape\n    keys = self.lin_key(x)  #bn, d\n    weights = self.lin_w(torch.tanh(keys))   #bn, 1\n    weights = weights.reshape(-1, self.n_tiles)   #b, n\n    weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n    return weights\n</code></pre>\n\n<p>```</p>\n\n<p>Then in the head you simply do <code>pooled = torch.matmul(x.transpose(1, 2), weights).squeeze(2)</code> where x is a b, n, c_in tensor representing the bags of tiles (output of CNN).</p>",
          "rawMarkdown": "True. Actually I use both anyway.\n\nRergarding attention pooling I just developpe my own based on a paper above. I get similar performance as with concat pooling. Main advantage is that it gives some interpretability as it shows which tiles it is focusing on. But I don't really understand what cancer really looks like so I'm not sure it's doing the right thing actually :D\n\nExpecting a BN, C tensor here is my attention based on Deep MIL paper.\n```\nclass AttentionPool(nn.Module):\n\n    def __init__(self, c_in, d, n_tiles):\n        super().__init__()\n        self.maxpool = AdaptiveConcatPool2d() # Edit: Please remove this\n        self.lin_key = nn.Linear(c_in, d)\n        self.lin_w = nn.Linear(d, 1)\n        self.n_tiles = n_tiles\n\n    def forward(self, x):\n        bn, c = x.shape\n        keys = self.lin_key(x)  #bn, d\n        weights = self.lin_w(torch.tanh(keys))   #bn, 1\n        weights = weights.reshape(-1, self.n_tiles)   #b, n\n        weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n        return weights\n```\n\nThen in the head you simply do `pooled = torch.matmul(x.transpose(1, 2), weights).squeeze(2)` where x is a b, n, c_in tensor representing the bags of tiles (output of CNN).",
          "votes": 5
        },
        {
          "id": 860041,
          "postDate": "2020-05-25T01:27:32.523Z",
          "content": "<p><a href=\"/analokamus\">@analokamus</a>  \" Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc\"</p>\n\n<p>how to test if attention pooling works?</p>\n\n<p>for the following experiments, you can use kaggle data, synthetic data or other opensource benchmark data(e.g. those that used in paper)\n0. assume you have 2 set of data : train and validation. for each image split into tiles (instance).\n1. baseline one: train on all instances \n2. baseline two: train on hand selected instances (e.g. based on mask). we assume human is the perfect model that can select the best discriminating instance. But this require prior knowledge or extra labels. (for this reason, it is sometime best to use synthetic or data that proved to work in paper)\n3. baseline three: train random instances (or equal weighing in pooling)\n4. now develop your attention-base model or instance based model and get results</p>\n\n<p>now if results2 is worse than results1, attention-based/multi-intsance based methods probably will not work. there may not exist subset of instances that perform better than all instances. a thumb of rule: \"performance of deep net is sightly better than human model\"</p>\n\n<p>if results2 is better than results1, then a well workflow of instance-based learning should make results4 better than results3. it is easier to debug the algorithm in  workflow4 since you have the ground truth best instances from workflow2. \"did your model select random instance (like in 3) or discrminative instance like the human(2)\"?</p>\n\n<p>once results4 is better than results3, then we can finetune and improve the algorithm the make results4 close to or better than results2.</p>",
          "rawMarkdown": "@analokamus  \" Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc\"\n\nhow to test if attention pooling works?\n\n\nfor the following experiments, you can use kaggle data, synthetic data or other opensource benchmark data(e.g. those that used in paper)\n0. assume you have 2 set of data : train and validation. for each image split into tiles (instance).\n1. baseline one: train on all instances \n2. baseline two: train on hand selected instances (e.g. based on mask). we assume human is the perfect model that can select the best discriminating instance. But this require prior knowledge or extra labels. (for this reason, it is sometime best to use synthetic or data that proved to work in paper)\n3. baseline three: train random instances (or equal weighing in pooling)\n4. now develop your attention-base model or instance based model and get results\n\nnow if results2 is worse than results1, attention-based/multi-intsance based methods probably will not work. there may not exist subset of instances that perform better than all instances. a thumb of rule: \"performance of deep net is sightly better than human model\"\n\nif results2 is better than results1, then a well workflow of instance-based learning should make results4 better than results3. it is easier to debug the algorithm in  workflow4 since you have the ground truth best instances from workflow2. \"did your model select random instance (like in 3) or discrminative instance like the human(2)\"?\n\nonce results4 is better than results3, then we can finetune and improve the algorithm the make results4 close to or better than results2.\n",
          "votes": 1
        },
        {
          "id": 860067,
          "postDate": "2020-05-25T02:12:49.073Z",
          "content": "<p>Very nice <a href=\"/arroqc\">@arroqc</a>, I've never tried such polling method! I will try to test on my model and see what are the results!</p>\n\n<p>Just a question, what does AdaptiveConcatPool2d() do in your code? This might be a stupid question haha</p>",
          "rawMarkdown": "Very nice @arroqc, I've never tried such polling method! I will try to test on my model and see what are the results!\n\nJust a question, what does AdaptiveConcatPool2d() do in your code? This might be a stupid question haha"
        },
        {
          "id": 860608,
          "postDate": "2020-05-25T13:14:38.590Z",
          "content": "<p>Sorry, that's some leftover code when I extracted the attnetion from the head module. It's not doing anything for attention, this is the classic concat pooling that I do on a per tile basis before attention. Going from BxN, C, H, W to BxN, C.</p>\n\n<p>So with attention I do: Pool per tile, Attention Pool, FC instead of Reshape, Pool for the bag, FC as in iafoss public kernel.</p>",
          "rawMarkdown": "Sorry, that's some leftover code when I extracted the attnetion from the head module. It's not doing anything for attention, this is the classic concat pooling that I do on a per tile basis before attention. Going from BxN, C, H, W to BxN, C.\n\nSo with attention I do: Pool per tile, Attention Pool, FC instead of Reshape, Pool for the bag, FC as in iafoss public kernel."
        },
        {
          "id": 860653,
          "postDate": "2020-05-25T13:50:03.007Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> ok so right after extracting features you do: pooling (AdaptiveConcatPool2d) and use attention method to weight each tile and then do the concat pooling like iafoss did? sorry im trying to understand :)</p>",
          "rawMarkdown": "@arroqc ok so right after extracting features you do: pooling (AdaptiveConcatPool2d) and use attention method to weight each tile and then do the concat pooling like iafoss did? sorry im trying to understand :)"
        },
        {
          "id": 861029,
          "postDate": "2020-05-25T19:28:55.297Z",
          "content": "<p>Almost but no. \nYou start with a B * N, C, H, W tensor after the CNN.\nThen you need to somehow get a representation per instance. You have multiple ways:\n* Max/Mean pooling (my suggestion)\n* A last convolution of H,W kernel size.</p>\n\n<p>So now you have a B*N, C tensor let's call it X and reshape it to B, N, C. Basically for each tile you have a vector C.</p>\n\n<p>Now you pass that to the attention layer I've given you. This will output WEIGHTS in a B, N (B, N, 1 actually) tensor. So you need to now compute the attended view of the bag by doing a weighted average of X according to the attention weights. This is done as matmul(x.transpose(1, 2), weights). You can try to convince yourself this will really give the weighted average according to weights.</p>\n\n<p>The result will be a B, C, 1 tensor. This is a bag level representation. You now can remove dim=2 and pass it to some fully connected network. Note that if you use average-max concatenation you will often have to deal with 2C instead of C. Here is the full head I use if it makes it easier for you to follow:\n```\nclass AttentionPoolHead(nn.Module):</p>\n\n<pre><code>def __init__(self, c_in, c_out, n_tiles):\n    super().__init__()\n    self.maxpool = AdaptiveConcatPool2d()\n    self.lin_key = nn.Linear(c_in * 2, c_in // 2)\n    self.lin_w = nn.Linear(c_in // 2, 1)\n    self.n_tiles = n_tiles\n    self.fc = nn.Sequential(nn.Dropout(0.5),\n                            nn.Linear(c_in * 2, 512),\n                            Mish(),\n                            nn.BatchNorm1d(512),\n                            nn.Dropout(0.5),\n                            nn.Linear(512, c_out))\n\ndef compute_attention(self, x):\n    keys = self.lin_key(x)\n    weights = self.lin_w(torch.tanh(keys))\n    weights = weights.reshape(-1, self.n_tiles)\n    weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n    return weights\n\ndef forward(self, x):\n    bn, c, h, w = x.shape\n    h = self.maxpool(x).squeeze(2).squeeze(2)  # bn, c\n    weights = self.compute_attention(h)\n    h = h.reshape(-1, self.n_tiles, c * 2)\n    h = h.transpose(1, 2)\n    pooled = torch.matmul(h, weights).squeeze(2)\n    return self.fc(pooled)\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Almost but no. \nYou start with a B * N, C, H, W tensor after the CNN.\nThen you need to somehow get a representation per instance. You have multiple ways:\n* Max/Mean pooling (my suggestion)\n* A last convolution of H,W kernel size.\n\nSo now you have a B*N, C tensor let's call it X and reshape it to B, N, C. Basically for each tile you have a vector C.\n\nNow you pass that to the attention layer I've given you. This will output WEIGHTS in a B, N (B, N, 1 actually) tensor. So you need to now compute the attended view of the bag by doing a weighted average of X according to the attention weights. This is done as matmul(x.transpose(1, 2), weights). You can try to convince yourself this will really give the weighted average according to weights.\n\nThe result will be a B, C, 1 tensor. This is a bag level representation. You now can remove dim=2 and pass it to some fully connected network. Note that if you use average-max concatenation you will often have to deal with 2C instead of C. Here is the full head I use if it makes it easier for you to follow:\n```\nclass AttentionPoolHead(nn.Module):\n\n    def __init__(self, c_in, c_out, n_tiles):\n        super().__init__()\n        self.maxpool = AdaptiveConcatPool2d()\n        self.lin_key = nn.Linear(c_in * 2, c_in // 2)\n        self.lin_w = nn.Linear(c_in // 2, 1)\n        self.n_tiles = n_tiles\n        self.fc = nn.Sequential(nn.Dropout(0.5),\n                                nn.Linear(c_in * 2, 512),\n                                Mish(),\n                                nn.BatchNorm1d(512),\n                                nn.Dropout(0.5),\n                                nn.Linear(512, c_out))\n\n    def compute_attention(self, x):\n        keys = self.lin_key(x)\n        weights = self.lin_w(torch.tanh(keys))\n        weights = weights.reshape(-1, self.n_tiles)\n        weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n        return weights\n\n    def forward(self, x):\n        bn, c, h, w = x.shape\n        h = self.maxpool(x).squeeze(2).squeeze(2)  # bn, c\n        weights = self.compute_attention(h)\n        h = h.reshape(-1, self.n_tiles, c * 2)\n        h = h.transpose(1, 2)\n        pooled = torch.matmul(h, weights).squeeze(2)\n        return self.fc(pooled)\n```",
          "votes": 5
        },
        {
          "id": 861080,
          "postDate": "2020-05-25T20:35:53.073Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Merci Arnaud, j'apprécie vrm ton explication, j'ai bcp appris! Je viens de voir que tu viens de Montréal tout comme moi nice!</p>",
          "rawMarkdown": "@arroqc Merci Arnaud, j'apprécie vrm ton explication, j'ai bcp appris! Je viens de voir que tu viens de Montréal tout comme moi nice!",
          "votes": 1
        },
        {
          "id": 861103,
          "postDate": "2020-05-25T21:07:29.380Z",
          "content": "<p>Let me know if it improves your results :) Or at least if like me it's approx the same as a simple Avg/max pooling like iafoss.</p>",
          "rawMarkdown": "Let me know if it improves your results :) Or at least if like me it's approx the same as a simple Avg/max pooling like iafoss."
        },
        {
          "id": 862343,
          "postDate": "2020-05-26T14:48:44.187Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Unfortunately it did not give me better results! I learned a lot though thanks! :) </p>",
          "rawMarkdown": "@arroqc Unfortunately it did not give me better results! I learned a lot though thanks! :) "
        },
        {
          "id": 862350,
          "postDate": "2020-05-26T14:52:35.437Z",
          "content": "<p>Did it get worse or similar ?</p>",
          "rawMarkdown": "Did it get worse or similar ?"
        },
        {
          "id": 862360,
          "postDate": "2020-05-26T14:58:59.757Z",
          "content": "<p>A little worse, i think its because of the number and size of my tiles, not sure though! Also i don't usually use AdaptiveConcatPool2d in my other models, might be an other reason</p>",
          "rawMarkdown": "A little worse, i think its because of the number and size of my tiles, not sure though! Also i don't usually use AdaptiveConcatPool2d in my other models, might be an other reason"
        },
        {
          "id": 868427,
          "postDate": "2020-05-31T08:19:22.763Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> thanks for your explanations, the idea is quite attractive indeed. Did it improve the score so far ?  I'm pretty sure there's something to do with bagging but my intuition is that we need much more tiles to make it work properly </p>",
          "rawMarkdown": "@arroqc thanks for your explanations, the idea is quite attractive indeed. Did it improve the score so far ?  I'm pretty sure there's something to do with bagging but my intuition is that we need much more tiles to make it work properly "
        },
        {
          "id": 869016,
          "postDate": "2020-05-31T16:39:49.400Z",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> No I get approximately same performance for now. Although I still think it could be useful. I'm exploring some ideas that either will pay off or will have been a huge time sink :D</p>",
          "rawMarkdown": "@alexj21 No I get approximately same performance for now. Although I still think it could be useful. I'm exploring some ideas that either will pay off or will have been a huge time sink :D"
        }
      ]
    },
    {
      "id": 868012,
      "postDate": "2020-05-30T20:16:36.067Z",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> I am a beginner, so there might be stupid questions. Why are you using softmax activation for attention-based pooling? It looks more natural for me to use sigmoid function, because our goal is jsut to assing weight to the output. The softmax function also introduces weights, however, now weights sum to 1. What't the purpose of this?</p>",
      "rawMarkdown": "@arroqc I am a beginner, so there might be stupid questions. Why are you using softmax activation for attention-based pooling? It looks more natural for me to use sigmoid function, because our goal is jsut to assing weight to the output. The softmax function also introduces weights, however, now weights sum to 1. What't the purpose of this?",
      "replies": [
        {
          "id": 868148,
          "postDate": "2020-05-31T01:14:49.843Z",
          "content": "<p><a href=\"/xtonny13\">@xtonny13</a> Softmax is standard for attention. Because when you do the weighted average you want normalized weights. If you don't then your average will have varrying degrees of magnitude depending on the elements present in the bag... Let's say you have tiles that are all cancer. You expect your attention score to get high for all of them. Softmax will therefore give weights like an average. If you send only one cancer tile and the rest is garbage on the other hand then you will get a weight of almost 1 for the cancer and 0 elsewhere. Both example will therefore give a similar pooled representation.</p>\n\n<p>If you use a sigmoid or any other un-normalized weights you will end up with 2 very different representations for these 2 examples.</p>",
          "rawMarkdown": "@xtonny13 Softmax is standard for attention. Because when you do the weighted average you want normalized weights. If you don't then your average will have varrying degrees of magnitude depending on the elements present in the bag... Let's say you have tiles that are all cancer. You expect your attention score to get high for all of them. Softmax will therefore give weights like an average. If you send only one cancer tile and the rest is garbage on the other hand then you will get a weight of almost 1 for the cancer and 0 elsewhere. Both example will therefore give a similar pooled representation.\n\nIf you use a sigmoid or any other un-normalized weights you will end up with 2 very different representations for these 2 examples."
        },
        {
          "id": 868544,
          "postDate": "2020-05-31T10:20:16.900Z",
          "content": "<p>Oh, I see, thank you!</p>",
          "rawMarkdown": "Oh, I see, thank you!"
        }
      ]
    },
    {
      "id": 862611,
      "postDate": "2020-05-26T17:29:55.920Z",
      "content": "<p>Does anyone got improved results based on \n(1) Attention based pooling or \n(2) Applying a selection network first to identify the 'important' tiles and then train the network on these selected tiles?</p>",
      "rawMarkdown": "Does anyone got improved results based on \n(1) Attention based pooling or \n(2) Applying a selection network first to identify the 'important' tiles and then train the network on these selected tiles?\n\n"
    },
    {
      "id": 859076,
      "postDate": "2020-05-24T05:29:20.717Z",
      "rawMarkdown": ""
    },
    {
      "id": 854628,
      "postDate": "2020-05-20T06:55:34.190Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>Could I have 2 questions if you don't mind to share</p>\n\n<ol>\n<li><p>How did you combine attention maps (or attention weighting) with CNN feature map? Because original image shapes are not fixed, CNN feature map shape are fixed.</p></li>\n<li><p>What is attention based pooling? Is it normal pooling?</p></li>\n</ol>\n\n<p>Thank you</p>",
      "rawMarkdown": "@hengck23 \n\nCould I have 2 questions if you don't mind to share\n\n1. How did you combine attention maps (or attention weighting) with CNN feature map? Because original image shapes are not fixed, CNN feature map shape are fixed.\n\n2. What is attention based pooling? Is it normal pooling?\n\nThank you",
      "replies": [
        {
          "id": 855462,
          "postDate": "2020-05-21T00:23:55.177Z",
          "content": "<p>Attention pooling is when you take each embedding of every element (tile) of the bag (the collection of tile) and compute a single embedding as a weighted sum of all others with the weights being attention weights. The weights come from a softmax normalization of some scores computed for each tile. It's an alternative to max and mean pooling. I implemented it in my code based on the reference below but only got a small improvement so far. Which is not a surprise since the authors only note a small improvement compared to mean OR max pooling and I was using mean + max pooling before.</p>\n\n<p>Here is a reference <a href=\"https://arxiv.org/pdf/1802.04712.pdf\">https://arxiv.org/pdf/1802.04712.pdf</a></p>\n\n<p>If you get your tile like in the deep attention sampling paper, you can directly use the attention values for your attention pooling.</p>",
          "rawMarkdown": "Attention pooling is when you take each embedding of every element (tile) of the bag (the collection of tile) and compute a single embedding as a weighted sum of all others with the weights being attention weights. The weights come from a softmax normalization of some scores computed for each tile. It's an alternative to max and mean pooling. I implemented it in my code based on the reference below but only got a small improvement so far. Which is not a surprise since the authors only note a small improvement compared to mean OR max pooling and I was using mean + max pooling before.\n\nHere is a reference https://arxiv.org/pdf/1802.04712.pdf\n\nIf you get your tile like in the deep attention sampling paper, you can directly use the attention values for your attention pooling.",
          "votes": 1
        },
        {
          "id": 855640,
          "postDate": "2020-05-21T04:47:43.823Z",
          "content": "<p>Thank you for your detailed information!\nAppreciate it</p>",
          "rawMarkdown": "Thank you for your detailed information!\nAppreciate it"
        }
      ]
    },
    {
      "id": 853804,
      "postDate": "2020-05-19T13:55:33.793Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> \nas usual u post great stuff\n1) what loss do you use ?\n2) Do you know about the issue with the Pen Marked images,what issue would they cause in learning.</p>",
      "rawMarkdown": "@hengck23 \nas usual u post great stuff\n1) what loss do you use ?\n2) Do you know about the issue with the Pen Marked images,what issue would they cause in learning.\n\n"
    },
    {
      "id": 853721,
      "postDate": "2020-05-19T12:38:56.970Z",
      "content": "<p>good!</p>",
      "rawMarkdown": "good!"
    },
    {
      "id": 853543,
      "postDate": "2020-05-19T09:15:36.653Z",
      "content": "<p>As far as I understand, you still train your first stage model on tiles. So you have to make a selection beforehand anyways. Am I correct ?</p>",
      "rawMarkdown": "As far as I understand, you still train your first stage model on tiles. So you have to make a selection beforehand anyways. Am I correct ?",
      "replies": [
        {
          "id": 853711,
          "postDate": "2020-05-19T12:29:02.553Z",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> </p>\n\n<p>no. the selection is automatic. the challenge is how to make it \"learnable\", i.e. what is the supervisory signal?</p>\n\n<p>it can be weakly supervised using ground truth mask(noisy) or self-supervised using image label.</p>\n\n<p>i will put some papers that can do this later.</p>",
          "rawMarkdown": "@bdubreu \n\nno. the selection is automatic. the challenge is how to make it \"learnable\", i.e. what is the supervisory signal?\n\nit can be weakly supervised using ground truth mask(noisy) or self-supervised using image label.\n\ni will put some papers that can do this later.",
          "votes": 1
        },
        {
          "id": 853971,
          "postDate": "2020-05-19T16:37:25.153Z",
          "content": "<p>But then how do you propose to feed the data to the first model, since the input shapes vary so much ? </p>",
          "rawMarkdown": "But then how do you propose to feed the data to the first model, since the input shapes vary so much ? ",
          "votes": 1
        },
        {
          "id": 854363,
          "postDate": "2020-05-20T01:48:24.700Z",
          "content": "<p>it will be like object detection like faster-rcnn. each input image is of difference size in training.</p>\n\n<p>or you can pad all image to a good size.</p>",
          "rawMarkdown": "it will be like object detection like faster-rcnn. each input image is of difference size in training.\n\nor you can pad all image to a good size."
        },
        {
          "id": 855039,
          "postDate": "2020-05-20T14:23:48.633Z",
          "content": "<p>I was thinking about padding all the images to a good overall shape too, but after a short investigation I realized the idea was probably a dead-end: the aspect-ratios are all over the place, even when you crop everything beforehand. \n\"it will be like object detection like faster-rcnn\". I am too newbie to know this. I'll investigate how this works. If some models can indeed take different shapes inputs then that's actually promising</p>",
          "rawMarkdown": "I was thinking about padding all the images to a good overall shape too, but after a short investigation I realized the idea was probably a dead-end: the aspect-ratios are all over the place, even when you crop everything beforehand. \n\"it will be like object detection like faster-rcnn\". I am too newbie to know this. I'll investigate how this works. If some models can indeed take different shapes inputs then that's actually promising"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 853715,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-05-19T12:31:20.687000",
      "content": "<p>\"Processing Megapixel Images with Deep Attention-Sampling Models\"</p>\n\n<p><a href=\"https://arxiv.org/pdf/1905.03711.pdf\">https://arxiv.org/pdf/1905.03711.pdf</a>\n<a href=\"https://www.youtube.com/watch?v=H6Qiegq_36c\">https://www.youtube.com/watch?v=H6Qiegq_36c</a>\n<a href=\"https://github.com/idiap/attention-sampling\">https://github.com/idiap/attention-sampling</a></p>\n\n<p>abstract:\nExisting deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a*<em>fully differentiable end-to-end trainable</em>* model that samples and processes only a fraction of the full resolution input image.</p>\n\n<p>The <strong>locations to process are sampled from an attention distribution computed from a low resolution view of the input.</strong> We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 853717,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-19T12:33:26.017000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1ad9553af04fb8b2718ed6d960aaa8c5%2F694eed8d9e67cb1aa2d8ee410a5c97d0a1f36714.png?generation=1589956576501217&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 853720,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-19T12:37:39.603000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F00a8df7fbc3a8dbbbaa42c271edf8f0e%2FSelection_156.png?generation=1589891957077940&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 853738,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-19T12:54:13.103000",
          "content": "<p>related:\nStreaming convolutional neural networks for end-to-end learning with multi-megapixel images\n- Hans Pinckaers (Computational Pathology Group, Radboud University Medical Center)  </p>\n\n<p><a href=\"https://github.com/DIAGNijmegen/StreamingCNN\">https://github.com/DIAGNijmegen/StreamingCNN</a>\n<a href=\"https://www.computationalpathologygroup.eu/software/automated-gleason-grading/\">https://www.computationalpathologygroup.eu/software/automated-gleason-grading/</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 853845,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-19T14:36:00.993000",
          "content": "<p>Maybe you can use the masks to pretrain your attention network.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 854944,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-20T13:15:07.820000",
          "content": "<p>pytorch version! <a href=\"https://github.com/idiap/attention-sampling/issues/1\">https://github.com/idiap/attention-sampling/issues/1</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 854734,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2020-05-20T09:17:50.743000",
      "content": "<p>Thanks a lot for the references. \nI've also thought about improving iafoss tiling method and created a very basic kernel which uses a hybrid CNN-LSTM in order to learn long-term dependencies between features <a href=\"https://www.kaggle.com/alexj21/hybrid-cnn-lstm-starter\">here</a>\nHope it can provide some ideas too</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 853339,
      "author_name": "RabotniKuma",
      "author_url": "",
      "post_date": "2020-05-19T05:29:57.310000",
      "content": "<p>Links to reference papers:\n<a href=\"https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982\">https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982</a>\n<a href=\"https://arxiv.org/abs/1905.13208\">https://arxiv.org/abs/1905.13208</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 855472,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-05-21T00:41:51.673000",
      "content": "<p>Those approaches look exiting, but there may be an issue related to the nature of the labels: they are biased on the relative area. It would be interesting to think how to measure the area not having nearly all tissue tiles as an input. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 857925,
      "author_name": "RabotniKuma",
      "author_url": "",
      "post_date": "2020-05-23T04:07:23.953000",
      "content": "<p>I tried <code>attention based pooling</code> mentioned in <a href=\"https://jmtomczak.github.io/pdf/Tooploox_2018_06_27.pdf\">this paper</a>. Compared to my previous <code>concat pooling</code> approach (almost identical to @lafoss 's notebook), it didn't give any boost in both CV and LB. Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc. Anyone managed to make it work?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 858912,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-23T23:59:20.787000",
          "content": "<p>I just realized in <a href=\"/iafoss\">@iafoss</a> notebook what AdaptiveConcatPool2d does... I have just been using torch.mean on the feature vectors from the cnn backbone. Have you tried just using max or mean pooling and how much difference it makes?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 858918,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-05-24T00:21:46.913000",
          "content": "<p>Based on my experiment with different pooling layers so far <code>GeM</code> pooling layers works the best. </p>\n\n<p>What is GeM?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>)</p>\n\n<p>code below:</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 858962,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-24T01:57:10.480000",
          "content": "<p>Thanks for the tip. I will be trying this out</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 858982,
          "author_name": "RabotniKuma",
          "author_url": "",
          "post_date": "2020-05-24T02:50:18.297000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you! I'll give it a try.</p>\n\n<p><a href=\"/shujun717\">@shujun717</a> I tried avg pool and concat pool. They both performed almost the same .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 859922,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-05-24T21:14:39.157000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> from what I understand AdaptiveConcatPool2d is no different from concating the mean and max of the set of feature vectors produced by the cnn in your notebook correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 859937,
          "author_name": "Yousef Rabi",
          "author_url": "",
          "post_date": "2020-05-24T21:31:35.180000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a>  Yes. Precisely. It's taking both results of mean and max and concatenating them together. I think in this competition it wouldn't make that much of a difference because we are likely to be more interested in the results of the max pool (was there cancer or not? or what cancer was it?) as opposed to the average cancerness of the whole image. But in cases where we care about both, for example, what was the maximum haze in an image AND what was the overall haziness, then I guess concatenating both might give us more benefit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 859941,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-24T21:36:20.600000",
          "content": "<p><a href=\"/yousof9\">@yousof9</a> I would actually expect average to work better than max here considering grades are determined by the relative presence of gleason pattern in the sample. The average for that purpose is a better representation of the bag content.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 859969,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-05-24T22:45:11.923000",
          "content": "<p>Actually I would expect that for this particular labels both max and mean are important: mean gives the relative area of different patterns while max gives the worst cancer pattern. If you check the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/overview/additional-resources\">additional clarification</a> provided for this competition by organizers, you will see that the first Gleason score is based on the pattern having the maximum area, while the second corresponds to the worst pattern, if its score is higher than one for the most common pattern, or the second most common pattern having more than 5% of the tissue area. If the second pattern is not present, the second Gleason score is set to be equal to the first one, if I understand it correctly.</p>\n\n<p>For this particular labels in my tests generalized GeM pooling (the value of p is selected individually for each channel) worked worse than concatenate pooling. It's a little bit surprising since in last two competitions I participated in GeM worked better (though the difference was small in comparison with concat), but I'd think it is resulted by the way how labels are selected: based on mean and max values.</p>\n\n<p>And regarding attention pooling (CBAM block added right before pooling), it actually never worked for me. I tried it in 3-4 competitions I participated in during the last year without much success. Probably, the model needs to be trained in some special manner, but somehow I couldn't make it work better than just simple concat or GeM pooling. From the first glance I wd think that attention applied right before pooling could really boost the results, but the reality is a little bit counterintuitive. Probably, some aux added to the attention could guide the model into the right direction, I don't know.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 860001,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-24T23:57:54.957000",
          "content": "<p>True. Actually I use both anyway.</p>\n\n<p>Rergarding attention pooling I just developpe my own based on a paper above. I get similar performance as with concat pooling. Main advantage is that it gives some interpretability as it shows which tiles it is focusing on. But I don't really understand what cancer really looks like so I'm not sure it's doing the right thing actually :D</p>\n\n<p>Expecting a BN, C tensor here is my attention based on Deep MIL paper.\n```\nclass AttentionPool(nn.Module):</p>\n\n<pre><code>def __init__(self, c_in, d, n_tiles):\n    super().__init__()\n    self.maxpool = AdaptiveConcatPool2d() # Edit: Please remove this\n    self.lin_key = nn.Linear(c_in, d)\n    self.lin_w = nn.Linear(d, 1)\n    self.n_tiles = n_tiles\n\ndef forward(self, x):\n    bn, c = x.shape\n    keys = self.lin_key(x)  #bn, d\n    weights = self.lin_w(torch.tanh(keys))   #bn, 1\n    weights = weights.reshape(-1, self.n_tiles)   #b, n\n    weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n    return weights\n</code></pre>\n\n<p>```</p>\n\n<p>Then in the head you simply do <code>pooled = torch.matmul(x.transpose(1, 2), weights).squeeze(2)</code> where x is a b, n, c_in tensor representing the bags of tiles (output of CNN).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 860041,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-25T01:27:32.523000",
          "content": "<p><a href=\"/analokamus\">@analokamus</a>  \" Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc\"</p>\n\n<p>how to test if attention pooling works?</p>\n\n<p>for the following experiments, you can use kaggle data, synthetic data or other opensource benchmark data(e.g. those that used in paper)\n0. assume you have 2 set of data : train and validation. for each image split into tiles (instance).\n1. baseline one: train on all instances \n2. baseline two: train on hand selected instances (e.g. based on mask). we assume human is the perfect model that can select the best discriminating instance. But this require prior knowledge or extra labels. (for this reason, it is sometime best to use synthetic or data that proved to work in paper)\n3. baseline three: train random instances (or equal weighing in pooling)\n4. now develop your attention-base model or instance based model and get results</p>\n\n<p>now if results2 is worse than results1, attention-based/multi-intsance based methods probably will not work. there may not exist subset of instances that perform better than all instances. a thumb of rule: \"performance of deep net is sightly better than human model\"</p>\n\n<p>if results2 is better than results1, then a well workflow of instance-based learning should make results4 better than results3. it is easier to debug the algorithm in  workflow4 since you have the ground truth best instances from workflow2. \"did your model select random instance (like in 3) or discrminative instance like the human(2)\"?</p>\n\n<p>once results4 is better than results3, then we can finetune and improve the algorithm the make results4 close to or better than results2.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 860067,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-25T02:12:49.073000",
          "content": "<p>Very nice <a href=\"/arroqc\">@arroqc</a>, I've never tried such polling method! I will try to test on my model and see what are the results!</p>\n\n<p>Just a question, what does AdaptiveConcatPool2d() do in your code? This might be a stupid question haha</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 860608,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-25T13:14:38.590000",
          "content": "<p>Sorry, that's some leftover code when I extracted the attnetion from the head module. It's not doing anything for attention, this is the classic concat pooling that I do on a per tile basis before attention. Going from BxN, C, H, W to BxN, C.</p>\n\n<p>So with attention I do: Pool per tile, Attention Pool, FC instead of Reshape, Pool for the bag, FC as in iafoss public kernel.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 860653,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-25T13:50:03.007000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> ok so right after extracting features you do: pooling (AdaptiveConcatPool2d) and use attention method to weight each tile and then do the concat pooling like iafoss did? sorry im trying to understand :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 861029,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-25T19:28:55.297000",
          "content": "<p>Almost but no. \nYou start with a B * N, C, H, W tensor after the CNN.\nThen you need to somehow get a representation per instance. You have multiple ways:\n* Max/Mean pooling (my suggestion)\n* A last convolution of H,W kernel size.</p>\n\n<p>So now you have a B*N, C tensor let's call it X and reshape it to B, N, C. Basically for each tile you have a vector C.</p>\n\n<p>Now you pass that to the attention layer I've given you. This will output WEIGHTS in a B, N (B, N, 1 actually) tensor. So you need to now compute the attended view of the bag by doing a weighted average of X according to the attention weights. This is done as matmul(x.transpose(1, 2), weights). You can try to convince yourself this will really give the weighted average according to weights.</p>\n\n<p>The result will be a B, C, 1 tensor. This is a bag level representation. You now can remove dim=2 and pass it to some fully connected network. Note that if you use average-max concatenation you will often have to deal with 2C instead of C. Here is the full head I use if it makes it easier for you to follow:\n```\nclass AttentionPoolHead(nn.Module):</p>\n\n<pre><code>def __init__(self, c_in, c_out, n_tiles):\n    super().__init__()\n    self.maxpool = AdaptiveConcatPool2d()\n    self.lin_key = nn.Linear(c_in * 2, c_in // 2)\n    self.lin_w = nn.Linear(c_in // 2, 1)\n    self.n_tiles = n_tiles\n    self.fc = nn.Sequential(nn.Dropout(0.5),\n                            nn.Linear(c_in * 2, 512),\n                            Mish(),\n                            nn.BatchNorm1d(512),\n                            nn.Dropout(0.5),\n                            nn.Linear(512, c_out))\n\ndef compute_attention(self, x):\n    keys = self.lin_key(x)\n    weights = self.lin_w(torch.tanh(keys))\n    weights = weights.reshape(-1, self.n_tiles)\n    weights = torch.softmax(weights, dim=1).unsqueeze(2)  # b, n, 1\n    return weights\n\ndef forward(self, x):\n    bn, c, h, w = x.shape\n    h = self.maxpool(x).squeeze(2).squeeze(2)  # bn, c\n    weights = self.compute_attention(h)\n    h = h.reshape(-1, self.n_tiles, c * 2)\n    h = h.transpose(1, 2)\n    pooled = torch.matmul(h, weights).squeeze(2)\n    return self.fc(pooled)\n</code></pre>\n\n<p>```</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 861080,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-25T20:35:53.073000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Merci Arnaud, j'apprécie vrm ton explication, j'ai bcp appris! Je viens de voir que tu viens de Montréal tout comme moi nice!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 861103,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-25T21:07:29.380000",
          "content": "<p>Let me know if it improves your results :) Or at least if like me it's approx the same as a simple Avg/max pooling like iafoss.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 862343,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-26T14:48:44.187000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Unfortunately it did not give me better results! I learned a lot though thanks! :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 862350,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-26T14:52:35.437000",
          "content": "<p>Did it get worse or similar ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 862360,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-05-26T14:58:59.757000",
          "content": "<p>A little worse, i think its because of the number and size of my tiles, not sure though! Also i don't usually use AdaptiveConcatPool2d in my other models, might be an other reason</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 868427,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-05-31T08:19:22.763000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> thanks for your explanations, the idea is quite attractive indeed. Did it improve the score so far ?  I'm pretty sure there's something to do with bagging but my intuition is that we need much more tiles to make it work properly </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869016,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-31T16:39:49.400000",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> No I get approximately same performance for now. Although I still think it could be useful. I'm exploring some ideas that either will pay off or will have been a huge time sink :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868012,
      "author_name": "AntonZ",
      "author_url": "",
      "post_date": "2020-05-30T20:16:36.067000",
      "content": "<p><a href=\"/arroqc\">@arroqc</a> I am a beginner, so there might be stupid questions. Why are you using softmax activation for attention-based pooling? It looks more natural for me to use sigmoid function, because our goal is jsut to assing weight to the output. The softmax function also introduces weights, however, now weights sum to 1. What't the purpose of this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 868148,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-31T01:14:49.843000",
          "content": "<p><a href=\"/xtonny13\">@xtonny13</a> Softmax is standard for attention. Because when you do the weighted average you want normalized weights. If you don't then your average will have varrying degrees of magnitude depending on the elements present in the bag... Let's say you have tiles that are all cancer. You expect your attention score to get high for all of them. Softmax will therefore give weights like an average. If you send only one cancer tile and the rest is garbage on the other hand then you will get a weight of almost 1 for the cancer and 0 elsewhere. Both example will therefore give a similar pooled representation.</p>\n\n<p>If you use a sigmoid or any other un-normalized weights you will end up with 2 very different representations for these 2 examples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 868544,
          "author_name": "AntonZ",
          "author_url": "",
          "post_date": "2020-05-31T10:20:16.900000",
          "content": "<p>Oh, I see, thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 862611,
      "author_name": "DeepLearner",
      "author_url": "",
      "post_date": "2020-05-26T17:29:55.920000",
      "content": "<p>Does anyone got improved results based on \n(1) Attention based pooling or \n(2) Applying a selection network first to identify the 'important' tiles and then train the network on these selected tiles?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 859076,
      "author_name": "Ashish Jangra",
      "author_url": "",
      "post_date": "2020-05-24T05:29:20.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 854628,
      "author_name": "YS",
      "author_url": "",
      "post_date": "2020-05-20T06:55:34.190000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>Could I have 2 questions if you don't mind to share</p>\n\n<ol>\n<li><p>How did you combine attention maps (or attention weighting) with CNN feature map? Because original image shapes are not fixed, CNN feature map shape are fixed.</p></li>\n<li><p>What is attention based pooling? Is it normal pooling?</p></li>\n</ol>\n\n<p>Thank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 855462,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-21T00:23:55.177000",
          "content": "<p>Attention pooling is when you take each embedding of every element (tile) of the bag (the collection of tile) and compute a single embedding as a weighted sum of all others with the weights being attention weights. The weights come from a softmax normalization of some scores computed for each tile. It's an alternative to max and mean pooling. I implemented it in my code based on the reference below but only got a small improvement so far. Which is not a surprise since the authors only note a small improvement compared to mean OR max pooling and I was using mean + max pooling before.</p>\n\n<p>Here is a reference <a href=\"https://arxiv.org/pdf/1802.04712.pdf\">https://arxiv.org/pdf/1802.04712.pdf</a></p>\n\n<p>If you get your tile like in the deep attention sampling paper, you can directly use the attention values for your attention pooling.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 855640,
          "author_name": "YS",
          "author_url": "",
          "post_date": "2020-05-21T04:47:43.823000",
          "content": "<p>Thank you for your detailed information!\nAppreciate it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 853804,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-05-19T13:55:33.793000",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> \nas usual u post great stuff\n1) what loss do you use ?\n2) Do you know about the issue with the Pen Marked images,what issue would they cause in learning.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 853721,
      "author_name": "Ali krg",
      "author_url": "",
      "post_date": "2020-05-19T12:38:56.970000",
      "content": "<p>good!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 853543,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2020-05-19T09:15:36.653000",
      "content": "<p>As far as I understand, you still train your first stage model on tiles. So you have to make a selection beforehand anyways. Am I correct ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 853711,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-19T12:29:02.553000",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> </p>\n\n<p>no. the selection is automatic. the challenge is how to make it \"learnable\", i.e. what is the supervisory signal?</p>\n\n<p>it can be weakly supervised using ground truth mask(noisy) or self-supervised using image label.</p>\n\n<p>i will put some papers that can do this later.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 853971,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-19T16:37:25.153000",
          "content": "<p>But then how do you propose to feed the data to the first model, since the input shapes vary so much ? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 854363,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-05-20T01:48:24.700000",
          "content": "<p>it will be like object detection like faster-rcnn. each input image is of difference size in training.</p>\n\n<p>or you can pad all image to a good size.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 855039,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-20T14:23:48.633000",
          "content": "<p>I was thinking about padding all the images to a good overall shape too, but after a short investigation I realized the idea was probably a dead-end: the aspect-ratios are all over the place, even when you crop everything beforehand. \n\"it will be like object detection like faster-rcnn\". I am too newbie to know this. I'll investigate how this works. If some models can indeed take different shapes inputs then that's actually promising</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "853305": "as attached\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbd5ca75ad58e41520efe7b3ecb587cd1%2FSlide1.png?generation=1589863350980222&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe81ae8bb26f98f33140c85c615d93%2FSlide2.png?generation=1589863399965273&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc894777e08d0a57a4e7bcf52a3ab0fe0%2FSlide3.png?generation=1589863398695926&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fee0d3d1eb9cf0e4828dd69d53878a021%2FSlide4.png?generation=1589863402368753&amp;alt=media)\n",
    "853715": "\"Processing Megapixel Images with Deep Attention-Sampling Models\"\n\nhttps://arxiv.org/pdf/1905.03711.pdf\nhttps://www.youtube.com/watch?v=H6Qiegq_36c\nhttps://github.com/idiap/attention-sampling\n\n\nabstract:\nExisting deep architectures cannot operate on very large signals such as megapixel images due to computational and memory constraints. To tackle this limitation, we propose a**fully differentiable end-to-end trainable** model that samples and processes only a fraction of the full resolution input image.\n\nThe **locations to process are sampled from an attention distribution computed from a low resolution view of the input.** We refer to our method as attention sampling and it can process images of several megapixels with a standard single GPU setup.",
    "854734": "Thanks a lot for the references. \nI've also thought about improving iafoss tiling method and created a very basic kernel which uses a hybrid CNN-LSTM in order to learn long-term dependencies between features [here](https://www.kaggle.com/alexj21/hybrid-cnn-lstm-starter)\nHope it can provide some ideas too",
    "853339": "Links to reference papers:\nhttps://jamanetwork.com/journals/jamanetworkopen/fullarticle/2753982\nhttps://arxiv.org/abs/1905.13208",
    "855472": "Those approaches look exiting, but there may be an issue related to the nature of the labels: they are biased on the relative area. It would be interesting to think how to measure the area not having nearly all tissue tiles as an input. ",
    "857925": "I tried `attention based pooling` mentioned in [this paper](https://jmtomczak.github.io/pdf/Tooploox_2018_06_27.pdf). Compared to my previous `concat pooling` approach (almost identical to @lafoss 's notebook), it didn't give any boost in both CV and LB. Probably whether it works depends on how many instances(tiles) in a bag(WSI) etc. Anyone managed to make it work?",
    "868012": "@arroqc I am a beginner, so there might be stupid questions. Why are you using softmax activation for attention-based pooling? It looks more natural for me to use sigmoid function, because our goal is jsut to assing weight to the output. The softmax function also introduces weights, however, now weights sum to 1. What't the purpose of this?",
    "862611": "Does anyone got improved results based on \n(1) Attention based pooling or \n(2) Applying a selection network first to identify the 'important' tiles and then train the network on these selected tiles?\n\n",
    "859076": "",
    "854628": "@hengck23 \n\nCould I have 2 questions if you don't mind to share\n\n1. How did you combine attention maps (or attention weighting) with CNN feature map? Because original image shapes are not fixed, CNN feature map shape are fixed.\n\n2. What is attention based pooling? Is it normal pooling?\n\nThank you",
    "853804": "@hengck23 \nas usual u post great stuff\n1) what loss do you use ?\n2) Do you know about the issue with the Pen Marked images,what issue would they cause in learning.\n\n",
    "853721": "good!",
    "853543": "As far as I understand, you still train your first stage model on tiles. So you have to make a selection beforehand anyways. Am I correct ?"
  }
}