{
  "id": 229641,
  "title": "Best No Finding Classifiers",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229641",
  "author_name": "beluga",
  "post_date": "2021-03-31T04:15:55.094000",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>What was your best Normal/Abnormal LB score?</p>\n<p>I have seen lots of AUC CV 0.99+ models but I wonder how they translate on the LB?</p>\n<p>As <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> wrote you can test it by just submitting boxes for 14.</p>\n<blockquote>\n  <p>For each image just add the string 14 PR 0 0 1 1 where PR = prob(no finding) from your classifier. If your classifier is perfect, then you will score full mAP = 1.0 for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP</p>\n</blockquote>\n<p>We tried to build pretty robust classifiers trained on external data, but could not go above 0.064 ~ 0.96 AP</p>",
  "messages": [
    {
      "id": 1257733,
      "postDate": "2021-03-31T04:15:55.093Z",
      "content": "<p>What was your best Normal/Abnormal LB score?</p>\n<p>I have seen lots of AUC CV 0.99+ models but I wonder how they translate on the LB?</p>\n<p>As <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> wrote you can test it by just submitting boxes for 14.</p>\n<blockquote>\n  <p>For each image just add the string 14 PR 0 0 1 1 where PR = prob(no finding) from your classifier. If your classifier is perfect, then you will score full mAP = 1.0 for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP</p>\n</blockquote>\n<p>We tried to build pretty robust classifiers trained on external data, but could not go above 0.064 ~ 0.96 AP</p>",
      "rawMarkdown": "What was your best Normal/Abnormal LB score?\n\n\nI have seen lots of AUC CV 0.99+ models but I wonder how they translate on the LB?\n\nAs @cdeotte wrote you can test it by just submitting boxes for 14.\n\n> For each image just add the string 14 PR 0 0 1 1 where PR = prob(no finding) from your classifier. If your classifier is perfect, then you will score full mAP = 1.0 for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP\n\nWe tried to build pretty robust classifiers trained on external data, but could not go above 0.064 ~ 0.96 AP",
      "votes": 11
    },
    {
      "id": 1258377,
      "postDate": "2021-03-31T15:00:12.807Z",
      "content": "<p>0.065</p>\n<p>we dont use any separate classifier but just multiply the box predictions for an image to get to the class 14 probability<br>\nacross models this is averaged</p>",
      "rawMarkdown": "0.065\n\nwe dont use any separate classifier but just multiply the box predictions for an image to get to the class 14 probability\nacross models this is averaged",
      "votes": 4,
      "replies": [
        {
          "id": 1258409,
          "postDate": "2021-03-31T15:25:39.547Z",
          "content": "<p>0.065, nice job. Smart technique of using object detection bbox scores.</p>",
          "rawMarkdown": "0.065, nice job. Smart technique of using object detection bbox scores.",
          "votes": 2
        },
        {
          "id": 1258410,
          "postDate": "2021-03-31T15:26:59.470Z",
          "content": "<p>Do you train your object detection with all images both with and without finding?</p>",
          "rawMarkdown": "Do you train your object detection with all images both with and without finding?"
        },
        {
          "id": 1258416,
          "postDate": "2021-03-31T15:31:02.747Z",
          "content": "<p>Yes we train on all images</p>",
          "rawMarkdown": "Yes we train on all images",
          "votes": 1
        },
        {
          "id": 1259241,
          "postDate": "2021-04-01T08:45:58.020Z",
          "content": "<p>Did you calculate the 14 class probability something like this?<br>\nprob(class_14) = product(1-<em>p_i</em>) for <em>i</em> in [0, 13], where <em>p_i</em> is the maximum of box score for class <em>i</em>.</p>\n<p>And why didn't you want to train a separate binary classifier? Is it because the result from this calculation was good enough?</p>",
          "rawMarkdown": "Did you calculate the 14 class probability something like this?\nprob(class_14) = product(1-*p_i*) for *i* in [0, 13], where *p_i* is the maximum of box score for class *i*.\n\nAnd why didn't you want to train a separate binary classifier? Is it because the result from this calculation was good enough?",
          "votes": 3
        },
        {
          "id": 1259248,
          "postDate": "2021-04-01T08:58:04.620Z",
          "content": "<p>Basically this, but instead of taking the max box score of each class we just multiplied all (1-p) box probabilities for that image.</p>\n<p>Yes, there was no need for a second 2-stage classifier as this was close to perfection already.</p>",
          "rawMarkdown": "Basically this, but instead of taking the max box score of each class we just multiplied all (1-p) box probabilities for that image.\n\nYes, there was no need for a second 2-stage classifier as this was close to perfection already.",
          "votes": 4
        },
        {
          "id": 1259262,
          "postDate": "2021-04-01T09:15:38.293Z",
          "content": "<p>Ok, got it! Thanks! I think this trick can be useful in another competition! </p>",
          "rawMarkdown": "Ok, got it! Thanks! I think this trick can be useful in another competition! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1257812,
      "postDate": "2021-03-31T05:47:33.720Z",
      "content": "<p>Ours is LB 0.064</p>",
      "rawMarkdown": "Ours is LB 0.064",
      "votes": 2
    },
    {
      "id": 1258695,
      "postDate": "2021-03-31T20:10:34.143Z",
      "content": "<p>In this competition class 14 is not so important. If you get 0.064, that's ok, close to the maximum.</p>",
      "rawMarkdown": "In this competition class 14 is not so important. If you get 0.064, that's ok, close to the maximum.",
      "replies": [
        {
          "id": 1258708,
          "postDate": "2021-03-31T20:21:40.593Z",
          "content": "<p>Yeah I am sure now. Before the LB flip I thought the final supervisiors cleaned and fixed the automated No Findings more. That could have hurt some of the overfitted 2 class filter tricks…</p>",
          "rawMarkdown": "Yeah I am sure now. Before the LB flip I thought the final supervisiors cleaned and fixed the automated No Findings more. That could have hurt some of the overfitted 2 class filter tricks..."
        }
      ]
    },
    {
      "id": 1258417,
      "postDate": "2021-03-31T15:31:53.053Z",
      "content": "<p>I just want to point out that imo these high results for a \"no finding\" binary classifier are representing the most important practical consideration from this competition. This is unseen territory for chest xray in the research litterature. Most reported binary classifiers are usually trained on full image classification labels which consequently has trouble to deal with spatially small but significant abnormalities. These results are directly related to the huge multi-class bbox annotation effort done by Vinbigdata and their associated radiologists.</p>",
      "rawMarkdown": "I just want to point out that imo these high results for a \"no finding\" binary classifier are representing the most important practical consideration from this competition. This is unseen territory for chest xray in the research litterature. Most reported binary classifiers are usually trained on full image classification labels which consequently has trouble to deal with spatially small but significant abnormalities. These results are directly related to the huge multi-class bbox annotation effort done by Vinbigdata and their associated radiologists.",
      "replies": [
        {
          "id": 1258433,
          "postDate": "2021-03-31T15:50:20.557Z",
          "content": "<p>I believe there is quite some bias in the samples in this dataset, which is why this classification is so good. I am not sure it will generalize well to new samples.</p>",
          "rawMarkdown": "I believe there is quite some bias in the samples in this dataset, which is why this classification is so good. I am not sure it will generalize well to new samples.",
          "votes": 2,
          "replies": [
            {
              "id": 1258467,
              "postDate": "2021-03-31T16:23:35.607Z",
              "content": "<p>One clear bias is the selection of PA images only which are inherently better quality images than AP images. I plan to test the results on an external but real clinical distribution to see the robustness. I can give an update when done.</p>",
              "rawMarkdown": "One clear bias is the selection of PA images only which are inherently better quality images than AP images. I plan to test the results on an external but real clinical distribution to see the robustness. I can give an update when done.",
              "votes": 1
            }
          ]
        },
        {
          "id": 1258440,
          "postDate": "2021-03-31T15:58:12.480Z",
          "content": "<p>We tried to eliminated this bias as much as possible. We even trained No Finding models with external data where we removed all the original VinBigData No Finding samples as we did not trust  the 100% agreeable annotators. </p>\n<p>It looks like it did not give much benefit on the private LB.</p>",
          "rawMarkdown": "We tried to eliminated this bias as much as possible. We even trained No Finding models with external data where we removed all the original VinBigData No Finding samples as we did not trust  the 100% agreeable annotators. \n\nIt looks like it did not give much benefit on the private LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1258316,
      "postDate": "2021-03-31T14:27:35.650Z",
      "content": "<p>You mean 0.064 ~ 0.096 AP?  (0.064-0.96 is a pretty wide interval imho ^^)</p>",
      "rawMarkdown": "You mean 0.064 ~ 0.096 AP?  (0.064-0.96 is a pretty wide interval imho ^^)",
      "replies": [
        {
          "id": 1258326,
          "postDate": "2021-03-31T14:34:49.673Z",
          "content": "<p>0.064 score on private public LB for a single class</p>\n<p>Since we have 15 classes  0.064 x 15 = 0.96  ;)</p>",
          "rawMarkdown": "0.064 score on private public LB for a single class\n\nSince we have 15 classes  0.064 x 15 = 0.96  ;)",
          "votes": 1
        },
        {
          "id": 1258354,
          "postDate": "2021-03-31T14:49:46.837Z",
          "content": "<p>Ah, I see, sorry ^^</p>",
          "rawMarkdown": "Ah, I see, sorry ^^"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1258377,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-03-31T15:00:12.807000",
      "content": "<p>0.065</p>\n<p>we dont use any separate classifier but just multiply the box predictions for an image to get to the class 14 probability<br>\nacross models this is averaged</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1258409,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T15:25:39.547000",
          "content": "<p>0.065, nice job. Smart technique of using object detection bbox scores.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1258410,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-03-31T15:26:59.470000",
          "content": "<p>Do you train your object detection with all images both with and without finding?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258416,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-03-31T15:31:02.747000",
          "content": "<p>Yes we train on all images</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259241,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-04-01T08:45:58.020000",
          "content": "<p>Did you calculate the 14 class probability something like this?<br>\nprob(class_14) = product(1-<em>p_i</em>) for <em>i</em> in [0, 13], where <em>p_i</em> is the maximum of box score for class <em>i</em>.</p>\n<p>And why didn't you want to train a separate binary classifier? Is it because the result from this calculation was good enough?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1259248,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-01T08:58:04.620000",
          "content": "<p>Basically this, but instead of taking the max box score of each class we just multiplied all (1-p) box probabilities for that image.</p>\n<p>Yes, there was no need for a second 2-stage classifier as this was close to perfection already.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1259262,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2021-04-01T09:15:38.293000",
          "content": "<p>Ok, got it! Thanks! I think this trick can be useful in another competition! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1257812,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-03-31T05:47:33.720000",
      "content": "<p>Ours is LB 0.064</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1258695,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2021-03-31T20:10:34.143000",
      "content": "<p>In this competition class 14 is not so important. If you get 0.064, that's ok, close to the maximum.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1258708,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-31T20:21:40.593000",
          "content": "<p>Yeah I am sure now. Before the LB flip I thought the final supervisiors cleaned and fixed the automated No Findings more. That could have hurt some of the overfitted 2 class filter tricks…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258417,
      "author_name": "Alexandre Cadrin-Chênevert",
      "author_url": "",
      "post_date": "2021-03-31T15:31:53.053000",
      "content": "<p>I just want to point out that imo these high results for a \"no finding\" binary classifier are representing the most important practical consideration from this competition. This is unseen territory for chest xray in the research litterature. Most reported binary classifiers are usually trained on full image classification labels which consequently has trouble to deal with spatially small but significant abnormalities. These results are directly related to the huge multi-class bbox annotation effort done by Vinbigdata and their associated radiologists.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1258433,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-03-31T15:50:20.557000",
          "content": "<p>I believe there is quite some bias in the samples in this dataset, which is why this classification is so good. I am not sure it will generalize well to new samples.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1258467,
              "author_name": "Alexandre Cadrin-Chênevert",
              "author_url": "",
              "post_date": "2021-03-31T16:23:35.607000",
              "content": "<p>One clear bias is the selection of PA images only which are inherently better quality images than AP images. I plan to test the results on an external but real clinical distribution to see the robustness. I can give an update when done.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1258440,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-31T15:58:12.480000",
          "content": "<p>We tried to eliminated this bias as much as possible. We even trained No Finding models with external data where we removed all the original VinBigData No Finding samples as we did not trust  the 100% agreeable annotators. </p>\n<p>It looks like it did not give much benefit on the private LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1258316,
      "author_name": "AmorfEvo",
      "author_url": "",
      "post_date": "2021-03-31T14:27:35.650000",
      "content": "<p>You mean 0.064 ~ 0.096 AP?  (0.064-0.96 is a pretty wide interval imho ^^)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1258326,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-31T14:34:49.673000",
          "content": "<p>0.064 score on private public LB for a single class</p>\n<p>Since we have 15 classes  0.064 x 15 = 0.96  ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1258354,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2021-03-31T14:49:46.837000",
          "content": "<p>Ah, I see, sorry ^^</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1257733": "What was your best Normal/Abnormal LB score?\n\n\nI have seen lots of AUC CV 0.99+ models but I wonder how they translate on the LB?\n\nAs @cdeotte wrote you can test it by just submitting boxes for 14.\n\n> For each image just add the string 14 PR 0 0 1 1 where PR = prob(no finding) from your classifier. If your classifier is perfect, then you will score full mAP = 1.0 for class 14. Most classifiers have CV AUC 0.99+, so most teams will maximize this class mAP\n\nWe tried to build pretty robust classifiers trained on external data, but could not go above 0.064 ~ 0.96 AP",
    "1258377": "0.065\n\nwe dont use any separate classifier but just multiply the box predictions for an image to get to the class 14 probability\nacross models this is averaged",
    "1257812": "Ours is LB 0.064",
    "1258695": "In this competition class 14 is not so important. If you get 0.064, that's ok, close to the maximum.",
    "1258417": "I just want to point out that imo these high results for a \"no finding\" binary classifier are representing the most important practical consideration from this competition. This is unseen territory for chest xray in the research litterature. Most reported binary classifiers are usually trained on full image classification labels which consequently has trouble to deal with spatially small but significant abnormalities. These results are directly related to the huge multi-class bbox annotation effort done by Vinbigdata and their associated radiologists.",
    "1258316": "You mean 0.064 ~ 0.096 AP?  (0.064-0.96 is a pretty wide interval imho ^^)"
  }
}