{
  "id": 220761,
  "title": "Experiment on CV/Public LB relationship",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/220761",
  "author_name": "Hannes Öhler",
  "post_date": "2021-02-19T13:09:17.378000",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Following up on <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219090\" target=\"_blank\">this</a> discussion I investigated the idea that the CV-Public LB relationship should be more stable for the more frequent classes and run a little experiment:</p>\n<p>I submitted three models with predictions restricted to the four most frequent classes (Aortic enlargement, Cardiomegaly, Pleural thickening and Pulmonary fibrosis). As you can see from the table below, in case of all class predictions submitted the Public LB is significantly lower for the second and third model although their CV scores are a bit higher than the one of the baseline model. When restricting the submission to the four most frequent classes, the LB score does not decrease as before and the CV/Public LB relationship appears more stable :</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th></th>\n<th></th>\n<th>All classes</th>\n<th></th>\n<th>Most freq. classes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>CV</td>\n<td>Diff CV</td>\n<td>Public LB</td>\n<td>Diff LB</td>\n<td>Public LB</td>\n<td>Diff LB</td>\n</tr>\n<tr>\n<td>baseline model</td>\n<td>0.334</td>\n<td>baseline</td>\n<td>0.248</td>\n<td>baseline</td>\n<td>0.146</td>\n<td>baseline</td>\n</tr>\n<tr>\n<td>second model</td>\n<td>0.346</td>\n<td>0.012</td>\n<td>0.222</td>\n<td>-0.026</td>\n<td>0.145</td>\n<td>-0.001</td>\n</tr>\n<tr>\n<td>third model</td>\n<td>0.342</td>\n<td>0.008</td>\n<td>0.232</td>\n<td>-0.016</td>\n<td>0.145</td>\n<td>-0.001</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 1210471,
      "postDate": "2021-02-19T13:09:17.377Z",
      "content": "<p>Following up on <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219090\" target=\"_blank\">this</a> discussion I investigated the idea that the CV-Public LB relationship should be more stable for the more frequent classes and run a little experiment:</p>\n<p>I submitted three models with predictions restricted to the four most frequent classes (Aortic enlargement, Cardiomegaly, Pleural thickening and Pulmonary fibrosis). As you can see from the table below, in case of all class predictions submitted the Public LB is significantly lower for the second and third model although their CV scores are a bit higher than the one of the baseline model. When restricting the submission to the four most frequent classes, the LB score does not decrease as before and the CV/Public LB relationship appears more stable :</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n<th></th>\n<th></th>\n<th>All classes</th>\n<th></th>\n<th>Most freq. classes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>CV</td>\n<td>Diff CV</td>\n<td>Public LB</td>\n<td>Diff LB</td>\n<td>Public LB</td>\n<td>Diff LB</td>\n</tr>\n<tr>\n<td>baseline model</td>\n<td>0.334</td>\n<td>baseline</td>\n<td>0.248</td>\n<td>baseline</td>\n<td>0.146</td>\n<td>baseline</td>\n</tr>\n<tr>\n<td>second model</td>\n<td>0.346</td>\n<td>0.012</td>\n<td>0.222</td>\n<td>-0.026</td>\n<td>0.145</td>\n<td>-0.001</td>\n</tr>\n<tr>\n<td>third model</td>\n<td>0.342</td>\n<td>0.008</td>\n<td>0.232</td>\n<td>-0.016</td>\n<td>0.145</td>\n<td>-0.001</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Following up on [this](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219090) discussion I investigated the idea that the CV-Public LB relationship should be more stable for the more frequent classes and run a little experiment:\n\nI submitted three models with predictions restricted to the four most frequent classes (Aortic enlargement, Cardiomegaly, Pleural thickening and Pulmonary fibrosis). As you can see from the table below, in case of all class predictions submitted the Public LB is significantly lower for the second and third model although their CV scores are a bit higher than the one of the baseline model. When restricting the submission to the four most frequent classes, the LB score does not decrease as before and the CV/Public LB relationship appears more stable ~~(although it does not refect the subtle increases in CV scores)~~:\n\n||||| All classes|| Most freq. classes |\n| --- | --- | --- |\n|  | CV|Diff CV|\tPublic LB\t|Diff LB\t|Public LB\t|Diff LB\n|baseline model |0.334|\tbaseline\t|0.248\t|baseline\t|0.146\t|baseline\n|second model |0.346|\t0.012\t|0.222\t|-0.026\t|0.145\t|-0.001\n|third model |0.342\t|0.008\t|0.232\t|-0.016\t|0.145\t|-0.001\n\n\n\n",
      "votes": 17
    },
    {
      "id": 1233529,
      "postDate": "2021-03-10T13:54:12.310Z",
      "content": "<blockquote>\n  <p>(although it does not refect the subtle increases in CV scores):</p>\n</blockquote>\n<p>I think to be able to come up with this statement you should also check the AP score for each top class for all three models instead of checking overall CVs. Maybe your AP scores for these four classes are already similar for the three models.</p>",
      "rawMarkdown": "> (although it does not refect the subtle increases in CV scores):\n\nI think to be able to come up with this statement you should also check the AP score for each top class for all three models instead of checking overall CVs. Maybe your AP scores for these four classes are already similar for the three models.",
      "votes": 3,
      "replies": [
        {
          "id": 1233897,
          "postDate": "2021-03-10T19:00:22.857Z",
          "content": "<p>You are obviously right. Thanks!</p>",
          "rawMarkdown": "You are obviously right. Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1231220,
      "postDate": "2021-03-08T18:47:59.587Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a>!<br>\nOur CV-LB scores are also instable :)</p>\n<p>Did you use the 14 - No finding class too in your restricted experiments or just the 4 mentioned classes?</p>",
      "rawMarkdown": "Thanks @hannes82!\nOur CV-LB scores are also instable :)\n\nDid you use the 14 - No finding class too in your restricted experiments or just the 4 mentioned classes?",
      "votes": 3,
      "replies": [
        {
          "id": 1231324,
          "postDate": "2021-03-08T21:55:06.450Z",
          "content": "<p>Oh I didn't think about the No finding class. I actually just used the 2-class prediction model for the No finding class in the experiment. So in this sense, it is included.</p>",
          "rawMarkdown": "Oh I didn't think about the No finding class. I actually just used the 2-class prediction model for the No finding class in the experiment. So in this sense, it is included.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1232882,
      "postDate": "2021-03-10T03:08:25.680Z",
      "content": "<p>I see many people using YOLOv5 could obtain up to 0.4 CV but the score is still around ~0.23x </p>",
      "rawMarkdown": "I see many people using YOLOv5 could obtain up to 0.4 CV but the score is still around ~0.23x ",
      "replies": [
        {
          "id": 1233466,
          "postDate": "2021-03-10T12:47:55.340Z",
          "content": "<p>I think that is because they use weighted box fusion or some other preprocessing, which increases CV but not LB score.</p>",
          "rawMarkdown": "I think that is because they use weighted box fusion or some other preprocessing, which increases CV but not LB score."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1233529,
      "author_name": "Fatih Öztürk",
      "author_url": "",
      "post_date": "2021-03-10T13:54:12.310000",
      "content": "<blockquote>\n  <p>(although it does not refect the subtle increases in CV scores):</p>\n</blockquote>\n<p>I think to be able to come up with this statement you should also check the AP score for each top class for all three models instead of checking overall CVs. Maybe your AP scores for these four classes are already similar for the three models.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1233897,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-10T19:00:22.857000",
          "content": "<p>You are obviously right. Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1231220,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2021-03-08T18:47:59.587000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a>!<br>\nOur CV-LB scores are also instable :)</p>\n<p>Did you use the 14 - No finding class too in your restricted experiments or just the 4 mentioned classes?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1231324,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-08T21:55:06.450000",
          "content": "<p>Oh I didn't think about the No finding class. I actually just used the 2-class prediction model for the No finding class in the experiment. So in this sense, it is included.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1232882,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2021-03-10T03:08:25.680000",
      "content": "<p>I see many people using YOLOv5 could obtain up to 0.4 CV but the score is still around ~0.23x </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1233466,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-10T12:47:55.340000",
          "content": "<p>I think that is because they use weighted box fusion or some other preprocessing, which increases CV but not LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210471": "Following up on [this](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/219090) discussion I investigated the idea that the CV-Public LB relationship should be more stable for the more frequent classes and run a little experiment:\n\nI submitted three models with predictions restricted to the four most frequent classes (Aortic enlargement, Cardiomegaly, Pleural thickening and Pulmonary fibrosis). As you can see from the table below, in case of all class predictions submitted the Public LB is significantly lower for the second and third model although their CV scores are a bit higher than the one of the baseline model. When restricting the submission to the four most frequent classes, the LB score does not decrease as before and the CV/Public LB relationship appears more stable ~~(although it does not refect the subtle increases in CV scores)~~:\n\n||||| All classes|| Most freq. classes |\n| --- | --- | --- |\n|  | CV|Diff CV|\tPublic LB\t|Diff LB\t|Public LB\t|Diff LB\n|baseline model |0.334|\tbaseline\t|0.248\t|baseline\t|0.146\t|baseline\n|second model |0.346|\t0.012\t|0.222\t|-0.026\t|0.145\t|-0.001\n|third model |0.342\t|0.008\t|0.232\t|-0.016\t|0.145\t|-0.001\n\n\n\n",
    "1233529": "> (although it does not refect the subtle increases in CV scores):\n\nI think to be able to come up with this statement you should also check the AP score for each top class for all three models instead of checking overall CVs. Maybe your AP scores for these four classes are already similar for the three models.",
    "1231220": "Thanks @hannes82!\nOur CV-LB scores are also instable :)\n\nDid you use the 14 - No finding class too in your restricted experiments or just the 4 mentioned classes?",
    "1232882": "I see many people using YOLOv5 could obtain up to 0.4 CV but the score is still around ~0.23x "
  }
}