{
  "id": 67997,
  "title": "gap between local and public LB",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/67997",
  "author_name": "outrunner",
  "post_date": "2018-10-08T09:10:02.730000",
  "votes": 12,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Local val: Top1-acc = 0.835, Top3-acc = 0.946, MAP@3 = 0.886</p>\n\n<p>Public LB: 0.940</p>\n\n<p>Does someone has the same problem?</p>\n\n<p>[edited:]</p>\n\n<p>it seems normal because:</p>\n\n<blockquote>\n  <p>You’ll need to build a recognizer that can effectively learn from this noisy data and perform well on a manually-labeled test set from a different distribution.</p>\n</blockquote>",
  "messages": [
    {
      "id": 400407,
      "postDate": "2018-10-08T09:10:02.730Z",
      "content": "<p>Local val: Top1-acc = 0.835, Top3-acc = 0.946, MAP@3 = 0.886</p>\n\n<p>Public LB: 0.940</p>\n\n<p>Does someone has the same problem?</p>\n\n<p>[edited:]</p>\n\n<p>it seems normal because:</p>\n\n<blockquote>\n  <p>You’ll need to build a recognizer that can effectively learn from this noisy data and perform well on a manually-labeled test set from a different distribution.</p>\n</blockquote>",
      "rawMarkdown": "Local val: Top1-acc = 0.835, Top3-acc = 0.946, MAP@3 = 0.886\n\nPublic LB: 0.940\n\nDoes someone has the same problem?\n\n[edited:]\n\nit seems normal because:\n\n&gt; You’ll need to build a recognizer that can effectively learn from this noisy data and perform well on a manually-labeled test set from a different distribution.",
      "votes": 12
    },
    {
      "id": 419550,
      "postDate": "2018-11-12T07:11:08.740Z",
      "content": "<p>Update~~~\nFYI, my 5 submissions (80 images/class for validation):</p>\n\n<ol>\n<li>Local: 0.859 Public LB: 0.906 GAP: 0.047</li>\n<li>Local: 0.866 Public LB: 0.914 GAP: 0.048</li>\n<li>Local: 0.876 Public LB: 0.926 GAP: 0.050</li>\n<li>Local: 0.885 Public LB: 0.932 GAP: 0.047</li>\n<li>Local: 0.887 Public LB: 0.935 GAP: 0.048  ==add some magic things: ) </li>\n<li>to be updated...</li>\n</ol>",
      "rawMarkdown": "Update~~~\nFYI, my 5 submissions (80 images/class for validation):\n\n1. Local: 0.859 Public LB: 0.906 GAP: 0.047\n2. Local: 0.866 Public LB: 0.914 GAP: 0.048\n3. Local: 0.876 Public LB: 0.926 GAP: 0.050\n4. Local: 0.885 Public LB: 0.932 GAP: 0.047\n5. Local: 0.887 Public LB: 0.935 GAP: 0.048  ==add some magic things: ) \n6. to be updated...",
      "votes": 3,
      "replies": [
        {
          "id": 419807,
          "postDate": "2018-11-12T15:30:27.150Z",
          "content": "<p>actually, i find that ce loss is a better judge for LB score.</p>\n\n<p>because you my get high local LB but the probability confidence may be low. low probability gives lower public LB.</p>",
          "rawMarkdown": "actually, i find that ce loss is a better judge for LB score.\n\nbecause you my get high local LB but the probability confidence may be low. low probability gives lower public LB.",
          "votes": 3
        },
        {
          "id": 420066,
          "postDate": "2018-11-13T02:19:58.387Z",
          "content": "<p>Thanks, Heng. I'll have a try.</p>",
          "rawMarkdown": "Thanks, Heng. I'll have a try.",
          "votes": 1
        }
      ]
    },
    {
      "id": 417281,
      "postDate": "2018-11-08T03:06:35.757Z",
      "content": "<p>train = all \nvalidation = all, but print your metrics separately for all, recognised,non-recognised samples</p>\n\n<p>some of the 'non-recognised' samples are correct, that is why it helps the LB. But some are wrong.\nIf you can relabel the the 'non-recognised' samples, maybe results improve. Each class has about 10% 'non-recognised' samples , it would be easy to manually relabel them. or you can use your trained classifier to rank and select them.</p>\n\n<p>for more fancy approach, refer to :<a href=\"http://boqinggong.info/papers/wacv18.pdf\">http://boqinggong.info/papers/wacv18.pdf</a></p>\n\n<p>cleaning up noisy label, can improve your score a bit.</p>\n\n<hr>\n\n<p>on a separate note, from the 'non-recognised' vs 'recognised' samples, it is possible to train a classifier (per class) to predict if the sample is recognised or not. hence there are 340x2=680 target labels: e.g. 'recognised cat', 'non recognised cat, 'recognised lion', 'non recognised lion'</p>\n\n<p>or you can train 340+1 label: 340 softmax class, + 1 sigmoid class</p>",
      "rawMarkdown": "train = all \nvalidation = all, but print your metrics separately for all, recognised,non-recognised samples\n\nsome of the 'non-recognised' samples are correct, that is why it helps the LB. But some are wrong.\nIf you can relabel the the 'non-recognised' samples, maybe results improve. Each class has about 10% 'non-recognised' samples , it would be easy to manually relabel them. or you can use your trained classifier to rank and select them.\n\nfor more fancy approach, refer to :http://boqinggong.info/papers/wacv18.pdf\n\ncleaning up noisy label, can improve your score a bit.\n\n---\n\non a separate note, from the 'non-recognised' vs 'recognised' samples, it is possible to train a classifier (per class) to predict if the sample is recognised or not. hence there are 340x2=680 target labels: e.g. 'recognised cat', 'non recognised cat, 'recognised lion', 'non recognised lion'\n\nor you can train 340+1 label: 340 softmax class, + 1 sigmoid class",
      "votes": 3,
      "replies": [
        {
          "id": 419594,
          "postDate": "2018-11-12T09:05:40.377Z",
          "content": "<p>Is that 340 softmax class, + 1 sigmoid class can get a better socre?</p>",
          "rawMarkdown": "Is that 340 softmax class, + 1 sigmoid class can get a better socre?"
        }
      ]
    },
    {
      "id": 414527,
      "postDate": "2018-11-02T23:21:45.977Z",
      "content": "<p>My model has the training and val loss curves overlapping really well, so no sign of overfitting. it achieves a local val accuracy of 0.846 and top 3 accuracy of 0.96, yet the LB score is quite disappointing 0.88. Can you possibly think of why? </p>",
      "rawMarkdown": "My model has the training and val loss curves overlapping really well, so no sign of overfitting. it achieves a local val accuracy of 0.846 and top 3 accuracy of 0.96, yet the LB score is quite disappointing 0.88. Can you possibly think of why? ",
      "votes": 1,
      "replies": [
        {
          "id": 414540,
          "postDate": "2018-11-02T23:45:14.133Z",
          "content": "<p>Maybe you use recognized image only, or your train/val set mixed.</p>",
          "rawMarkdown": "Maybe you use recognized image only, or your train/val set mixed."
        },
        {
          "id": 414712,
          "postDate": "2018-11-03T12:45:02.760Z",
          "content": "<p>I used all recognized &amp; unrecognized images :-/</p>",
          "rawMarkdown": "I used all recognized &amp; unrecognized images :-/"
        },
        {
          "id": 415015,
          "postDate": "2018-11-04T06:05:37.350Z",
          "content": "<p>@HuyenNguyen, the metric of this competition is MAP@3, you should to measure your MAP@3 to know you are overfitted or not. </p>",
          "rawMarkdown": "@HuyenNguyen, the metric of this competition is MAP@3, you should to measure your MAP@3 to know you are overfitted or not. "
        },
        {
          "id": 415245,
          "postDate": "2018-11-04T18:33:58.957Z",
          "content": "<p>@HuyenNguyen</p>\n\n<p>i suspect some implementation bug, in code or data, etc</p>\n\n<p>\"top 3 accuracy of 0.96\" is very high score and i think it is higher than current top kagglers</p>",
          "rawMarkdown": "@HuyenNguyen\n\ni suspect some implementation bug, in code or data, etc\n\n\"top 3 accuracy of 0.96\" is very high score and i think it is higher than current top kagglers"
        },
        {
          "id": 415284,
          "postDate": "2018-11-04T20:10:06.277Z",
          "content": "<p>\"the training and val loss curves overlapping really well\"</p>\n\n<p>Do not look further, your validation data are used in the training set. </p>",
          "rawMarkdown": "\"the training and val loss curves overlapping really well\"\n\nDo not look further, your validation data are used in the training set. "
        },
        {
          "id": 415457,
          "postDate": "2018-11-05T06:58:44.457Z",
          "content": "<p>Or, did you ensure that your image size in test set is consistent with training set? I made this kind of fault yesterday..</p>",
          "rawMarkdown": "Or, did you ensure that your image size in test set is consistent with training set? I made this kind of fault yesterday.."
        },
        {
          "id": 415642,
          "postDate": "2018-11-05T12:41:32.103Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 406150,
      "postDate": "2018-10-18T18:29:40.730Z",
      "content": "<p>Local val: Top1-acc = 0.77, Top3-acc = 0.91, MAP@3 = 0.84</p>\n\n<p>Public LB: 0.89</p>",
      "rawMarkdown": "Local val: Top1-acc = 0.77, Top3-acc = 0.91, MAP@3 = 0.84\n\nPublic LB: 0.89",
      "votes": 1
    },
    {
      "id": 423127,
      "postDate": "2018-11-17T15:02:53.167Z",
      "content": "<p>Hi,\nwith 80 sample / class\nI get, validate: error 0.498, top1: 0.846, top3: 0.963, Private LB: 0.900, Public LB: 0.893\nVery different result, but LB score is fair to evalute public LB, \nIs that because not enough vlidation data,\nTherefore, there is high variance between each case?</p>",
      "rawMarkdown": "Hi,\nwith 80 sample / class\nI get, validate: error 0.498, top1: 0.846, top3: 0.963, Private LB: 0.900, Public LB: 0.893\nVery different result, but LB score is fair to evalute public LB, \nIs that because not enough vlidation data,\nTherefore, there is high variance between each case?"
    },
    {
      "id": 419535,
      "postDate": "2018-11-12T06:15:44.070Z",
      "content": "<p>Local val: Top-1 Acc= 0.828, Top-3 Acc = 0.939, MAP@3 = 0.874\nPublic LB: 0.929</p>",
      "rawMarkdown": "Local val: Top-1 Acc= 0.828, Top-3 Acc = 0.939, MAP@3 = 0.874\nPublic LB: 0.929"
    },
    {
      "id": 417216,
      "postDate": "2018-11-08T00:26:36.360Z",
      "content": "<p>Is there a gap between your train and val loss? Is it consistent? I'm using the full dataset and it opens up a huge gap. </p>",
      "rawMarkdown": "Is there a gap between your train and val loss? Is it consistent? I'm using the full dataset and it opens up a huge gap. ",
      "replies": [
        {
          "id": 417238,
          "postDate": "2018-11-08T01:18:55.567Z",
          "content": "<p>The loss gap between train and val is less than 0.05.</p>",
          "rawMarkdown": "The loss gap between train and val is less than 0.05."
        },
        {
          "id": 417266,
          "postDate": "2018-11-08T02:37:18.817Z",
          "content": "<p>Yes that's generally been my experience. But when I use more data, this gap gets bigger. When I used 30k images/class, the train and val losses were very close together. </p>",
          "rawMarkdown": "Yes that's generally been my experience. But when I use more data, this gap gets bigger. When I used 30k images/class, the train and val losses were very close together. "
        }
      ]
    },
    {
      "id": 414509,
      "postDate": "2018-11-02T22:41:24.820Z",
      "content": "<p>Do you find that for the same model, when you use different image sizes, the gap between local and LB scores changes? I keep getting this result when I increase the image size, the local val score looks better, but the LB score gets worse?</p>",
      "rawMarkdown": "Do you find that for the same model, when you use different image sizes, the gap between local and LB scores changes? I keep getting this result when I increase the image size, the local val score looks better, but the LB score gets worse?",
      "replies": [
        {
          "id": 414539,
          "postDate": "2018-11-02T23:41:24.507Z",
          "content": "<p>The gap will converge to zero when accuracy achieve 1.</p>\n\n<p>In my case, the relation is:</p>\n\n<pre><code>LB ~= 1 - ( 1 - Local_MAP@3 ) / 1.9\n</code></pre>",
          "rawMarkdown": "The gap will converge to zero when accuracy achieve 1.\n\nIn my case, the relation is:\n\n    LB ~= 1 - ( 1 - Local_MAP@3 ) / 1.9\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 407946,
      "postDate": "2018-10-22T03:16:15.053Z",
      "content": "<p>Local val: Top1-acc = 0.836, Top3-acc = 0.944</p>\n\n<p>Public LB: 0.926</p>",
      "rawMarkdown": "Local val: Top1-acc = 0.836, Top3-acc = 0.944\n\nPublic LB: 0.926"
    },
    {
      "id": 407839,
      "postDate": "2018-10-21T21:25:56.427Z",
      "content": "<p>Local val: Top1-acc = 0.8041, Top3-acc = 0.9259</p>\n\n<p>Public LB: 0.897</p>\n\n<p>(validation set = 2 Million samples distributed by the several classes)</p>\n\n<p>Not sure why my public LB score is so far from my local score.</p>",
      "rawMarkdown": "Local val: Top1-acc = 0.8041, Top3-acc = 0.9259\n\nPublic LB: 0.897\n\n(validation set = 2 Million samples distributed by the several classes)\n\nNot sure why my public LB score is so far from my local score."
    },
    {
      "id": 403612,
      "postDate": "2018-10-14T05:02:41.163Z",
      "content": "<p>Local val: MAP@3 = 0.86  (80 samples / class) <br>\nPublic: 0.917  </p>",
      "rawMarkdown": "Local val: MAP@3 = 0.86  (80 samples / class)  \nPublic: 0.917  "
    },
    {
      "id": 403469,
      "postDate": "2018-10-13T17:18:23.403Z",
      "content": "<p>Local val: Top1-acc = 0.800, MAP@3 = 0.855 (80samples per class)</p>\n\n<p>Public: 0.897</p>\n\n<p>hmmm, somebody must know something I don't.</p>",
      "rawMarkdown": "Local val: Top1-acc = 0.800, MAP@3 = 0.855 (80samples per class)\n\n\nPublic: 0.897\n\nhmmm, somebody must know something I don't."
    },
    {
      "id": 400533,
      "postDate": "2018-10-08T13:37:43.030Z",
      "content": "<p>Yes, I have something similar as well</p>",
      "rawMarkdown": "Yes, I have something similar as well"
    },
    {
      "id": 400415,
      "postDate": "2018-10-08T09:19:46.100Z",
      "content": "<p>yes, my values are about same as yours</p>\n\n<p>Local val: Top1-acc = 0.804, Top3-acc = 0.925, MAP@3 = 0.860</p>\n\n<p>Public LB: 0.916</p>\n\n<p>(validation set = 80 randoms sample per class)</p>",
      "rawMarkdown": "yes, my values are about same as yours\n\n\nLocal val: Top1-acc = 0.804, Top3-acc = 0.925, MAP@3 = 0.860\n\nPublic LB: 0.916\n\n(validation set = 80 randoms sample per class)",
      "replies": [
        {
          "id": 400534,
          "postDate": "2018-10-08T13:38:04.813Z",
          "content": "<p>On how many images you train the network ?</p>",
          "rawMarkdown": "On how many images you train the network ?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 419550,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2018-11-12T07:11:08.740000",
      "content": "<p>Update~~~\nFYI, my 5 submissions (80 images/class for validation):</p>\n\n<ol>\n<li>Local: 0.859 Public LB: 0.906 GAP: 0.047</li>\n<li>Local: 0.866 Public LB: 0.914 GAP: 0.048</li>\n<li>Local: 0.876 Public LB: 0.926 GAP: 0.050</li>\n<li>Local: 0.885 Public LB: 0.932 GAP: 0.047</li>\n<li>Local: 0.887 Public LB: 0.935 GAP: 0.048  ==add some magic things: ) </li>\n<li>to be updated...</li>\n</ol>",
      "votes": 3,
      "replies": [
        {
          "id": 419807,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-12T15:30:27.150000",
          "content": "<p>actually, i find that ce loss is a better judge for LB score.</p>\n\n<p>because you my get high local LB but the probability confidence may be low. low probability gives lower public LB.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 420066,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2018-11-13T02:19:58.387000",
          "content": "<p>Thanks, Heng. I'll have a try.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 417281,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-08T03:06:35.757000",
      "content": "<p>train = all \nvalidation = all, but print your metrics separately for all, recognised,non-recognised samples</p>\n\n<p>some of the 'non-recognised' samples are correct, that is why it helps the LB. But some are wrong.\nIf you can relabel the the 'non-recognised' samples, maybe results improve. Each class has about 10% 'non-recognised' samples , it would be easy to manually relabel them. or you can use your trained classifier to rank and select them.</p>\n\n<p>for more fancy approach, refer to :<a href=\"http://boqinggong.info/papers/wacv18.pdf\">http://boqinggong.info/papers/wacv18.pdf</a></p>\n\n<p>cleaning up noisy label, can improve your score a bit.</p>\n\n<hr>\n\n<p>on a separate note, from the 'non-recognised' vs 'recognised' samples, it is possible to train a classifier (per class) to predict if the sample is recognised or not. hence there are 340x2=680 target labels: e.g. 'recognised cat', 'non recognised cat, 'recognised lion', 'non recognised lion'</p>\n\n<p>or you can train 340+1 label: 340 softmax class, + 1 sigmoid class</p>",
      "votes": 3,
      "replies": [
        {
          "id": 419594,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2018-11-12T09:05:40.377000",
          "content": "<p>Is that 340 softmax class, + 1 sigmoid class can get a better socre?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414527,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-11-02T23:21:45.977000",
      "content": "<p>My model has the training and val loss curves overlapping really well, so no sign of overfitting. it achieves a local val accuracy of 0.846 and top 3 accuracy of 0.96, yet the LB score is quite disappointing 0.88. Can you possibly think of why? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 414540,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-11-02T23:45:14.133000",
          "content": "<p>Maybe you use recognized image only, or your train/val set mixed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414712,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-03T12:45:02.760000",
          "content": "<p>I used all recognized &amp; unrecognized images :-/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415015,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2018-11-04T06:05:37.350000",
          "content": "<p>@HuyenNguyen, the metric of this competition is MAP@3, you should to measure your MAP@3 to know you are overfitted or not. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415245,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-04T18:33:58.957000",
          "content": "<p>@HuyenNguyen</p>\n\n<p>i suspect some implementation bug, in code or data, etc</p>\n\n<p>\"top 3 accuracy of 0.96\" is very high score and i think it is higher than current top kagglers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415284,
          "author_name": "jeandebleau",
          "author_url": "",
          "post_date": "2018-11-04T20:10:06.277000",
          "content": "<p>\"the training and val loss curves overlapping really well\"</p>\n\n<p>Do not look further, your validation data are used in the training set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415457,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2018-11-05T06:58:44.457000",
          "content": "<p>Or, did you ensure that your image size in test set is consistent with training set? I made this kind of fault yesterday..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415642,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-05T12:41:32.103000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 406150,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2018-10-18T18:29:40.730000",
      "content": "<p>Local val: Top1-acc = 0.77, Top3-acc = 0.91, MAP@3 = 0.84</p>\n\n<p>Public LB: 0.89</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423127,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "2018-11-17T15:02:53.167000",
      "content": "<p>Hi,\nwith 80 sample / class\nI get, validate: error 0.498, top1: 0.846, top3: 0.963, Private LB: 0.900, Public LB: 0.893\nVery different result, but LB score is fair to evalute public LB, \nIs that because not enough vlidation data,\nTherefore, there is high variance between each case?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 419535,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2018-11-12T06:15:44.070000",
      "content": "<p>Local val: Top-1 Acc= 0.828, Top-3 Acc = 0.939, MAP@3 = 0.874\nPublic LB: 0.929</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 417216,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-11-08T00:26:36.360000",
      "content": "<p>Is there a gap between your train and val loss? Is it consistent? I'm using the full dataset and it opens up a huge gap. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 417238,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-11-08T01:18:55.567000",
          "content": "<p>The loss gap between train and val is less than 0.05.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417266,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-08T02:37:18.817000",
          "content": "<p>Yes that's generally been my experience. But when I use more data, this gap gets bigger. When I used 30k images/class, the train and val losses were very close together. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 414509,
      "author_name": "HuyenNguyen",
      "author_url": "",
      "post_date": "2018-11-02T22:41:24.820000",
      "content": "<p>Do you find that for the same model, when you use different image sizes, the gap between local and LB scores changes? I keep getting this result when I increase the image size, the local val score looks better, but the LB score gets worse?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 414539,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-11-02T23:41:24.507000",
          "content": "<p>The gap will converge to zero when accuracy achieve 1.</p>\n\n<p>In my case, the relation is:</p>\n\n<pre><code>LB ~= 1 - ( 1 - Local_MAP@3 ) / 1.9\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 407946,
      "author_name": "luudactam",
      "author_url": "",
      "post_date": "2018-10-22T03:16:15.053000",
      "content": "<p>Local val: Top1-acc = 0.836, Top3-acc = 0.944</p>\n\n<p>Public LB: 0.926</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 407839,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2018-10-21T21:25:56.427000",
      "content": "<p>Local val: Top1-acc = 0.8041, Top3-acc = 0.9259</p>\n\n<p>Public LB: 0.897</p>\n\n<p>(validation set = 2 Million samples distributed by the several classes)</p>\n\n<p>Not sure why my public LB score is so far from my local score.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 403612,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2018-10-14T05:02:41.163000",
      "content": "<p>Local val: MAP@3 = 0.86  (80 samples / class) <br>\nPublic: 0.917  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 403469,
      "author_name": "Miha Skalic",
      "author_url": "",
      "post_date": "2018-10-13T17:18:23.403000",
      "content": "<p>Local val: Top1-acc = 0.800, MAP@3 = 0.855 (80samples per class)</p>\n\n<p>Public: 0.897</p>\n\n<p>hmmm, somebody must know something I don't.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400533,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2018-10-08T13:37:43.030000",
      "content": "<p>Yes, I have something similar as well</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 400415,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-10-08T09:19:46.100000",
      "content": "<p>yes, my values are about same as yours</p>\n\n<p>Local val: Top1-acc = 0.804, Top3-acc = 0.925, MAP@3 = 0.860</p>\n\n<p>Public LB: 0.916</p>\n\n<p>(validation set = 80 randoms sample per class)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 400534,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2018-10-08T13:38:04.813000",
          "content": "<p>On how many images you train the network ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "400407": "Local val: Top1-acc = 0.835, Top3-acc = 0.946, MAP@3 = 0.886\n\nPublic LB: 0.940\n\nDoes someone has the same problem?\n\n[edited:]\n\nit seems normal because:\n\n&gt; You’ll need to build a recognizer that can effectively learn from this noisy data and perform well on a manually-labeled test set from a different distribution.",
    "419550": "Update~~~\nFYI, my 5 submissions (80 images/class for validation):\n\n1. Local: 0.859 Public LB: 0.906 GAP: 0.047\n2. Local: 0.866 Public LB: 0.914 GAP: 0.048\n3. Local: 0.876 Public LB: 0.926 GAP: 0.050\n4. Local: 0.885 Public LB: 0.932 GAP: 0.047\n5. Local: 0.887 Public LB: 0.935 GAP: 0.048  ==add some magic things: ) \n6. to be updated...",
    "417281": "train = all \nvalidation = all, but print your metrics separately for all, recognised,non-recognised samples\n\nsome of the 'non-recognised' samples are correct, that is why it helps the LB. But some are wrong.\nIf you can relabel the the 'non-recognised' samples, maybe results improve. Each class has about 10% 'non-recognised' samples , it would be easy to manually relabel them. or you can use your trained classifier to rank and select them.\n\nfor more fancy approach, refer to :http://boqinggong.info/papers/wacv18.pdf\n\ncleaning up noisy label, can improve your score a bit.\n\n---\n\non a separate note, from the 'non-recognised' vs 'recognised' samples, it is possible to train a classifier (per class) to predict if the sample is recognised or not. hence there are 340x2=680 target labels: e.g. 'recognised cat', 'non recognised cat, 'recognised lion', 'non recognised lion'\n\nor you can train 340+1 label: 340 softmax class, + 1 sigmoid class",
    "414527": "My model has the training and val loss curves overlapping really well, so no sign of overfitting. it achieves a local val accuracy of 0.846 and top 3 accuracy of 0.96, yet the LB score is quite disappointing 0.88. Can you possibly think of why? ",
    "406150": "Local val: Top1-acc = 0.77, Top3-acc = 0.91, MAP@3 = 0.84\n\nPublic LB: 0.89",
    "423127": "Hi,\nwith 80 sample / class\nI get, validate: error 0.498, top1: 0.846, top3: 0.963, Private LB: 0.900, Public LB: 0.893\nVery different result, but LB score is fair to evalute public LB, \nIs that because not enough vlidation data,\nTherefore, there is high variance between each case?",
    "419535": "Local val: Top-1 Acc= 0.828, Top-3 Acc = 0.939, MAP@3 = 0.874\nPublic LB: 0.929",
    "417216": "Is there a gap between your train and val loss? Is it consistent? I'm using the full dataset and it opens up a huge gap. ",
    "414509": "Do you find that for the same model, when you use different image sizes, the gap between local and LB scores changes? I keep getting this result when I increase the image size, the local val score looks better, but the LB score gets worse?",
    "407946": "Local val: Top1-acc = 0.836, Top3-acc = 0.944\n\nPublic LB: 0.926",
    "407839": "Local val: Top1-acc = 0.8041, Top3-acc = 0.9259\n\nPublic LB: 0.897\n\n(validation set = 2 Million samples distributed by the several classes)\n\nNot sure why my public LB score is so far from my local score.",
    "403612": "Local val: MAP@3 = 0.86  (80 samples / class)  \nPublic: 0.917  ",
    "403469": "Local val: Top1-acc = 0.800, MAP@3 = 0.855 (80samples per class)\n\n\nPublic: 0.897\n\nhmmm, somebody must know something I don't.",
    "400533": "Yes, I have something similar as well",
    "400415": "yes, my values are about same as yours\n\n\nLocal val: Top1-acc = 0.804, Top3-acc = 0.925, MAP@3 = 0.860\n\nPublic LB: 0.916\n\n(validation set = 80 randoms sample per class)"
  }
}