{
  "id": 70553,
  "title": "Better val accuracy with smaller train dataset?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70553",
  "author_name": "Dileep Patchigolla",
  "post_date": "2018-11-05T07:05:23.950000",
  "votes": -1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi\nI was training an LSTM model on the data, and built the working code on my local machine - with a mere 4% of training data. On that, I was getting about 50% as validation top 3 accuracy after Four epochs. But while I try the same model on kaggle kernel on full dataset, I get only 24% accuracy after 4th epoch. When can it happen that validation accuracy will get worse if trained on more data? On this locally trained model, after 16 epochs, here's the top 3 accuracies:\nTrain: 0.78\nVal: 0.67\nLB: 0.61</p>\n\n<p>Looking at the huge drop from train to LB, I think I was overfitting, hence the artificial high accuracy. Is there any other possible answer why small train dataset can give higher accuracy?</p>",
  "messages": [
    {
      "id": 415471,
      "postDate": "2018-11-05T07:30:43.233Z",
      "content": "<p>Small dataset means smaller validation set maybe this is why you are getting better accuracy on smaller dataset locally. And it is not generalizing on a larger dataset. My CNN+LSTM model scored 0.84 MAP@3 locally and 0.87 on public LB when I am training on around 100000 rows per file using then shuffle csv approach by <a href=\"/beluga\">@beluga</a> . You can try using more data and see if accuracy improves.</p>",
      "rawMarkdown": "Small dataset means smaller validation set maybe this is why you are getting better accuracy on smaller dataset locally. And it is not generalizing on a larger dataset. My CNN+LSTM model scored 0.84 MAP@3 locally and 0.87 on public LB when I am training on around 100000 rows per file using then shuffle csv approach by @beluga . You can try using more data and see if accuracy improves.",
      "votes": 1,
      "replies": [
        {
          "id": 415532,
          "postDate": "2018-11-05T09:34:09.463Z",
          "content": "<p>Thank you. I am trying on larger sample size, but wanted to know if this phenomenon is something that's observed generally. </p>",
          "rawMarkdown": "Thank you. I am trying on larger sample size, but wanted to know if this phenomenon is something that's observed generally. "
        }
      ]
    },
    {
      "id": 415459,
      "postDate": "2018-11-05T07:05:23.950Z",
      "content": "<p>Hi\nI was training an LSTM model on the data, and built the working code on my local machine - with a mere 4% of training data. On that, I was getting about 50% as validation top 3 accuracy after Four epochs. But while I try the same model on kaggle kernel on full dataset, I get only 24% accuracy after 4th epoch. When can it happen that validation accuracy will get worse if trained on more data? On this locally trained model, after 16 epochs, here's the top 3 accuracies:\nTrain: 0.78\nVal: 0.67\nLB: 0.61</p>\n\n<p>Looking at the huge drop from train to LB, I think I was overfitting, hence the artificial high accuracy. Is there any other possible answer why small train dataset can give higher accuracy?</p>",
      "rawMarkdown": "Hi\nI was training an LSTM model on the data, and built the working code on my local machine - with a mere 4% of training data. On that, I was getting about 50% as validation top 3 accuracy after Four epochs. But while I try the same model on kaggle kernel on full dataset, I get only 24% accuracy after 4th epoch. When can it happen that validation accuracy will get worse if trained on more data? On this locally trained model, after 16 epochs, here's the top 3 accuracies:\nTrain: 0.78\nVal: 0.67\nLB: 0.61\n\nLooking at the huge drop from train to LB, I think I was overfitting, hence the artificial high accuracy. Is there any other possible answer why small train dataset can give higher accuracy?",
      "votes": -1
    }
  ],
  "comments": [
    {
      "id": 415471,
      "author_name": "Ankit Sati",
      "author_url": "",
      "post_date": "2018-11-05T07:30:43.233000",
      "content": "<p>Small dataset means smaller validation set maybe this is why you are getting better accuracy on smaller dataset locally. And it is not generalizing on a larger dataset. My CNN+LSTM model scored 0.84 MAP@3 locally and 0.87 on public LB when I am training on around 100000 rows per file using then shuffle csv approach by <a href=\"/beluga\">@beluga</a> . You can try using more data and see if accuracy improves.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 415532,
          "author_name": "Dileep Patchigolla",
          "author_url": "",
          "post_date": "2018-11-05T09:34:09.463000",
          "content": "<p>Thank you. I am trying on larger sample size, but wanted to know if this phenomenon is something that's observed generally. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "415471": "Small dataset means smaller validation set maybe this is why you are getting better accuracy on smaller dataset locally. And it is not generalizing on a larger dataset. My CNN+LSTM model scored 0.84 MAP@3 locally and 0.87 on public LB when I am training on around 100000 rows per file using then shuffle csv approach by @beluga . You can try using more data and see if accuracy improves.",
    "415459": "Hi\nI was training an LSTM model on the data, and built the working code on my local machine - with a mere 4% of training data. On that, I was getting about 50% as validation top 3 accuracy after Four epochs. But while I try the same model on kaggle kernel on full dataset, I get only 24% accuracy after 4th epoch. When can it happen that validation accuracy will get worse if trained on more data? On this locally trained model, after 16 epochs, here's the top 3 accuracies:\nTrain: 0.78\nVal: 0.67\nLB: 0.61\n\nLooking at the huge drop from train to LB, I think I was overfitting, hence the artificial high accuracy. Is there any other possible answer why small train dataset can give higher accuracy?"
  }
}