{
  "id": 70925,
  "title": "How to train using all data?",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/70925",
  "author_name": "HuyenNguyen",
  "post_date": "2018-11-08T13:18:53.001000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am using this data generator method to be able to use as much data as possible. I use the first 3000 rows from each file as the validation set, the rest for training. But I don't know why it creates a huge gap between the training and validation losses. Can you think of why? </p>\n\n<pre><code>out_df_list = []\nfor y, cat in enumerate(list_all_categories()):\n    c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n    c_df = pd.read_csv(c_path, nrows=3000, skiprows=0)\n    c_df.columns=COL_NAMES\n    c_df['y'] = y\n    out_df_list += [c_df[['drawing',  'y']]]\nval = pd.concat(out_df_list)\n\nfor j in range(100):\n    out_df_list = []\n    for y, cat in enumerate(list_all_categories()):\n        c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n        c_df = pd.read_csv(c_path, nrows=1000, skiprows=3000+1000*j)\n        c_df.columns=COL_NAMES\n        c_df['y'] = y\n        out_df_list += [c_df[['drawing',  'y']]]\n    train = pd.concat(out_df_list)\n\n    def data_gen(bs = 340, data = train):\n        while True:\n            # The val data is made of the first 3000 rows from each category\n            # so when sampling the train data, avoid the first 3000 rows\n            df = data.sample(bs)\n            df['drawing'] = df['drawing'].apply(ast.literal_eval)\n            x = np.zeros((len(df), size, size, 1))\n            for i, raw_strokes in enumerate(df.drawing.values):\n                x[i, :, :, 0] = draw_cv2(raw_strokes, size=size, lw=6,\n                                         time_color=True)\n            x = np.repeat(x, 3, axis =3)\n            x = preprocess_input(x).astype(np.float32)\n            y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n            yield x, y\n\n    train_datagen = data_gen(bs = 340, data = train)\n    val_datagen = data_gen(bs = 340, data = val)\n\n    hist = model.fit_generator(\n        train_datagen, steps_per_epoch=STEPS, epochs=1, verbose=1,\n        validation_data=val_datagen, validation_steps = 100,\n        callbacks = callbacks\n    )\n    hists.append(hist)\n</code></pre>",
  "messages": [
    {
      "id": 417567,
      "postDate": "2018-11-08T13:18:53Z",
      "content": "<p>I am using this data generator method to be able to use as much data as possible. I use the first 3000 rows from each file as the validation set, the rest for training. But I don't know why it creates a huge gap between the training and validation losses. Can you think of why? </p>\n\n<pre><code>out_df_list = []\nfor y, cat in enumerate(list_all_categories()):\n    c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n    c_df = pd.read_csv(c_path, nrows=3000, skiprows=0)\n    c_df.columns=COL_NAMES\n    c_df['y'] = y\n    out_df_list += [c_df[['drawing',  'y']]]\nval = pd.concat(out_df_list)\n\nfor j in range(100):\n    out_df_list = []\n    for y, cat in enumerate(list_all_categories()):\n        c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n        c_df = pd.read_csv(c_path, nrows=1000, skiprows=3000+1000*j)\n        c_df.columns=COL_NAMES\n        c_df['y'] = y\n        out_df_list += [c_df[['drawing',  'y']]]\n    train = pd.concat(out_df_list)\n\n    def data_gen(bs = 340, data = train):\n        while True:\n            # The val data is made of the first 3000 rows from each category\n            # so when sampling the train data, avoid the first 3000 rows\n            df = data.sample(bs)\n            df['drawing'] = df['drawing'].apply(ast.literal_eval)\n            x = np.zeros((len(df), size, size, 1))\n            for i, raw_strokes in enumerate(df.drawing.values):\n                x[i, :, :, 0] = draw_cv2(raw_strokes, size=size, lw=6,\n                                         time_color=True)\n            x = np.repeat(x, 3, axis =3)\n            x = preprocess_input(x).astype(np.float32)\n            y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n            yield x, y\n\n    train_datagen = data_gen(bs = 340, data = train)\n    val_datagen = data_gen(bs = 340, data = val)\n\n    hist = model.fit_generator(\n        train_datagen, steps_per_epoch=STEPS, epochs=1, verbose=1,\n        validation_data=val_datagen, validation_steps = 100,\n        callbacks = callbacks\n    )\n    hists.append(hist)\n</code></pre>",
      "rawMarkdown": "I am using this data generator method to be able to use as much data as possible. I use the first 3000 rows from each file as the validation set, the rest for training. But I don't know why it creates a huge gap between the training and validation losses. Can you think of why? \n\n    out_df_list = []\n    for y, cat in enumerate(list_all_categories()):\n        c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n        c_df = pd.read_csv(c_path, nrows=3000, skiprows=0)\n        c_df.columns=COL_NAMES\n        c_df['y'] = y\n        out_df_list += [c_df[['drawing',  'y']]]\n    val = pd.concat(out_df_list)\n\n    for j in range(100):\n        out_df_list = []\n        for y, cat in enumerate(list_all_categories()):\n            c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n            c_df = pd.read_csv(c_path, nrows=1000, skiprows=3000+1000*j)\n            c_df.columns=COL_NAMES\n            c_df['y'] = y\n            out_df_list += [c_df[['drawing',  'y']]]\n        train = pd.concat(out_df_list)\n\n        def data_gen(bs = 340, data = train):\n            while True:\n                # The val data is made of the first 3000 rows from each category\n                # so when sampling the train data, avoid the first 3000 rows\n                df = data.sample(bs)\n                df['drawing'] = df['drawing'].apply(ast.literal_eval)\n                x = np.zeros((len(df), size, size, 1))\n                for i, raw_strokes in enumerate(df.drawing.values):\n                    x[i, :, :, 0] = draw_cv2(raw_strokes, size=size, lw=6,\n                                             time_color=True)\n                x = np.repeat(x, 3, axis =3)\n                x = preprocess_input(x).astype(np.float32)\n                y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n                yield x, y\n\n        train_datagen = data_gen(bs = 340, data = train)\n        val_datagen = data_gen(bs = 340, data = val)\n\n        hist = model.fit_generator(\n            train_datagen, steps_per_epoch=STEPS, epochs=1, verbose=1,\n            validation_data=val_datagen, validation_steps = 100,\n            callbacks = callbacks\n        )\n        hists.append(hist)",
      "votes": 3
    },
    {
      "id": 418058,
      "postDate": "2018-11-09T08:08:07.820Z",
      "content": "<p>As far as I can see, you are not shuffling your training data.</p>\n\n<p>Hence your neural network (nn) learns to change its mind quickly during training phase. When it first see \"airplane\" in a mini-batch, it will answer \"airplane\" in next mini-batch and it will be right... If batch size is 100, nn will be right 9 times out of 10... until next batch moved to \"alarm clock\", so it knows it has to answer \"alarm clock\" now, and so on.</p>\n\n<p>When switching to validation phase, as there is no back-prop, the nn will probably answer always the same thing and thus a worst accuracy.</p>\n\n<p>I would suggest to shuffle \"train\" right after the \"for\" loop.</p>",
      "rawMarkdown": "As far as I can see, you are not shuffling your training data.\n\nHence your neural network (nn) learns to change its mind quickly during training phase. When it first see \"airplane\" in a mini-batch, it will answer \"airplane\" in next mini-batch and it will be right... If batch size is 100, nn will be right 9 times out of 10... until next batch moved to \"alarm clock\", so it knows it has to answer \"alarm clock\" now, and so on.\n\nWhen switching to validation phase, as there is no back-prop, the nn will probably answer always the same thing and thus a worst accuracy.\n\nI would suggest to shuffle \"train\" right after the \"for\" loop.",
      "replies": [
        {
          "id": 418175,
          "postDate": "2018-11-09T12:15:29.057Z",
          "content": "<p>Hi Eric, but in this step df = data.sample(bs), it's taking 340 random images out of the dataframe, that's a mix of many different categories, not just one. </p>",
          "rawMarkdown": "Hi Eric, but in this step df = data.sample(bs), it's taking 340 random images out of the dataframe, that's a mix of many different categories, not just one. "
        },
        {
          "id": 418265,
          "postDate": "2018-11-09T15:11:33.957Z",
          "content": "<p>My bad, I missed it.</p>\n\n<p>I would have shuffled <code>train</code> outside the batch generator and designed the batch generator to pick a different set of image at each iteration. e.g. supposing <code>train</code> is shuffled, at first iteration get first <code>bs</code> images from <code>data</code>, then the next <code>bs</code> images...</p>\n\n<p>I supposed the above and read your code too fast. Sorry.</p>\n\n<p>Anyway, sampling <code>data</code> at each mini-batch does not ensure you will have different images at each mini-batch from previous mini-batches. I guess the more you are getting close to <code>len(data)/bs</code> mini-batches generated, the more likely you will have 100% of already seen images while the above method will ensure it won't happen.</p>\n\n<p>Maybe it worth to try this change.</p>",
          "rawMarkdown": "My bad, I missed it.\n\nI would have shuffled `train` outside the batch generator and designed the batch generator to pick a different set of image at each iteration. e.g. supposing `train` is shuffled, at first iteration get first `bs` images from `data`, then the next `bs` images...\n\nI supposed the above and read your code too fast. Sorry.\n\nAnyway, sampling `data` at each mini-batch does not ensure you will have different images at each mini-batch from previous mini-batches. I guess the more you are getting close to `len(data)/bs` mini-batches generated, the more likely you will have 100% of already seen images while the above method will ensure it won't happen.\n\nMaybe it worth to try this change."
        }
      ]
    },
    {
      "id": 417797,
      "postDate": "2018-11-08T19:45:19.947Z",
      "content": "<p>Hello, usually when the validation loss is much bigger than training loss, it signals that your model is overfitting. Since you generate data, it can add disbalance to the dataset and disturb the learning process. I'm not sure what model you use, but if it's a neural network, you can add a dropout layer which should help with overfitting.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Hello, usually when the validation loss is much bigger than training loss, it signals that your model is overfitting. Since you generate data, it can add disbalance to the dataset and disturb the learning process. I'm not sure what model you use, but if it's a neural network, you can add a dropout layer which should help with overfitting.\n\nGood luck!",
      "replies": [
        {
          "id": 417857,
          "postDate": "2018-11-08T21:54:10.970Z",
          "content": "<p>no, there's something dodgy about this generator that I can't see. because exactly the same model, if I generate the data from the Shuffled CSVs from Beluga's kernels, this gap disappears. </p>",
          "rawMarkdown": "no, there's something dodgy about this generator that I can't see. because exactly the same model, if I generate the data from the Shuffled CSVs from Beluga's kernels, this gap disappears. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 418058,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2018-11-09T08:08:07.820000",
      "content": "<p>As far as I can see, you are not shuffling your training data.</p>\n\n<p>Hence your neural network (nn) learns to change its mind quickly during training phase. When it first see \"airplane\" in a mini-batch, it will answer \"airplane\" in next mini-batch and it will be right... If batch size is 100, nn will be right 9 times out of 10... until next batch moved to \"alarm clock\", so it knows it has to answer \"alarm clock\" now, and so on.</p>\n\n<p>When switching to validation phase, as there is no back-prop, the nn will probably answer always the same thing and thus a worst accuracy.</p>\n\n<p>I would suggest to shuffle \"train\" right after the \"for\" loop.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 418175,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-09T12:15:29.057000",
          "content": "<p>Hi Eric, but in this step df = data.sample(bs), it's taking 340 random images out of the dataframe, that's a mix of many different categories, not just one. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418265,
          "author_name": "Eric Bouteillon",
          "author_url": "",
          "post_date": "2018-11-09T15:11:33.957000",
          "content": "<p>My bad, I missed it.</p>\n\n<p>I would have shuffled <code>train</code> outside the batch generator and designed the batch generator to pick a different set of image at each iteration. e.g. supposing <code>train</code> is shuffled, at first iteration get first <code>bs</code> images from <code>data</code>, then the next <code>bs</code> images...</p>\n\n<p>I supposed the above and read your code too fast. Sorry.</p>\n\n<p>Anyway, sampling <code>data</code> at each mini-batch does not ensure you will have different images at each mini-batch from previous mini-batches. I guess the more you are getting close to <code>len(data)/bs</code> mini-batches generated, the more likely you will have 100% of already seen images while the above method will ensure it won't happen.</p>\n\n<p>Maybe it worth to try this change.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 417797,
      "author_name": "Mark Tselikov",
      "author_url": "",
      "post_date": "2018-11-08T19:45:19.947000",
      "content": "<p>Hello, usually when the validation loss is much bigger than training loss, it signals that your model is overfitting. Since you generate data, it can add disbalance to the dataset and disturb the learning process. I'm not sure what model you use, but if it's a neural network, you can add a dropout layer which should help with overfitting.</p>\n\n<p>Good luck!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 417857,
          "author_name": "HuyenNguyen",
          "author_url": "",
          "post_date": "2018-11-08T21:54:10.970000",
          "content": "<p>no, there's something dodgy about this generator that I can't see. because exactly the same model, if I generate the data from the Shuffled CSVs from Beluga's kernels, this gap disappears. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417567": "I am using this data generator method to be able to use as much data as possible. I use the first 3000 rows from each file as the validation set, the rest for training. But I don't know why it creates a huge gap between the training and validation losses. Can you think of why? \n\n    out_df_list = []\n    for y, cat in enumerate(list_all_categories()):\n        c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n        c_df = pd.read_csv(c_path, nrows=3000, skiprows=0)\n        c_df.columns=COL_NAMES\n        c_df['y'] = y\n        out_df_list += [c_df[['drawing',  'y']]]\n    val = pd.concat(out_df_list)\n\n    for j in range(100):\n        out_df_list = []\n        for y, cat in enumerate(list_all_categories()):\n            c_path = os.path.join('../input/quickdraw-doodle-recognition', 'train_simplified',cat+'.csv')\n            c_df = pd.read_csv(c_path, nrows=1000, skiprows=3000+1000*j)\n            c_df.columns=COL_NAMES\n            c_df['y'] = y\n            out_df_list += [c_df[['drawing',  'y']]]\n        train = pd.concat(out_df_list)\n\n        def data_gen(bs = 340, data = train):\n            while True:\n                # The val data is made of the first 3000 rows from each category\n                # so when sampling the train data, avoid the first 3000 rows\n                df = data.sample(bs)\n                df['drawing'] = df['drawing'].apply(ast.literal_eval)\n                x = np.zeros((len(df), size, size, 1))\n                for i, raw_strokes in enumerate(df.drawing.values):\n                    x[i, :, :, 0] = draw_cv2(raw_strokes, size=size, lw=6,\n                                             time_color=True)\n                x = np.repeat(x, 3, axis =3)\n                x = preprocess_input(x).astype(np.float32)\n                y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n                yield x, y\n\n        train_datagen = data_gen(bs = 340, data = train)\n        val_datagen = data_gen(bs = 340, data = val)\n\n        hist = model.fit_generator(\n            train_datagen, steps_per_epoch=STEPS, epochs=1, verbose=1,\n            validation_data=val_datagen, validation_steps = 100,\n            callbacks = callbacks\n        )\n        hists.append(hist)",
    "418058": "As far as I can see, you are not shuffling your training data.\n\nHence your neural network (nn) learns to change its mind quickly during training phase. When it first see \"airplane\" in a mini-batch, it will answer \"airplane\" in next mini-batch and it will be right... If batch size is 100, nn will be right 9 times out of 10... until next batch moved to \"alarm clock\", so it knows it has to answer \"alarm clock\" now, and so on.\n\nWhen switching to validation phase, as there is no back-prop, the nn will probably answer always the same thing and thus a worst accuracy.\n\nI would suggest to shuffle \"train\" right after the \"for\" loop.",
    "417797": "Hello, usually when the validation loss is much bigger than training loss, it signals that your model is overfitting. Since you generate data, it can add disbalance to the dataset and disturb the learning process. I'm not sure what model you use, but if it's a neural network, you can add a dropout layer which should help with overfitting.\n\nGood luck!"
  }
}