{
  "id": 218254,
  "title": "Tips for proper data augmentation",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/218254",
  "author_name": "Oleksiy",
  "post_date": "2021-02-09T22:41:40.220000",
  "votes": -2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyody. I am pretty new to this topic and doing as a school project.</p>\n<p><strong>1. I am trying to do classification on this dataset.</strong></p>\n<p>**# Read data<br>\ntrain_generator = train_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/train',<br>\n                                                    target_size = (224, 224),<br>\n                                                    batch_size = 128,<br>\n                                                    class_mode = 'categorical')</p>\n<p>Here I need to create a validation dataset<br>\nvalidation_generator = test_datagen.flow_from_directory('/content/drive/My Drive/DL Pneumonia data/val/', <br>\n                                                         target_size = (224, 224),<br>\n                                                         batch_size = 128,<br>\n                                                         class_mode = 'categorical')</p>\n<p>test_generator = eval_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/test', <br>\n                                                  target_size = (224, 224),<br>\n                                                  batch_size = 128, <br>\n                                                  class_mode = 'categorical')**</p>\n<p>I receive <code>Found 0 images belonging to 0 classes</code>. </p>\n<p>If I understand correctly, I need to have data in different subdirectories in my train. But I have 15000 resized images in <code>train</code> folder. </p>\n<p>How to approach it?</p>\n<p><strong>2. Also what is the good practice to split the training dataset into train and val datasets?</strong></p>\n<p><strong>3. What are the appropriate parameters for <code>datagen</code> in this dataset?</strong></p>",
  "messages": [
    {
      "id": 1193860,
      "postDate": "2021-02-09T22:41:40.220Z",
      "content": "<p>Hi everyody. I am pretty new to this topic and doing as a school project.</p>\n<p><strong>1. I am trying to do classification on this dataset.</strong></p>\n<p>**# Read data<br>\ntrain_generator = train_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/train',<br>\n                                                    target_size = (224, 224),<br>\n                                                    batch_size = 128,<br>\n                                                    class_mode = 'categorical')</p>\n<p>Here I need to create a validation dataset<br>\nvalidation_generator = test_datagen.flow_from_directory('/content/drive/My Drive/DL Pneumonia data/val/', <br>\n                                                         target_size = (224, 224),<br>\n                                                         batch_size = 128,<br>\n                                                         class_mode = 'categorical')</p>\n<p>test_generator = eval_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/test', <br>\n                                                  target_size = (224, 224),<br>\n                                                  batch_size = 128, <br>\n                                                  class_mode = 'categorical')**</p>\n<p>I receive <code>Found 0 images belonging to 0 classes</code>. </p>\n<p>If I understand correctly, I need to have data in different subdirectories in my train. But I have 15000 resized images in <code>train</code> folder. </p>\n<p>How to approach it?</p>\n<p><strong>2. Also what is the good practice to split the training dataset into train and val datasets?</strong></p>\n<p><strong>3. What are the appropriate parameters for <code>datagen</code> in this dataset?</strong></p>",
      "rawMarkdown": "Hi everyody. I am pretty new to this topic and doing as a school project.\n\n**1. I am trying to do classification on this dataset.**\n\n**# Read data\ntrain_generator = train_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/train',\n                                                    target_size = (224, 224),\n                                                    batch_size = 128,\n                                                    class_mode = 'categorical')\n\nHere I need to create a validation dataset\nvalidation_generator = test_datagen.flow_from_directory('/content/drive/My Drive/DL Pneumonia data/val/', \n                                                         target_size = (224, 224),\n                                                         batch_size = 128,\n                                                         class_mode = 'categorical')\n\ntest_generator = eval_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/test', \n                                                  target_size = (224, 224),\n                                                  batch_size = 128, \n                                                  class_mode = 'categorical')**\n\nI receive `Found 0 images belonging to 0 classes`. \n\nIf I understand correctly, I need to have data in different subdirectories in my train. But I have 15000 resized images in `train` folder. \n\nHow to approach it?\n\n**2. Also what is the good practice to split the training dataset into train and val datasets?**\n\n**3. What are the appropriate parameters for `datagen` in this dataset?**",
      "votes": -2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1193860": "Hi everyody. I am pretty new to this topic and doing as a school project.\n\n**1. I am trying to do classification on this dataset.**\n\n**# Read data\ntrain_generator = train_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/train',\n                                                    target_size = (224, 224),\n                                                    batch_size = 128,\n                                                    class_mode = 'categorical')\n\nHere I need to create a validation dataset\nvalidation_generator = test_datagen.flow_from_directory('/content/drive/My Drive/DL Pneumonia data/val/', \n                                                         target_size = (224, 224),\n                                                         batch_size = 128,\n                                                         class_mode = 'categorical')\n\ntest_generator = eval_datagen.flow_from_directory('/content/drive/My Drive/MScA Capstone/Classification/test', \n                                                  target_size = (224, 224),\n                                                  batch_size = 128, \n                                                  class_mode = 'categorical')**\n\nI receive `Found 0 images belonging to 0 classes`. \n\nIf I understand correctly, I need to have data in different subdirectories in my train. But I have 15000 resized images in `train` folder. \n\nHow to approach it?\n\n**2. Also what is the good practice to split the training dataset into train and val datasets?**\n\n**3. What are the appropriate parameters for `datagen` in this dataset?**"
  }
}