{
  "id": 393473,
  "title": "How can I split the Train image dataset in kaggle into train, test and validation ?",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/393473",
  "author_name": "Pablo Palacios",
  "post_date": "2023-03-09T14:41:42.392000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, I am working on <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256</a> and trying to split the Training data into train, test and validation in the output directory. Can someone please help me out with it.<br>\nI've used \"splitfolders\" but It doesn't do it in the right way, It only copy the first folder without the 2nd folder and the images.</p>\n<p>This is the code I used:</p>\n<p><code>train_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'</code></p>\n<pre><code>splitfolders.ratio(train_dir, output=, ratio=(,,))\nwarnings.filterwarnings()\n</code></pre>",
  "messages": [
    {
      "id": 2175246,
      "postDate": "2023-03-09T18:33:55.473Z",
      "content": "<p>Something like this:</p>\n<pre><code> os\n sklearn.model_selection  train_test_split\n\ntrain_dir = \n\n\nfile_names = os.listdir(train_dir)\n\n\ntrain_files, test_files = train_test_split(file_names, test_size=, random_state=)\n\n\ntrain_files, val_files = train_test_split(train_files, test_size=, random_state=)\n\n\n(, (train_files))\n(, (val_files))\n(, (test_files))\n</code></pre>\n<p>Another way could be done via <code>keras.preprocessing.image.ImageDataGenerator</code>.</p>",
      "rawMarkdown": "Something like this:\n```python\nimport os\nfrom sklearn.model_selection import train_test_split\n\ntrain_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'\n\n# Get the list of all file names in the train_dir\nfile_names = os.listdir(train_dir)\n\n# Split the data into train and test sets\ntrain_files, test_files = train_test_split(file_names, test_size=0.2, random_state=42)\n\n# Split the train set further into train and validation sets\ntrain_files, val_files = train_test_split(train_files, test_size=0.2, random_state=42)\n\n# Print the number of files in each set\nprint(\"Number of files in train set:\", len(train_files))\nprint(\"Number of files in validation set:\", len(val_files))\nprint(\"Number of files in test set:\", len(test_files))\n```\n\nAnother way could be done via `keras.preprocessing.image.ImageDataGenerator`.",
      "replies": [
        {
          "id": 2176091,
          "postDate": "2023-03-10T11:48:01.090Z",
          "content": "<p>Thank you, but I think it doesn't work well, because how do you know where are the images? I mean, train_files fo example has all the names of the first folder, but where is the second folder and the images? That's why I try to use splitfolders beacuse I saw so many people do it right, and the images where store in the Output in kaggle but I don't know if I'm making a mistake. </p>",
          "rawMarkdown": "Thank you, but I think it doesn't work well, because how do you know where are the images? I mean, train_files fo example has all the names of the first folder, but where is the second folder and the images? That's why I try to use splitfolders beacuse I saw so many people do it right, and the images where store in the Output in kaggle but I don't know if I'm making a mistake. ",
          "replies": [
            {
              "id": 2177693,
              "postDate": "2023-03-11T17:30:05.257Z",
              "content": "<p>Thank you so much, It was my bad that I didn't understand at first. Now it is working well.</p>",
              "rawMarkdown": "Thank you so much, It was my bad that I didn't understand at first. Now it is working well."
            }
          ]
        }
      ]
    },
    {
      "id": 2174964,
      "postDate": "2023-03-09T14:41:42.393Z",
      "content": "<p>Hi, I am working on <a href=\"https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256\" target=\"_blank\">https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256</a> and trying to split the Training data into train, test and validation in the output directory. Can someone please help me out with it.<br>\nI've used \"splitfolders\" but It doesn't do it in the right way, It only copy the first folder without the 2nd folder and the images.</p>\n<p>This is the code I used:</p>\n<p><code>train_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'</code></p>\n<pre><code>splitfolders.ratio(train_dir, output=, ratio=(,,))\nwarnings.filterwarnings()\n</code></pre>",
      "rawMarkdown": "Hi, I am working on https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256 and trying to split the Training data into train, test and validation in the output directory. Can someone please help me out with it.\nI've used \"splitfolders\" but It doesn't do it in the right way, It only copy the first folder without the 2nd folder and the images.\n\nThis is the code I used:\n\n`train_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'`\n```python\nsplitfolders.ratio(train_dir, output=\"output\", ratio=(0.6,0.2,0.2))\nwarnings.filterwarnings(\"ignore\")\n```\n\n\n"
    }
  ],
  "comments": [
    {
      "id": 2175246,
      "author_name": "Steven Van Ingelgem",
      "author_url": "",
      "post_date": "2023-03-09T18:33:55.473000",
      "content": "<p>Something like this:</p>\n<pre><code> os\n sklearn.model_selection  train_test_split\n\ntrain_dir = \n\n\nfile_names = os.listdir(train_dir)\n\n\ntrain_files, test_files = train_test_split(file_names, test_size=, random_state=)\n\n\ntrain_files, val_files = train_test_split(train_files, test_size=, random_state=)\n\n\n(, (train_files))\n(, (val_files))\n(, (test_files))\n</code></pre>\n<p>Another way could be done via <code>keras.preprocessing.image.ImageDataGenerator</code>.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2176091,
          "author_name": "Pablo Palacios",
          "author_url": "",
          "post_date": "2023-03-10T11:48:01.090000",
          "content": "<p>Thank you, but I think it doesn't work well, because how do you know where are the images? I mean, train_files fo example has all the names of the first folder, but where is the second folder and the images? That's why I try to use splitfolders beacuse I saw so many people do it right, and the images where store in the Output in kaggle but I don't know if I'm making a mistake. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2177693,
              "author_name": "Pablo Palacios",
              "author_url": "",
              "post_date": "2023-03-11T17:30:05.257000",
              "content": "<p>Thank you so much, It was my bad that I didn't understand at first. Now it is working well.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2175246": "Something like this:\n```python\nimport os\nfrom sklearn.model_selection import train_test_split\n\ntrain_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'\n\n# Get the list of all file names in the train_dir\nfile_names = os.listdir(train_dir)\n\n# Split the data into train and test sets\ntrain_files, test_files = train_test_split(file_names, test_size=0.2, random_state=42)\n\n# Split the train set further into train and validation sets\ntrain_files, val_files = train_test_split(train_files, test_size=0.2, random_state=42)\n\n# Print the number of files in each set\nprint(\"Number of files in train set:\", len(train_files))\nprint(\"Number of files in validation set:\", len(val_files))\nprint(\"Number of files in test set:\", len(test_files))\n```\n\nAnother way could be done via `keras.preprocessing.image.ImageDataGenerator`.",
    "2174964": "Hi, I am working on https://www.kaggle.com/datasets/vaillant/rsna-str-pe-detection-jpeg-256 and trying to split the Training data into train, test and validation in the output directory. Can someone please help me out with it.\nI've used \"splitfolders\" but It doesn't do it in the right way, It only copy the first folder without the 2nd folder and the images.\n\nThis is the code I used:\n\n`train_dir = r'../input/rsna-str-pe-detection-jpeg-256/train-jpegs/'`\n```python\nsplitfolders.ratio(train_dir, output=\"output\", ratio=(0.6,0.2,0.2))\nwarnings.filterwarnings(\"ignore\")\n```\n\n\n"
  }
}