{
  "id": 372161,
  "title": "How to download dataset file by file",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372161",
  "author_name": "Victor Rocco",
  "post_date": "2022-12-14T12:38:04.985000",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi!, i am trying to download the dataset file by file but i have a problem, the API only shows near 100 files instead of 54.7k files.</p>\n<ul>\n<li><p>command line:<br>\nkaggle competitions files rsna-breast-cancer-detection -v | wc -l<br>\nresult: 102</p></li>\n<li><p>python API:<br>\nfile_paths: list = kaggleapi.competition_list_files(competition='rsna-breast-cancer-detection')<br>\nprint(type(file_paths[0]), len(file_paths))<br>\nresult:  101</p></li>\n<li><p>python API:<br>\nfile_paths: list = kaggleapi.competitions_data_list_files(id='rsna-breast-cancer-detection')<br>\nprint(type(file_paths[0]), len(file_paths))<br>\nresult:  101</p></li>\n</ul>\n<p>I don't have enough space on disk, so i am thinking on processing each dicom one by one, substantially reducing the size of each image. Also, I also need to do a custom pre-processing, so the published resized datasets doesn't work for me.</p>\n<p>Thanks you in advance,<br>\nVictor</p>",
  "messages": [
    {
      "id": 2065241,
      "postDate": "2022-12-14T12:38:04.987Z",
      "content": "<p>Hi!, i am trying to download the dataset file by file but i have a problem, the API only shows near 100 files instead of 54.7k files.</p>\n<ul>\n<li><p>command line:<br>\nkaggle competitions files rsna-breast-cancer-detection -v | wc -l<br>\nresult: 102</p></li>\n<li><p>python API:<br>\nfile_paths: list = kaggleapi.competition_list_files(competition='rsna-breast-cancer-detection')<br>\nprint(type(file_paths[0]), len(file_paths))<br>\nresult:  101</p></li>\n<li><p>python API:<br>\nfile_paths: list = kaggleapi.competitions_data_list_files(id='rsna-breast-cancer-detection')<br>\nprint(type(file_paths[0]), len(file_paths))<br>\nresult:  101</p></li>\n</ul>\n<p>I don't have enough space on disk, so i am thinking on processing each dicom one by one, substantially reducing the size of each image. Also, I also need to do a custom pre-processing, so the published resized datasets doesn't work for me.</p>\n<p>Thanks you in advance,<br>\nVictor</p>",
      "rawMarkdown": "Hi!, i am trying to download the dataset file by file but i have a problem, the API only shows near 100 files instead of 54.7k files.\n\n- command line:\nkaggle competitions files rsna-breast-cancer-detection -v | wc -l\nresult: 102\n\n- python API:\nfile_paths: list = kaggleapi.competition_list_files(competition='rsna-breast-cancer-detection')\nprint(type(file_paths[0]), len(file_paths))\nresult: <class 'kaggle.models.kaggle_models_extended.File'> 101\n\n- python API:\nfile_paths: list = kaggleapi.competitions_data_list_files(id='rsna-breast-cancer-detection')\nprint(type(file_paths[0]), len(file_paths))\nresult: <class 'dict'> 101\n\n\nI don't have enough space on disk, so i am thinking on processing each dicom one by one, substantially reducing the size of each image. Also, I also need to do a custom pre-processing, so the published resized datasets doesn't work for me.\n\nThanks you in advance,\nVictor\n \n\n ",
      "votes": 5
    },
    {
      "id": 2066428,
      "postDate": "2022-12-15T17:14:36.927Z",
      "content": "<p>Yes, that's exactly what I did, as in the comment above. Here is a sample code. Then download the file.</p>\n<pre><code>input_path = \nstart = \ncount = \n\n zipfile.ZipFile(  \n    ,  \n    ,  \n    compression=zipfile.ZIP_DEFLATED,  \n    compresslevel=)  zf:  \n     i  (start, start + count):\n        r = df_train.iloc[i]\n        path_image = \n        path_to_zip = \n        file = \n        zf.write(file, path_to_zip)\n</code></pre>",
      "rawMarkdown": "Yes, that's exactly what I did, as in the comment above. Here is a sample code. Then download the file.\n\n```python\ninput_path = '/kaggle/input/rsna-breast-cancer-detection/'\nstart = 0\ncount = 350\n\nwith zipfile.ZipFile(  \n    f'/kaggle/working/part.zip',  \n    'w',  \n    compression=zipfile.ZIP_DEFLATED,  \n    compresslevel=9) as zf:  \n    for i in range(start, start + count):\n        r = df_train.iloc[i]\n        path_image = f'{r.patient_id}/{r.image_id}.dcm'\n        path_to_zip = f'{r.patient_id}_{r.image_id}.dcm'\n        file = f'{input_path}train_images/{path_image}'\n        zf.write(file, path_to_zip)\n```\n",
      "votes": 2,
      "replies": [
        {
          "id": 2066441,
          "postDate": "2022-12-15T17:40:42.087Z",
          "content": "<p>Thank you for the code, i will try it.</p>",
          "rawMarkdown": "Thank you for the code, i will try it."
        }
      ]
    },
    {
      "id": 2065878,
      "postDate": "2022-12-15T06:42:00.220Z",
      "content": "<p>Why not just use Kaggle Notebook, extract only the files you want to use, and compress them into a zip?</p>\n<p>It could be left as dicom files or converted to png images.</p>",
      "rawMarkdown": "Why not just use Kaggle Notebook, extract only the files you want to use, and compress them into a zip?\n\nIt could be left as dicom files or converted to png images.",
      "votes": 2,
      "replies": [
        {
          "id": 2066440,
          "postDate": "2022-12-15T17:39:55.510Z",
          "content": "<p>Thank you, you are right, that's a possible workaround.</p>",
          "rawMarkdown": "Thank you, you are right, that's a possible workaround.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2066428,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2022-12-15T17:14:36.927000",
      "content": "<p>Yes, that's exactly what I did, as in the comment above. Here is a sample code. Then download the file.</p>\n<pre><code>input_path = \nstart = \ncount = \n\n zipfile.ZipFile(  \n    ,  \n    ,  \n    compression=zipfile.ZIP_DEFLATED,  \n    compresslevel=)  zf:  \n     i  (start, start + count):\n        r = df_train.iloc[i]\n        path_image = \n        path_to_zip = \n        file = \n        zf.write(file, path_to_zip)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2066441,
          "author_name": "Victor Rocco",
          "author_url": "",
          "post_date": "2022-12-15T17:40:42.087000",
          "content": "<p>Thank you for the code, i will try it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2065878,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2022-12-15T06:42:00.220000",
      "content": "<p>Why not just use Kaggle Notebook, extract only the files you want to use, and compress them into a zip?</p>\n<p>It could be left as dicom files or converted to png images.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2066440,
          "author_name": "Victor Rocco",
          "author_url": "",
          "post_date": "2022-12-15T17:39:55.510000",
          "content": "<p>Thank you, you are right, that's a possible workaround.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2065241": "Hi!, i am trying to download the dataset file by file but i have a problem, the API only shows near 100 files instead of 54.7k files.\n\n- command line:\nkaggle competitions files rsna-breast-cancer-detection -v | wc -l\nresult: 102\n\n- python API:\nfile_paths: list = kaggleapi.competition_list_files(competition='rsna-breast-cancer-detection')\nprint(type(file_paths[0]), len(file_paths))\nresult: <class 'kaggle.models.kaggle_models_extended.File'> 101\n\n- python API:\nfile_paths: list = kaggleapi.competitions_data_list_files(id='rsna-breast-cancer-detection')\nprint(type(file_paths[0]), len(file_paths))\nresult: <class 'dict'> 101\n\n\nI don't have enough space on disk, so i am thinking on processing each dicom one by one, substantially reducing the size of each image. Also, I also need to do a custom pre-processing, so the published resized datasets doesn't work for me.\n\nThanks you in advance,\nVictor\n \n\n ",
    "2066428": "Yes, that's exactly what I did, as in the comment above. Here is a sample code. Then download the file.\n\n```python\ninput_path = '/kaggle/input/rsna-breast-cancer-detection/'\nstart = 0\ncount = 350\n\nwith zipfile.ZipFile(  \n    f'/kaggle/working/part.zip',  \n    'w',  \n    compression=zipfile.ZIP_DEFLATED,  \n    compresslevel=9) as zf:  \n    for i in range(start, start + count):\n        r = df_train.iloc[i]\n        path_image = f'{r.patient_id}/{r.image_id}.dcm'\n        path_to_zip = f'{r.patient_id}_{r.image_id}.dcm'\n        file = f'{input_path}train_images/{path_image}'\n        zf.write(file, path_to_zip)\n```\n",
    "2065878": "Why not just use Kaggle Notebook, extract only the files you want to use, and compress them into a zip?\n\nIt could be left as dicom files or converted to png images."
  }
}