{
  "id": 369301,
  "title": "question about getting data via kaggle API",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369301",
  "author_name": "Maxim Shatskiy",
  "post_date": "2022-11-29T17:11:19.382000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am using kaggle API for the first time.</p>\n<p><code>kaggle competitions files rsna-breast-cancer-detection</code></p>\n<p>list only 94 items from train_images. I would expect that it returns all files, which belong to this competition. Any suggestion on what can be wrong?</p>",
  "messages": [
    {
      "id": 2048780,
      "postDate": "2022-11-29T17:11:19.383Z",
      "content": "<p>I am using kaggle API for the first time.</p>\n<p><code>kaggle competitions files rsna-breast-cancer-detection</code></p>\n<p>list only 94 items from train_images. I would expect that it returns all files, which belong to this competition. Any suggestion on what can be wrong?</p>",
      "rawMarkdown": "I am using kaggle API for the first time.\n\n`kaggle competitions files rsna-breast-cancer-detection`\n\nlist only 94 items from train_images. I would expect that it returns all files, which belong to this competition. Any suggestion on what can be wrong?",
      "votes": 1
    },
    {
      "id": 2049054,
      "postDate": "2022-11-29T22:43:07.943Z",
      "content": "<p>I'll have to look into why it's only listing those files, but the actual download size is 270 GB so I think you will actually get the complete dataset via the API. I'll run another test on my end to verify if this is the case.</p>\n<p>Update: the API definitely delivers many more files than it lists. I would go ahead and use the API for the download.</p>",
      "rawMarkdown": "I'll have to look into why it's only listing those files, but the actual download size is 270 GB so I think you will actually get the complete dataset via the API. I'll run another test on my end to verify if this is the case.\n\nUpdate: the API definitely delivers many more files than it lists. I would go ahead and use the API for the download.",
      "replies": [
        {
          "id": 2050019,
          "postDate": "2022-11-30T13:38:17.770Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThanks for the reply. I know that I can get the entire 270GB in a single file with API. However, due to size it would make sense to download it based on some batches.<br>\nI hoped that the listing would provide me with all the files and I get as many as I needed.<br>\nHowever, as mentioned it returns some samples of the files available and it seems to me this is some kind of random sample. I guess it might be some mechanism to protect kaggle from many requests.</p>\n<p>So you can see the problems with using API at the following picture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3139%2F20bbd3f618deabed5508ffd6e65047c1%2FUnbenannt.PNG?generation=1669815331736605&amp;alt=media\" alt=\"\"><br>\nand if it gets interrupted then ….</p>",
          "rawMarkdown": "Hi @sohier \nThanks for the reply. I know that I can get the entire 270GB in a single file with API. However, due to size it would make sense to download it based on some batches.\nI hoped that the listing would provide me with all the files and I get as many as I needed.\nHowever, as mentioned it returns some samples of the files available and it seems to me this is some kind of random sample. I guess it might be some mechanism to protect kaggle from many requests.\n\nSo you can see the problems with using API at the following picture:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3139%2F20bbd3f618deabed5508ffd6e65047c1%2FUnbenannt.PNG?generation=1669815331736605&alt=media)\nand if it gets interrupted then ...."
        }
      ]
    },
    {
      "id": 2050016,
      "postDate": "2022-11-30T13:37:18.093Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2049054,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2022-11-29T22:43:07.943000",
      "content": "<p>I'll have to look into why it's only listing those files, but the actual download size is 270 GB so I think you will actually get the complete dataset via the API. I'll run another test on my end to verify if this is the case.</p>\n<p>Update: the API definitely delivers many more files than it lists. I would go ahead and use the API for the download.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2050019,
          "author_name": "Maxim Shatskiy",
          "author_url": "",
          "post_date": "2022-11-30T13:38:17.770000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThanks for the reply. I know that I can get the entire 270GB in a single file with API. However, due to size it would make sense to download it based on some batches.<br>\nI hoped that the listing would provide me with all the files and I get as many as I needed.<br>\nHowever, as mentioned it returns some samples of the files available and it seems to me this is some kind of random sample. I guess it might be some mechanism to protect kaggle from many requests.</p>\n<p>So you can see the problems with using API at the following picture:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3139%2F20bbd3f618deabed5508ffd6e65047c1%2FUnbenannt.PNG?generation=1669815331736605&amp;alt=media\" alt=\"\"><br>\nand if it gets interrupted then ….</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2050016,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-30T13:37:18.093000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2048780": "I am using kaggle API for the first time.\n\n`kaggle competitions files rsna-breast-cancer-detection`\n\nlist only 94 items from train_images. I would expect that it returns all files, which belong to this competition. Any suggestion on what can be wrong?",
    "2049054": "I'll have to look into why it's only listing those files, but the actual download size is 270 GB so I think you will actually get the complete dataset via the API. I'll run another test on my end to verify if this is the case.\n\nUpdate: the API definitely delivers many more files than it lists. I would go ahead and use the API for the download.",
    "2050016": ""
  }
}