{
  "id": 211693,
  "title": "How did you download the dataset? ",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/211693",
  "author_name": "Berkay Alan",
  "post_date": "2021-01-16T08:00:54.842000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>This question may be stupid but I am new at this level and try to understand.</p>\n<p>How did you download the dataset? It's so big. Do you directly work in Kaggle notebook or is there any other way to download?</p>",
  "messages": [
    {
      "id": 1155107,
      "postDate": "2021-01-16T08:00:54.843Z",
      "content": "<p>Hi all,</p>\n<p>This question may be stupid but I am new at this level and try to understand.</p>\n<p>How did you download the dataset? It's so big. Do you directly work in Kaggle notebook or is there any other way to download?</p>",
      "rawMarkdown": "Hi all,\n\nThis question may be stupid but I am new at this level and try to understand.\n\nHow did you download the dataset? It's so big. Do you directly work in Kaggle notebook or is there any other way to download?",
      "votes": 3
    },
    {
      "id": 1156001,
      "postDate": "2021-01-16T20:43:33.970Z",
      "content": "<p><a href=\"https://www.kaggle.com/berkayalan\" target=\"_blank\">@berkayalan</a> one technique I used was to rescale all the images into some standard format (say 400 x 500), convert them to 2D <code>numpy</code> arrays in memory, then create a dict that mapped <code>image_id</code> to the specific <code>numpy</code> array. I then used <code>pickle</code> to store the dict in a binary file that I could share between notebooks (or download). The resulting pickle file for the training images was 2.9 GB in size. This fit nicely into the VM's memory space, and was still detailed enough for me to see. I'm still experimenting with optimal scaling size to see if dropping resolution down this far is going to work. I haven't had a need to download it, but 2.9 GB for me is acceptable if I wanted to use my own hardware for experimentation.</p>",
      "rawMarkdown": "@berkayalan one technique I used was to rescale all the images into some standard format (say 400 x 500), convert them to 2D `numpy` arrays in memory, then create a dict that mapped `image_id` to the specific `numpy` array. I then used `pickle` to store the dict in a binary file that I could share between notebooks (or download). The resulting pickle file for the training images was 2.9 GB in size. This fit nicely into the VM's memory space, and was still detailed enough for me to see. I'm still experimenting with optimal scaling size to see if dropping resolution down this far is going to work. I haven't had a need to download it, but 2.9 GB for me is acceptable if I wanted to use my own hardware for experimentation.",
      "votes": 1,
      "replies": [
        {
          "id": 1156016,
          "postDate": "2021-01-16T21:10:26.527Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a> , I will try this way.</p>",
          "rawMarkdown": "Thanks @craigmthomas , I will try this way."
        }
      ]
    },
    {
      "id": 1155260,
      "postDate": "2021-01-16T10:38:04.980Z",
      "content": "<p>If you have a good wifi and computational power then you can download it. But if you don't have any of the things then you should try working directly in the kaggle notebook itself. <a href=\"https://www.kaggle.com/berkayalan\" target=\"_blank\">@berkayalan</a> </p>",
      "rawMarkdown": "If you have a good wifi and computational power then you can download it. But if you don't have any of the things then you should try working directly in the kaggle notebook itself. @berkayalan ",
      "votes": 1,
      "replies": [
        {
          "id": 1155385,
          "postDate": "2021-01-16T11:22:59.377Z",
          "content": "<p>Thanks so much for the info <a href=\"https://www.kaggle.com/saurabhshahane\" target=\"_blank\">@saurabhshahane</a> </p>",
          "rawMarkdown": "Thanks so much for the info @saurabhshahane ",
          "votes": 1
        },
        {
          "id": 1167706,
          "postDate": "2021-01-24T12:46:27.993Z",
          "content": "<p>There's no need of a good wifi to download a dataset. You can simply use <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a>. <br>\nIt is super fast even if you have poor connection. You can download any competition dataset, your own private dataset and public dataset to your local machine(or in colab session). <br>\nFollow <a href=\"https://medium.com/@ankushchoubey/how-to-download-dataset-from-kaggle-7f700d7f9198\" target=\"_blank\">this blog</a> if you haven't done it before.</p>",
          "rawMarkdown": "There's no need of a good wifi to download a dataset. You can simply use [Kaggle API](https://github.com/Kaggle/kaggle-api). \nIt is super fast even if you have poor connection. You can download any competition dataset, your own private dataset and public dataset to your local machine(or in colab session). \nFollow [this blog](https://medium.com/@ankushchoubey/how-to-download-dataset-from-kaggle-7f700d7f9198) if you haven't done it before.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1168075,
      "postDate": "2021-01-24T17:25:37.963Z",
      "content": "<p>You can also use Kaggle API<br>\n<a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">https://github.com/Kaggle/kaggle-api</a></p>",
      "rawMarkdown": "You can also use Kaggle API\nhttps://github.com/Kaggle/kaggle-api"
    }
  ],
  "comments": [
    {
      "id": 1156001,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-16T20:43:33.970000",
      "content": "<p><a href=\"https://www.kaggle.com/berkayalan\" target=\"_blank\">@berkayalan</a> one technique I used was to rescale all the images into some standard format (say 400 x 500), convert them to 2D <code>numpy</code> arrays in memory, then create a dict that mapped <code>image_id</code> to the specific <code>numpy</code> array. I then used <code>pickle</code> to store the dict in a binary file that I could share between notebooks (or download). The resulting pickle file for the training images was 2.9 GB in size. This fit nicely into the VM's memory space, and was still detailed enough for me to see. I'm still experimenting with optimal scaling size to see if dropping resolution down this far is going to work. I haven't had a need to download it, but 2.9 GB for me is acceptable if I wanted to use my own hardware for experimentation.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1156016,
          "author_name": "Berkay Alan",
          "author_url": "",
          "post_date": "2021-01-16T21:10:26.527000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a> , I will try this way.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1155260,
      "author_name": "Saurabh Shahane",
      "author_url": "",
      "post_date": "2021-01-16T10:38:04.980000",
      "content": "<p>If you have a good wifi and computational power then you can download it. But if you don't have any of the things then you should try working directly in the kaggle notebook itself. <a href=\"https://www.kaggle.com/berkayalan\" target=\"_blank\">@berkayalan</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1155385,
          "author_name": "Berkay Alan",
          "author_url": "",
          "post_date": "2021-01-16T11:22:59.377000",
          "content": "<p>Thanks so much for the info <a href=\"https://www.kaggle.com/saurabhshahane\" target=\"_blank\">@saurabhshahane</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1167706,
          "author_name": "Zaber Ibn Abdul Hakim",
          "author_url": "",
          "post_date": "2021-01-24T12:46:27.993000",
          "content": "<p>There's no need of a good wifi to download a dataset. You can simply use <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">Kaggle API</a>. <br>\nIt is super fast even if you have poor connection. You can download any competition dataset, your own private dataset and public dataset to your local machine(or in colab session). <br>\nFollow <a href=\"https://medium.com/@ankushchoubey/how-to-download-dataset-from-kaggle-7f700d7f9198\" target=\"_blank\">this blog</a> if you haven't done it before.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1168075,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-01-24T17:25:37.963000",
      "content": "<p>You can also use Kaggle API<br>\n<a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">https://github.com/Kaggle/kaggle-api</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1155107": "Hi all,\n\nThis question may be stupid but I am new at this level and try to understand.\n\nHow did you download the dataset? It's so big. Do you directly work in Kaggle notebook or is there any other way to download?",
    "1156001": "@berkayalan one technique I used was to rescale all the images into some standard format (say 400 x 500), convert them to 2D `numpy` arrays in memory, then create a dict that mapped `image_id` to the specific `numpy` array. I then used `pickle` to store the dict in a binary file that I could share between notebooks (or download). The resulting pickle file for the training images was 2.9 GB in size. This fit nicely into the VM's memory space, and was still detailed enough for me to see. I'm still experimenting with optimal scaling size to see if dropping resolution down this far is going to work. I haven't had a need to download it, but 2.9 GB for me is acceptable if I wanted to use my own hardware for experimentation.",
    "1155260": "If you have a good wifi and computational power then you can download it. But if you don't have any of the things then you should try working directly in the kaggle notebook itself. @berkayalan ",
    "1168075": "You can also use Kaggle API\nhttps://github.com/Kaggle/kaggle-api"
  }
}