{
  "id": 218311,
  "title": "Selectively downloading images from the data via filter variables.",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/218311",
  "author_name": "Erwin John T. Carpio",
  "post_date": "2021-02-10T05:11:02.056000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Greetings, is there a way to selectively download specific images? without downloadingthe entire 100+gigabytes? Let's just say I want to peruse  the data set online with kaggle (i can do this now) but it's manual and takes forever… Instead I want to filter out all the images with label \"A\" and download that first….</p>",
  "messages": [
    {
      "id": 1194213,
      "postDate": "2021-02-10T05:11:02.057Z",
      "content": "<p>Greetings, is there a way to selectively download specific images? without downloadingthe entire 100+gigabytes? Let's just say I want to peruse  the data set online with kaggle (i can do this now) but it's manual and takes forever… Instead I want to filter out all the images with label \"A\" and download that first….</p>",
      "rawMarkdown": "Greetings, is there a way to selectively download specific images? without downloadingthe entire 100+gigabytes? Let's just say I want to peruse  the data set online with kaggle (i can do this now) but it's manual and takes forever... Instead I want to filter out all the images with label \"A\" and download that first....",
      "votes": 1
    },
    {
      "id": 1194897,
      "postDate": "2021-02-10T12:54:16.793Z",
      "content": "<p>I don't think there's a simple way to do that via the website or the API. But here's one way of doing this: </p>\n<ol>\n<li>You create a Kaggle kernel that selects the relevant (e.g. <code>train[train['class_name']=='Cardiomegaly']</code> image IDs (<code>image_id</code>) from the .csv file.</li>\n<li>You write the relevant images to a zip file (there's e.g. the <code>zipfile</code> Python package).</li>\n<li>You selectively download that zip file (you can do that when you have run your notebook and look at the outputs section of the saved results in the viewer).</li>\n</ol>",
      "rawMarkdown": "I don't think there's a simple way to do that via the website or the API. But here's one way of doing this: \n1. You create a Kaggle kernel that selects the relevant (e.g. `train[train['class_name']=='Cardiomegaly']` image IDs (`image_id`) from the .csv file.\n2. You write the relevant images to a zip file (there's e.g. the `zipfile` Python package).\n3. You selectively download that zip file (you can do that when you have run your notebook and look at the outputs section of the saved results in the viewer).",
      "replies": [
        {
          "id": 1225158,
          "postDate": "2021-03-03T11:23:32.667Z",
          "content": "<p>Hi Bjorn,thank you for the reply. I haven't been able to return to the competition. Got buried in work. Really appreciate that you took the   time to answer my question…. Thanks again.</p>",
          "rawMarkdown": "Hi Bjorn,thank you for the reply. I haven't been able to return to the competition. Got buried in work. Really appreciate that you took the   time to answer my question.... Thanks again.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1194897,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-10T12:54:16.793000",
      "content": "<p>I don't think there's a simple way to do that via the website or the API. But here's one way of doing this: </p>\n<ol>\n<li>You create a Kaggle kernel that selects the relevant (e.g. <code>train[train['class_name']=='Cardiomegaly']</code> image IDs (<code>image_id</code>) from the .csv file.</li>\n<li>You write the relevant images to a zip file (there's e.g. the <code>zipfile</code> Python package).</li>\n<li>You selectively download that zip file (you can do that when you have run your notebook and look at the outputs section of the saved results in the viewer).</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 1225158,
          "author_name": "Erwin John T. Carpio",
          "author_url": "",
          "post_date": "2021-03-03T11:23:32.667000",
          "content": "<p>Hi Bjorn,thank you for the reply. I haven't been able to return to the competition. Got buried in work. Really appreciate that you took the   time to answer my question…. Thanks again.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1194213": "Greetings, is there a way to selectively download specific images? without downloadingthe entire 100+gigabytes? Let's just say I want to peruse  the data set online with kaggle (i can do this now) but it's manual and takes forever... Instead I want to filter out all the images with label \"A\" and download that first....",
    "1194897": "I don't think there's a simple way to do that via the website or the API. But here's one way of doing this: \n1. You create a Kaggle kernel that selects the relevant (e.g. `train[train['class_name']=='Cardiomegaly']` image IDs (`image_id`) from the .csv file.\n2. You write the relevant images to a zip file (there's e.g. the `zipfile` Python package).\n3. You selectively download that zip file (you can do that when you have run your notebook and look at the outputs section of the saved results in the viewer)."
  }
}