{
  "id": 66908,
  "title": "Welcome!",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/66908",
  "author_name": "inversion",
  "post_date": "2018-09-26T16:27:29.909000",
  "votes": 17,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Welcome to the Quick, Draw! Doodle Recognition Challenge!</p>\n\n<p>In this challenge, you are tasked with using vector drawing input to predict what the user is doodling. </p>\n\n<p>There are two versions of the data: Raw and Simplified. The Simplified data is much smaller, and contains nearly all of the relevant information. While you can access the Raw Train data from the competition data page, due to its size (205 Gb uncompressed), it is not included in the Kernels environment. Both the Raw and Simplified Test files are included in Kernels, as well as the Simplified Train data.</p>\n\n<p>There are labels in the Train data that are multi-word, delimited with a space. These align with the previously-released Quick Draw dataset. In order for the metric to parse these correctly, you will need to replace multi-word label spaces with an underscore (e.g., \"roller coaster\" to \"roller_coaster\").</p>\n\n<p>This <a href=\"https://www.kaggle.com/inversion/getting-started-viewing-quick-draw-doodles-etc\">Getting Started Kernel</a> will help you get started. </p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": 394301,
      "postDate": "2018-09-26T16:27:29.910Z",
      "content": "<p>Welcome to the Quick, Draw! Doodle Recognition Challenge!</p>\n\n<p>In this challenge, you are tasked with using vector drawing input to predict what the user is doodling. </p>\n\n<p>There are two versions of the data: Raw and Simplified. The Simplified data is much smaller, and contains nearly all of the relevant information. While you can access the Raw Train data from the competition data page, due to its size (205 Gb uncompressed), it is not included in the Kernels environment. Both the Raw and Simplified Test files are included in Kernels, as well as the Simplified Train data.</p>\n\n<p>There are labels in the Train data that are multi-word, delimited with a space. These align with the previously-released Quick Draw dataset. In order for the metric to parse these correctly, you will need to replace multi-word label spaces with an underscore (e.g., \"roller coaster\" to \"roller_coaster\").</p>\n\n<p>This <a href=\"https://www.kaggle.com/inversion/getting-started-viewing-quick-draw-doodles-etc\">Getting Started Kernel</a> will help you get started. </p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Welcome to the Quick, Draw! Doodle Recognition Challenge!\n\nIn this challenge, you are tasked with using vector drawing input to predict what the user is doodling. \n\nThere are two versions of the data: Raw and Simplified. The Simplified data is much smaller, and contains nearly all of the relevant information. While you can access the Raw Train data from the competition data page, due to its size (205 Gb uncompressed), it is not included in the Kernels environment. Both the Raw and Simplified Test files are included in Kernels, as well as the Simplified Train data.\n\nThere are labels in the Train data that are multi-word, delimited with a space. These align with the previously-released Quick Draw dataset. In order for the metric to parse these correctly, you will need to replace multi-word label spaces with an underscore (e.g., \"roller coaster\" to \"roller_coaster\").\n\nThis [Getting Started Kernel][1] will help you get started. \n\nPlease feel free to ask your questions in this thread.\n\nGood luck!\n\n\n  [1]: https://www.kaggle.com/inversion/getting-started-viewing-quick-draw-doodles-etc",
      "votes": 16
    },
    {
      "id": 404236,
      "postDate": "2018-10-15T13:17:39.277Z",
      "content": "<ol>\n<li><p>How is the test dataset actually labeled? Is it:\na. Show the annotator a picture and ask the annotator to pick the best label\nb. Show the annotator a picture and the original (noisy) label, accepting only when the annotator agrees with the original</p></li>\n<li><p>Some words have multiple meanings: e.g. bat. Some people draw baseball bats, others draw the mammal bat. In the test set, does this happen? I.e. do some of the baseball bats drawings have ground truth labels of \"bat\" while others have \"baseball bat\"?</p></li>\n<li><p>In the test set, are the \"written\" drawings supposed to be labeled too? Is there a point in training a model to identify if the drawing is in fact a written word, then passing the drawing to a handwriting model?</p></li>\n</ol>\n\n<p>Thanks!</p>",
      "rawMarkdown": "1. How is the test dataset actually labeled? Is it:\na. Show the annotator a picture and ask the annotator to pick the best label\nb. Show the annotator a picture and the original (noisy) label, accepting only when the annotator agrees with the original\n\n2. Some words have multiple meanings: e.g. bat. Some people draw baseball bats, others draw the mammal bat. In the test set, does this happen? I.e. do some of the baseball bats drawings have ground truth labels of \"bat\" while others have \"baseball bat\"?\n\n3. In the test set, are the \"written\" drawings supposed to be labeled too? Is there a point in training a model to identify if the drawing is in fact a written word, then passing the drawing to a handwriting model?\n\nThanks!",
      "votes": 5,
      "replies": [
        {
          "id": 404572,
          "postDate": "2018-10-16T01:44:32.190Z",
          "content": "<p>I find some \"written\" drawings in the test set:</p>\n\n<pre><code>9959710506773887\n9982335663788957\n9005411586993511\n</code></pre>\n\n<p>In the Cdiscount's competition, some people implement OCR in the pipeline, so I will not be surprised if someone use a handwriting model in this competition. :)</p>\n\n<p>About 2), I think \"cup, coffee cup, mug\" is harder than \"bat\".</p>",
          "rawMarkdown": "I find some \"written\" drawings in the test set:\n\n    9959710506773887\n    9982335663788957\n    9005411586993511\n\nIn the Cdiscount's competition, some people implement OCR in the pipeline, so I will not be surprised if someone use a handwriting model in this competition. :)\n\nAbout 2), I think \"cup, coffee cup, mug\" is harder than \"bat\".\n"
        }
      ]
    },
    {
      "id": 394352,
      "postDate": "2018-09-26T17:45:59.857Z",
      "content": "<p>Hi inversion, will Kaggle or Google AI provide free GCP credit this time?</p>",
      "rawMarkdown": "Hi inversion, will Kaggle or Google AI provide free GCP credit this time?",
      "votes": 3,
      "replies": [
        {
          "id": 394456,
          "postDate": "2018-09-26T21:40:00.083Z",
          "content": "<p>There won't be GCP credits for this competition. </p>",
          "rawMarkdown": "There won't be GCP credits for this competition. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 415268,
      "postDate": "2018-11-04T19:32:14.247Z",
      "content": "<p>I still do not get the metric completly. Regarding <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a>\nIs actual always of the length 1? So submitting 3 values is always worse than submitting the correct one only, right?</p>",
      "rawMarkdown": "I still do not get the metric completly. Regarding https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\nIs actual always of the length 1? So submitting 3 values is always worse than submitting the correct one only, right?",
      "votes": 1,
      "replies": [
        {
          "id": 426235,
          "postDate": "2018-11-23T00:04:07.137Z",
          "content": "<p>Tim - Sorry I missed this. The calculation stops when you get the correct label, so if you get the correct label in the first position, it scores 1, regardless of what is after.</p>",
          "rawMarkdown": "Tim - Sorry I missed this. The calculation stops when you get the correct label, so if you get the correct label in the first position, it scores 1, regardless of what is after."
        }
      ]
    },
    {
      "id": 395322,
      "postDate": "2018-09-28T10:15:33.373Z",
      "content": "<p>Can I participate in the competition just using the kaggle kernel? Have computational constraints :(</p>",
      "rawMarkdown": "Can I participate in the competition just using the kaggle kernel? Have computational constraints :(",
      "votes": 1
    },
    {
      "id": 396794,
      "postDate": "2018-10-01T11:42:32.967Z",
      "content": "<p>Hi, I have a question regarding the 'recognized' label. This information is contained in the training set, but not in the test set. So does the test set contain drawings which are not recognized?</p>",
      "rawMarkdown": "Hi, I have a question regarding the 'recognized' label. This information is contained in the training set, but not in the test set. So does the test set contain drawings which are not recognized?",
      "votes": 2
    },
    {
      "id": 426145,
      "postDate": "2018-11-22T18:21:40.877Z",
      "content": "<p>Train have timestamp. Where are timestamp for test set?</p>",
      "rawMarkdown": "Train have timestamp. Where are timestamp for test set?",
      "replies": [
        {
          "id": 426233,
          "postDate": "2018-11-23T00:00:38.347Z",
          "content": "<p>Timestamps are not provided in Test.</p>",
          "rawMarkdown": "Timestamps are not provided in Test."
        }
      ]
    },
    {
      "id": 412242,
      "postDate": "2018-10-29T20:45:55.940Z",
      "content": "<p>Hi,\nit seems like the submission system is broken? I get 0.033 public leaderboard when resubmitting an old file which got 0.902 before...</p>",
      "rawMarkdown": "Hi,\nit seems like the submission system is broken? I get 0.033 public leaderboard when resubmitting an old file which got 0.902 before...",
      "replies": [
        {
          "id": 412667,
          "postDate": "2018-10-30T15:29:24.223Z",
          "content": "<p>Seems like it is fixed...</p>",
          "rawMarkdown": "Seems like it is fixed..."
        },
        {
          "id": 412686,
          "postDate": "2018-10-30T15:55:53.550Z",
          "content": "<p>Yes, the issue has been fixed, and the affected submissions have been rescored. Sorry for the inconvenience.</p>",
          "rawMarkdown": "Yes, the issue has been fixed, and the affected submissions have been rescored. Sorry for the inconvenience."
        }
      ]
    },
    {
      "id": 410297,
      "postDate": "2018-10-25T19:29:48.230Z",
      "content": "<p>Can someone explain me this line: <a href=\"https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32\">https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32</a></p>\n\n<p>Isn't <code>p not in predicted[:i]</code> always true? (provided <code>predicted</code> only contains unique values)</p>",
      "rawMarkdown": "Can someone explain me this line: https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32\n\nIsn't `p not in predicted[:i]` always true? (provided `predicted` only contains unique values)",
      "replies": [
        {
          "id": 410327,
          "postDate": "2018-10-25T20:58:35.420Z",
          "content": "<p>Hi Tim, Yes it is. I assume that line would handle the general case in which you could have repeated items in the prediction.\nWithout that line you would get precisions larger than 1: e.g. if actual=1 and predicted=(1,1) you would have ap@2=2</p>",
          "rawMarkdown": "Hi Tim, Yes it is. I assume that line would handle the general case in which you could have repeated items in the prediction.\nWithout that line you would get precisions larger than 1: e.g. if actual=1 and predicted=(1,1) you would have ap@2=2",
          "votes": 1
        },
        {
          "id": 411282,
          "postDate": "2018-10-27T19:10:31.777Z",
          "content": "<p>Okay, good! Thank you :)</p>",
          "rawMarkdown": "Okay, good! Thank you :)"
        }
      ]
    },
    {
      "id": 409525,
      "postDate": "2018-10-24T12:30:23.777Z",
      "content": "<p>This is my first CV tournament, so forgive me if this has been asked before, but what stops people from hard coding values for the test set? Assuming 5 seconds to look at an image and write down your top guess, it would take about 155 hours to go through the whole training set, so that seems really long. But you could sort by your model's least confident predictions and spend a few hours hard coding the most difficult ones while letting your model handle the ones it predicts confidently. I figured something like this would be against the rules, but I didn't see anything that would forbid this.</p>",
      "rawMarkdown": "This is my first CV tournament, so forgive me if this has been asked before, but what stops people from hard coding values for the test set? Assuming 5 seconds to look at an image and write down your top guess, it would take about 155 hours to go through the whole training set, so that seems really long. But you could sort by your model's least confident predictions and spend a few hours hard coding the most difficult ones while letting your model handle the ones it predicts confidently. I figured something like this would be against the rules, but I didn't see anything that would forbid this.",
      "replies": [
        {
          "id": 409618,
          "postDate": "2018-10-24T15:47:59.747Z",
          "content": "<p>Public score is evaluated on some part of the test set. Not the whole.</p>",
          "rawMarkdown": "Public score is evaluated on some part of the test set. Not the whole."
        },
        {
          "id": 409807,
          "postDate": "2018-10-24T22:12:52.947Z",
          "content": "<p>From the Official Rules: </p>\n\n<blockquote>\n  <p>Entrants may re-annotate images in the training set, but may not\n  hand-label predictions, including having human observers rate and\n  evaluate the test data set.</p>\n</blockquote>",
          "rawMarkdown": "From the Official Rules: \n\n&gt; Entrants may re-annotate images in the training set, but may not\n&gt; hand-label predictions, including having human observers rate and\n&gt; evaluate the test data set.",
          "votes": 1
        },
        {
          "id": 411395,
          "postDate": "2018-10-28T02:47:23.967Z",
          "content": "<p>Ah, thanks! I must have missed that.</p>",
          "rawMarkdown": "Ah, thanks! I must have missed that."
        }
      ]
    },
    {
      "id": 407064,
      "postDate": "2018-10-20T10:12:35.953Z",
      "content": "<p>Cool</p>",
      "rawMarkdown": "Cool"
    },
    {
      "id": 396804,
      "postDate": "2018-10-01T11:58:57.317Z",
      "content": "<p>Is there any workaround for downloading the data....I don't have 73 GB of free space</p>",
      "rawMarkdown": "Is there any workaround for downloading the data....I don't have 73 GB of free space",
      "replies": [
        {
          "id": 396806,
          "postDate": "2018-10-01T12:01:15.130Z",
          "content": "<p>try the train_simplified one</p>",
          "rawMarkdown": "try the train_simplified one",
          "votes": 2
        }
      ]
    },
    {
      "id": 395079,
      "postDate": "2018-09-27T23:31:41.213Z",
      "content": "<p>Hi inversion! Thanks for hosting this competition it looks pretty cool :)</p>\n\n<p>I just had some general questions about the dataset. I have had a chance to at least go through some of the data from the simplified training set. From the overview we know that we have to assign more than one possible label for each image so I thought that over all the datasets I might see the same key id to possibly set up a multi-label classifier and use the key_ids to concatenate known labels for validation. </p>\n\n<p>It seems however that each key_id only occurs once. So my question is, how can we set up a multi-label classifier for images that have more than one label? Or is my understanding of the setup incorrect?</p>\n\n<p>Also, what does the recognized field in the data represent?</p>\n\n<p>Thanks again!</p>",
      "rawMarkdown": "Hi inversion! Thanks for hosting this competition it looks pretty cool :)\n\nI just had some general questions about the dataset. I have had a chance to at least go through some of the data from the simplified training set. From the overview we know that we have to assign more than one possible label for each image so I thought that over all the datasets I might see the same key id to possibly set up a multi-label classifier and use the key_ids to concatenate known labels for validation. \n\nIt seems however that each key_id only occurs once. So my question is, how can we set up a multi-label classifier for images that have more than one label? Or is my understanding of the setup incorrect?\n\nAlso, what does the recognized field in the data represent?\n\nThanks again!",
      "replies": [
        {
          "id": 395418,
          "postDate": "2018-09-28T14:45:45.097Z",
          "content": "<p>You get more than 1 <em>guess</em> for the <em>single</em> label of the drawing.</p>\n\n<p>You can find more about the dataset here:\n<a href=\"https://github.com/googlecreativelab/quickdraw-dataset\">https://github.com/googlecreativelab/quickdraw-dataset</a></p>",
          "rawMarkdown": "You get more than 1 _guess_ for the _single_ label of the drawing.\n\nYou can find more about the dataset here:\nhttps://github.com/googlecreativelab/quickdraw-dataset",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 404236,
      "author_name": "VinceTan",
      "author_url": "",
      "post_date": "2018-10-15T13:17:39.277000",
      "content": "<ol>\n<li><p>How is the test dataset actually labeled? Is it:\na. Show the annotator a picture and ask the annotator to pick the best label\nb. Show the annotator a picture and the original (noisy) label, accepting only when the annotator agrees with the original</p></li>\n<li><p>Some words have multiple meanings: e.g. bat. Some people draw baseball bats, others draw the mammal bat. In the test set, does this happen? I.e. do some of the baseball bats drawings have ground truth labels of \"bat\" while others have \"baseball bat\"?</p></li>\n<li><p>In the test set, are the \"written\" drawings supposed to be labeled too? Is there a point in training a model to identify if the drawing is in fact a written word, then passing the drawing to a handwriting model?</p></li>\n</ol>\n\n<p>Thanks!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 404572,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "2018-10-16T01:44:32.190000",
          "content": "<p>I find some \"written\" drawings in the test set:</p>\n\n<pre><code>9959710506773887\n9982335663788957\n9005411586993511\n</code></pre>\n\n<p>In the Cdiscount's competition, some people implement OCR in the pipeline, so I will not be surprised if someone use a handwriting model in this competition. :)</p>\n\n<p>About 2), I think \"cup, coffee cup, mug\" is harder than \"bat\".</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 394352,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-09-26T17:45:59.857000",
      "content": "<p>Hi inversion, will Kaggle or Google AI provide free GCP credit this time?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 394456,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-09-26T21:40:00.083000",
          "content": "<p>There won't be GCP credits for this competition. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415268,
      "author_name": "Tim Joseph",
      "author_url": "",
      "post_date": "2018-11-04T19:32:14.247000",
      "content": "<p>I still do not get the metric completly. Regarding <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a>\nIs actual always of the length 1? So submitting 3 values is always worse than submitting the correct one only, right?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 426235,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-23T00:04:07.137000",
          "content": "<p>Tim - Sorry I missed this. The calculation stops when you get the correct label, so if you get the correct label in the first position, it scores 1, regardless of what is after.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 395322,
      "author_name": "Hasib Zunair",
      "author_url": "",
      "post_date": "2018-09-28T10:15:33.373000",
      "content": "<p>Can I participate in the competition just using the kaggle kernel? Have computational constraints :(</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 396794,
      "author_name": "Yichen",
      "author_url": "",
      "post_date": "2018-10-01T11:42:32.967000",
      "content": "<p>Hi, I have a question regarding the 'recognized' label. This information is contained in the training set, but not in the test set. So does the test set contain drawings which are not recognized?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 426145,
      "author_name": "Leigh",
      "author_url": "",
      "post_date": "2018-11-22T18:21:40.877000",
      "content": "<p>Train have timestamp. Where are timestamp for test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 426233,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-23T00:00:38.347000",
          "content": "<p>Timestamps are not provided in Test.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 412242,
      "author_name": "Tim Joseph",
      "author_url": "",
      "post_date": "2018-10-29T20:45:55.940000",
      "content": "<p>Hi,\nit seems like the submission system is broken? I get 0.033 public leaderboard when resubmitting an old file which got 0.902 before...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 412667,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2018-10-30T15:29:24.223000",
          "content": "<p>Seems like it is fixed...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 412686,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-10-30T15:55:53.550000",
          "content": "<p>Yes, the issue has been fixed, and the affected submissions have been rescored. Sorry for the inconvenience.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 410297,
      "author_name": "Tim Joseph",
      "author_url": "",
      "post_date": "2018-10-25T19:29:48.230000",
      "content": "<p>Can someone explain me this line: <a href=\"https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32\">https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32</a></p>\n\n<p>Isn't <code>p not in predicted[:i]</code> always true? (provided <code>predicted</code> only contains unique values)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 410327,
          "author_name": "Fantomas",
          "author_url": "",
          "post_date": "2018-10-25T20:58:35.420000",
          "content": "<p>Hi Tim, Yes it is. I assume that line would handle the general case in which you could have repeated items in the prediction.\nWithout that line you would get precisions larger than 1: e.g. if actual=1 and predicted=(1,1) you would have ap@2=2</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411282,
          "author_name": "Tim Joseph",
          "author_url": "",
          "post_date": "2018-10-27T19:10:31.777000",
          "content": "<p>Okay, good! Thank you :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 409525,
      "author_name": "Allen",
      "author_url": "",
      "post_date": "2018-10-24T12:30:23.777000",
      "content": "<p>This is my first CV tournament, so forgive me if this has been asked before, but what stops people from hard coding values for the test set? Assuming 5 seconds to look at an image and write down your top guess, it would take about 155 hours to go through the whole training set, so that seems really long. But you could sort by your model's least confident predictions and spend a few hours hard coding the most difficult ones while letting your model handle the ones it predicts confidently. I figured something like this would be against the rules, but I didn't see anything that would forbid this.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 409618,
          "author_name": "Hasib Zunair",
          "author_url": "",
          "post_date": "2018-10-24T15:47:59.747000",
          "content": "<p>Public score is evaluated on some part of the test set. Not the whole.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409807,
          "author_name": "Jaden F",
          "author_url": "",
          "post_date": "2018-10-24T22:12:52.947000",
          "content": "<p>From the Official Rules: </p>\n\n<blockquote>\n  <p>Entrants may re-annotate images in the training set, but may not\n  hand-label predictions, including having human observers rate and\n  evaluate the test data set.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411395,
          "author_name": "Allen",
          "author_url": "",
          "post_date": "2018-10-28T02:47:23.967000",
          "content": "<p>Ah, thanks! I must have missed that.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 407064,
      "author_name": "Harshit Ahluwalia",
      "author_url": "",
      "post_date": "2018-10-20T10:12:35.953000",
      "content": "<p>Cool</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 396804,
      "author_name": "RevantTiwari",
      "author_url": "",
      "post_date": "2018-10-01T11:58:57.317000",
      "content": "<p>Is there any workaround for downloading the data....I don't have 73 GB of free space</p>",
      "votes": 0,
      "replies": [
        {
          "id": 396806,
          "author_name": "Hasib Zunair",
          "author_url": "",
          "post_date": "2018-10-01T12:01:15.130000",
          "content": "<p>try the train_simplified one</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 395079,
      "author_name": "RDizzl3",
      "author_url": "",
      "post_date": "2018-09-27T23:31:41.213000",
      "content": "<p>Hi inversion! Thanks for hosting this competition it looks pretty cool :)</p>\n\n<p>I just had some general questions about the dataset. I have had a chance to at least go through some of the data from the simplified training set. From the overview we know that we have to assign more than one possible label for each image so I thought that over all the datasets I might see the same key id to possibly set up a multi-label classifier and use the key_ids to concatenate known labels for validation. </p>\n\n<p>It seems however that each key_id only occurs once. So my question is, how can we set up a multi-label classifier for images that have more than one label? Or is my understanding of the setup incorrect?</p>\n\n<p>Also, what does the recognized field in the data represent?</p>\n\n<p>Thanks again!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 395418,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-09-28T14:45:45.097000",
          "content": "<p>You get more than 1 <em>guess</em> for the <em>single</em> label of the drawing.</p>\n\n<p>You can find more about the dataset here:\n<a href=\"https://github.com/googlecreativelab/quickdraw-dataset\">https://github.com/googlecreativelab/quickdraw-dataset</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "394301": "Welcome to the Quick, Draw! Doodle Recognition Challenge!\n\nIn this challenge, you are tasked with using vector drawing input to predict what the user is doodling. \n\nThere are two versions of the data: Raw and Simplified. The Simplified data is much smaller, and contains nearly all of the relevant information. While you can access the Raw Train data from the competition data page, due to its size (205 Gb uncompressed), it is not included in the Kernels environment. Both the Raw and Simplified Test files are included in Kernels, as well as the Simplified Train data.\n\nThere are labels in the Train data that are multi-word, delimited with a space. These align with the previously-released Quick Draw dataset. In order for the metric to parse these correctly, you will need to replace multi-word label spaces with an underscore (e.g., \"roller coaster\" to \"roller_coaster\").\n\nThis [Getting Started Kernel][1] will help you get started. \n\nPlease feel free to ask your questions in this thread.\n\nGood luck!\n\n\n  [1]: https://www.kaggle.com/inversion/getting-started-viewing-quick-draw-doodles-etc",
    "404236": "1. How is the test dataset actually labeled? Is it:\na. Show the annotator a picture and ask the annotator to pick the best label\nb. Show the annotator a picture and the original (noisy) label, accepting only when the annotator agrees with the original\n\n2. Some words have multiple meanings: e.g. bat. Some people draw baseball bats, others draw the mammal bat. In the test set, does this happen? I.e. do some of the baseball bats drawings have ground truth labels of \"bat\" while others have \"baseball bat\"?\n\n3. In the test set, are the \"written\" drawings supposed to be labeled too? Is there a point in training a model to identify if the drawing is in fact a written word, then passing the drawing to a handwriting model?\n\nThanks!",
    "394352": "Hi inversion, will Kaggle or Google AI provide free GCP credit this time?",
    "415268": "I still do not get the metric completly. Regarding https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\nIs actual always of the length 1? So submitting 3 values is always worse than submitting the correct one only, right?",
    "395322": "Can I participate in the competition just using the kaggle kernel? Have computational constraints :(",
    "396794": "Hi, I have a question regarding the 'recognized' label. This information is contained in the training set, but not in the test set. So does the test set contain drawings which are not recognized?",
    "426145": "Train have timestamp. Where are timestamp for test set?",
    "412242": "Hi,\nit seems like the submission system is broken? I get 0.033 public leaderboard when resubmitting an old file which got 0.902 before...",
    "410297": "Can someone explain me this line: https://github.com/benhamner/Metrics/blob/9a637aea795dc6f2333f022b0863398de0a1ca77/Python/ml_metrics/average_precision.py#L32\n\nIsn't `p not in predicted[:i]` always true? (provided `predicted` only contains unique values)",
    "409525": "This is my first CV tournament, so forgive me if this has been asked before, but what stops people from hard coding values for the test set? Assuming 5 seconds to look at an image and write down your top guess, it would take about 155 hours to go through the whole training set, so that seems really long. But you could sort by your model's least confident predictions and spend a few hours hard coding the most difficult ones while letting your model handle the ones it predicts confidently. I figured something like this would be against the rules, but I didn't see anything that would forbid this.",
    "407064": "Cool",
    "396804": "Is there any workaround for downloading the data....I don't have 73 GB of free space",
    "395079": "Hi inversion! Thanks for hosting this competition it looks pretty cool :)\n\nI just had some general questions about the dataset. I have had a chance to at least go through some of the data from the simplified training set. From the overview we know that we have to assign more than one possible label for each image so I thought that over all the datasets I might see the same key id to possibly set up a multi-label classifier and use the key_ids to concatenate known labels for validation. \n\nIt seems however that each key_id only occurs once. So my question is, how can we set up a multi-label classifier for images that have more than one label? Or is my understanding of the setup incorrect?\n\nAlso, what does the recognized field in the data represent?\n\nThanks again!"
  }
}