{
  "id": 410464,
  "title": "Supplemental landmarks",
  "url": "/competitions/asl-fingerspelling/discussion/410464",
  "author_name": "Mathieu De Coster",
  "post_date": "2023-05-15T14:03:50.522000",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In the data there is <code>train.csv</code> and <code>supplemental.csv</code>, as well as <code>train_landmarks/</code> and <code>supplemental_landmarks/</code>. I did not find an explanation about the reason for splitting the dataset examples into these two CSV files and folders.</p>\n<p>What is the purpose of the supplemental files?</p>\n<p><strong>Update</strong> The data description says:</p>\n<blockquote>\n  <p>The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. </p>\n</blockquote>\n<p>Are we allowed to use the supplemental dataset for training?</p>",
  "messages": [
    {
      "id": 2260174,
      "postDate": "2023-05-15T14:03:50.523Z",
      "content": "<p>In the data there is <code>train.csv</code> and <code>supplemental.csv</code>, as well as <code>train_landmarks/</code> and <code>supplemental_landmarks/</code>. I did not find an explanation about the reason for splitting the dataset examples into these two CSV files and folders.</p>\n<p>What is the purpose of the supplemental files?</p>\n<p><strong>Update</strong> The data description says:</p>\n<blockquote>\n  <p>The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. </p>\n</blockquote>\n<p>Are we allowed to use the supplemental dataset for training?</p>",
      "rawMarkdown": "In the data there is `train.csv` and `supplemental.csv`, as well as `train_landmarks/` and `supplemental_landmarks/`. I did not find an explanation about the reason for splitting the dataset examples into these two CSV files and folders.\n\nWhat is the purpose of the supplemental files?\n\n**Update** The data description says:\n\n> The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. \n\nAre we allowed to use the supplemental dataset for training?",
      "votes": 11
    },
    {
      "id": 2260232,
      "postDate": "2023-05-15T14:45:48.943Z",
      "content": "<p>Based on first analysis, it seems like train data is mostly unstructured or barely structured, i.e., things like url's, phone numbers, names. There are also many more unique sequences (occasionally, there are duplicate labels by different signers but not always). <br>\nSupplemental data seems to contain phrases or short sentences. It also far fewer unique sequences and many duplicates (from different signers) per sequence.</p>\n<p>Based on the names ('train' versus 'supplemental') I would guess that the unseen (test) data is similar to the train data while the supplemental is easier to analyse (but also much more prone to overfitting).</p>\n<p>Could the organisers confirm my assumption about the nature of the test data?</p>",
      "rawMarkdown": "Based on first analysis, it seems like train data is mostly unstructured or barely structured, i.e., things like url's, phone numbers, names. There are also many more unique sequences (occasionally, there are duplicate labels by different signers but not always). \nSupplemental data seems to contain phrases or short sentences. It also far fewer unique sequences and many duplicates (from different signers) per sequence.\n\nBased on the names ('train' versus 'supplemental') I would guess that the unseen (test) data is similar to the train data while the supplemental is easier to analyse (but also much more prone to overfitting).\n\nCould the organisers confirm my assumption about the nature of the test data?",
      "votes": 7
    },
    {
      "id": 2281202,
      "postDate": "2023-05-30T17:00:46.307Z",
      "content": "<p>Yes, you can use the supplemental dataset for training.</p>",
      "rawMarkdown": "Yes, you can use the supplemental dataset for training.",
      "votes": 1
    },
    {
      "id": 2285032,
      "postDate": "2023-06-02T12:09:02.837Z",
      "content": "<p>Does this mean that the data used for <strong>scoring</strong> is also randomly generated addresses, phone numbers, and urls?</p>",
      "rawMarkdown": "Does this mean that the data used for **scoring** is also randomly generated addresses, phone numbers, and urls?",
      "replies": [
        {
          "id": 2285439,
          "postDate": "2023-06-02T17:54:10.183Z",
          "content": "<p>That's correct.</p>",
          "rawMarkdown": "That's correct."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2260232,
      "author_name": "Wondering Alice",
      "author_url": "",
      "post_date": "2023-05-15T14:45:48.943000",
      "content": "<p>Based on first analysis, it seems like train data is mostly unstructured or barely structured, i.e., things like url's, phone numbers, names. There are also many more unique sequences (occasionally, there are duplicate labels by different signers but not always). <br>\nSupplemental data seems to contain phrases or short sentences. It also far fewer unique sequences and many duplicates (from different signers) per sequence.</p>\n<p>Based on the names ('train' versus 'supplemental') I would guess that the unseen (test) data is similar to the train data while the supplemental is easier to analyse (but also much more prone to overfitting).</p>\n<p>Could the organisers confirm my assumption about the nature of the test data?</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2281202,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2023-05-30T17:00:46.307000",
      "content": "<p>Yes, you can use the supplemental dataset for training.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2285032,
      "author_name": "stk2666",
      "author_url": "",
      "post_date": "2023-06-02T12:09:02.837000",
      "content": "<p>Does this mean that the data used for <strong>scoring</strong> is also randomly generated addresses, phone numbers, and urls?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2285439,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-06-02T17:54:10.183000",
          "content": "<p>That's correct.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2260174": "In the data there is `train.csv` and `supplemental.csv`, as well as `train_landmarks/` and `supplemental_landmarks/`. I did not find an explanation about the reason for splitting the dataset examples into these two CSV files and folders.\n\nWhat is the purpose of the supplemental files?\n\n**Update** The data description says:\n\n> The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. \n\nAre we allowed to use the supplemental dataset for training?",
    "2260232": "Based on first analysis, it seems like train data is mostly unstructured or barely structured, i.e., things like url's, phone numbers, names. There are also many more unique sequences (occasionally, there are duplicate labels by different signers but not always). \nSupplemental data seems to contain phrases or short sentences. It also far fewer unique sequences and many duplicates (from different signers) per sequence.\n\nBased on the names ('train' versus 'supplemental') I would guess that the unseen (test) data is similar to the train data while the supplemental is easier to analyse (but also much more prone to overfitting).\n\nCould the organisers confirm my assumption about the nature of the test data?",
    "2281202": "Yes, you can use the supplemental dataset for training.",
    "2285032": "Does this mean that the data used for **scoring** is also randomly generated addresses, phone numbers, and urls?"
  }
}