{
  "id": 410747,
  "title": "Weird samples?",
  "url": "/competitions/asl-fingerspelling/discussion/410747",
  "author_name": "Vadim Irtlach",
  "post_date": "2023-05-16T12:47:34.926000",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.</p>\n<p><a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
  "messages": [
    {
      "id": 2261593,
      "postDate": "2023-05-16T12:47:34.927Z",
      "content": "<p>What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.</p>\n<p><a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
      "rawMarkdown": "What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.\n\n<a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
      "votes": 11
    },
    {
      "id": 2261813,
      "postDate": "2023-05-16T15:06:55.530Z",
      "content": "<p>Please fix your image, it says \"image not found\"</p>",
      "rawMarkdown": "Please fix your image, it says \"image not found\"",
      "votes": 2,
      "replies": [
        {
          "id": 2261829,
          "postDate": "2023-05-16T15:14:33.620Z",
          "content": "<p>Sorry, Mark. Here is the new image</p>\n<p><a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
          "rawMarkdown": "Sorry, Mark. Here is the new image\n\n<a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
          "votes": 1,
          "replies": [
            {
              "id": 2262118,
              "postDate": "2023-05-16T18:34:30.787Z",
              "content": "<p>Thanks and sharp observation.</p>\n<p>There seem to be plenty of single frame samples labelled with long phrases as shown below.<br>\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&amp;alt=media\" alt=\"\"></p>\n<p>It might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.</p>\n<p>The general take away from your observation I believe is the dataset is quite noisy.</p>\n<p>I hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thanks and sharp observation.\n\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&alt=media)\n\nIt might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.\n\nThe general take away from your observation I believe is the dataset is quite noisy.\n\nI hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&alt=media)",
              "votes": 5
            },
            {
              "id": 2262141,
              "postDate": "2023-05-16T18:50:57.470Z",
              "content": "<pre><code>There seem to be plenty of single frame samples labelled with long phrases as shown below.\n</code></pre>\n<p>Yep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. </p>\n<pre><code>The general take away from your observation I believe is the dataset is quite noisy.\n</code></pre>\n<p>Yes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"</p>\n<pre><code>I hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n</code></pre>\n<p>I hope so. </p>",
              "rawMarkdown": "```\nThere seem to be plenty of single frame samples labelled with long phrases as shown below.\n```\n\nYep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. \n\n\n```\nThe general take away from your observation I believe is the dataset is quite noisy.\n```\n\nYes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"\n\n```\nI hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n```\n\nI hope so. ",
              "votes": 1
            },
            {
              "id": 2295935,
              "postDate": "2023-06-11T12:06:09.260Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@sohier \nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&alt=media)",
              "votes": 2
            },
            {
              "id": 2296130,
              "postDate": "2023-06-11T15:06:06Z",
              "content": "<p>there are also samples with only 1 frame 👀</p>",
              "rawMarkdown": "there are also samples with only 1 frame 👀",
              "votes": 1
            },
            {
              "id": 2369334,
              "postDate": "2023-08-01T17:07:51.867Z",
              "content": "<p>Hey! Did you find an answer if the test dataset would have a similar pattern?</p>",
              "rawMarkdown": "Hey! Did you find an answer if the test dataset would have a similar pattern?"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2261813,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-05-16T15:06:55.530000",
      "content": "<p>Please fix your image, it says \"image not found\"</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2261829,
          "author_name": "Vadim Irtlach",
          "author_url": "",
          "post_date": "2023-05-16T15:14:33.620000",
          "content": "<p>Sorry, Mark. Here is the new image</p>\n<p><a href=\"https://ibb.co/RzqQvcb\"><img src=\"https://i.ibb.co/86HPK9s/Screenshot-4.png\" alt=\"Screenshot-4\"></a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2262118,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-05-16T18:34:30.787000",
              "content": "<p>Thanks and sharp observation.</p>\n<p>There seem to be plenty of single frame samples labelled with long phrases as shown below.<br>\nThese are only recordings with a single frame from 30 parquet files, meaning there can be many more samples containing a few frames labelled, incorrectly, with long phrases.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F7e57ce00464d1b244087daec8b7c7768%2Ftemp.png?generation=1684261467580951&amp;alt=media\" alt=\"\"></p>\n<p>It might be an idea to exclude the samples from the first peak shown below, as it is it is unlikely phrases can be fingerspelled in a few frames.</p>\n<p>The general take away from your observation I believe is the dataset is quite noisy.</p>\n<p>I hope, and expect, the test set will contain a manually picked subset with only correctly labelled samples, but this should be confirmed by the competition host.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F724f6ab398a3f59dfc33e52b22e8fa99%2Fabc.png?generation=1684262126411883&amp;alt=media\" alt=\"\"></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2262141,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-16T18:50:57.470000",
              "content": "<pre><code>There seem to be plenty of single frame samples labelled with long phrases as shown below.\n</code></pre>\n<p>Yep, I expected this. Thank you for your \"deep\" and almost full (only 30 parquet files) analysis, I just did an analysis on the subset of data (~1000 samples from the first parquet files). I appreciate such community intersections. </p>\n<pre><code>The general take away from your observation I believe is the dataset is quite noisy.\n</code></pre>\n<p>Yes, of course. We need some time to check it out, but the first hypothesis: \"maybe something was missed during preparation? Bug in software?\"</p>\n<pre><code>I hope, and expect, the test set will contain a manually picked subset with only correctly labeled samples, but this should be confirmed by the competition host.\n</code></pre>\n<p>I hope so. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2295935,
              "author_name": "kurupical",
              "author_url": "",
              "post_date": "2023-06-11T12:06:09.260000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThere are more than 1000 frames of length = 6. Can we assume that such a strange data set is included in the test data set?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1146523%2Fce2ee3765184be1bbf33c5fbf49cf95b%2F2023-06-11%20210355.png?generation=1686485056328741&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2296130,
              "author_name": "yuanzhe zhou",
              "author_url": "",
              "post_date": "2023-06-11T15:06:06",
              "content": "<p>there are also samples with only 1 frame 👀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2369334,
              "author_name": "Quentin Parrenin",
              "author_url": "",
              "post_date": "2023-08-01T17:07:51.867000",
              "content": "<p>Hey! Did you find an answer if the test dataset would have a similar pattern?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2261593": "What are these samples? It looks very weird that there are only a few frames, but the labels are very long to show for such a small number of frames.\n\n<a href=\"https://ibb.co/tYyNrLb\"><img src=\"https://i.ibb.co/gtsXK4D/Screenshot-4.png\" alt=\"Screenshot-4\" border=\"0\"></a>",
    "2261813": "Please fix your image, it says \"image not found\""
  }
}