{
  "id": 416148,
  "title": "Update to a small fraction of the train set",
  "url": "/competitions/asl-fingerspelling/discussion/416148",
  "author_name": "Sohier Dane",
  "post_date": "2023-06-09T20:22:50.708000",
  "votes": 6,
  "comment_count": 17,
  "views": 0,
  "content": "<p>We intentionally included adult-themed content in our training dataset to ensure the Deaf and Hard of Hearing community can engage with technology on equal footing with other adults. Out of an abundance of caution, we scrubbed the dataset and removed phrases that were excessively inappropriate, implied illegal content, or wrongly (and randomly) associated names with questionable content. This update accounts for approximately 0.1% of the original train set. We've also shuffled the order of the data to ensure the preview of the training set meets our policy. </p>\n<p>Thanks for your understanding and happy modeling!</p>",
  "messages": [
    {
      "id": 2294158,
      "postDate": "2023-06-09T20:22:50.710Z",
      "content": "<p>We intentionally included adult-themed content in our training dataset to ensure the Deaf and Hard of Hearing community can engage with technology on equal footing with other adults. Out of an abundance of caution, we scrubbed the dataset and removed phrases that were excessively inappropriate, implied illegal content, or wrongly (and randomly) associated names with questionable content. This update accounts for approximately 0.1% of the original train set. We've also shuffled the order of the data to ensure the preview of the training set meets our policy. </p>\n<p>Thanks for your understanding and happy modeling!</p>",
      "rawMarkdown": "We intentionally included adult-themed content in our training dataset to ensure the Deaf and Hard of Hearing community can engage with technology on equal footing with other adults. Out of an abundance of caution, we scrubbed the dataset and removed phrases that were excessively inappropriate, implied illegal content, or wrongly (and randomly) associated names with questionable content. This update accounts for approximately 0.1% of the original train set. We've also shuffled the order of the data to ensure the preview of the training set meets our policy. \n\nThanks for your understanding and happy modeling!",
      "votes": 6
    },
    {
      "id": 2294567,
      "postDate": "2023-06-10T06:52:23.537Z",
      "content": "<p>Great to see that Kaggle took appropriate actions in addressing the concerns. </p>\n<p>While the dataset change may introduce some fairness concerns, I personally think it would be great for users to voluntarily make transition to the updated dataset just to ensure a level playing field.😁</p>",
      "rawMarkdown": "Great to see that Kaggle took appropriate actions in addressing the concerns. \n\nWhile the dataset change may introduce some fairness concerns, I personally think it would be great for users to voluntarily make transition to the updated dataset just to ensure a level playing field.😁",
      "votes": 1
    },
    {
      "id": 2294357,
      "postDate": "2023-06-10T02:39:44.293Z",
      "content": "<p>Did you also changed the test set accordingly, so that they come from the same distribution? If not, it seems to me that it's better to use the old training set…even 0.1% can make a difference between winning and losing solutions. And if so, I hope you will allow us to download and save the old set before reapplying the change so that we have the same chance at winning as those who already did so. </p>",
      "rawMarkdown": "Did you also changed the test set accordingly, so that they come from the same distribution? If not, it seems to me that it's better to use the old training set...even 0.1% can make a difference between winning and losing solutions. And if so, I hope you will allow us to download and save the old set before reapplying the change so that we have the same chance at winning as those who already did so. ",
      "votes": 1,
      "replies": [
        {
          "id": 2294369,
          "postDate": "2023-06-10T02:44:09.473Z",
          "content": "<p>On second thought, even if you did change the test set, It would probably still be better to use the old data set since it has more data…this abrupt change looks like a potentially unfair advantage for those that would keep training with the old set. </p>",
          "rawMarkdown": "On second thought, even if you did change the test set, It would probably still be better to use the old data set since it has more data...this abrupt change looks like a potentially unfair advantage for those that would keep training with the old set. "
        },
        {
          "id": 2294377,
          "postDate": "2023-06-10T02:47:14.750Z",
          "content": "<p>That's an interesting point. Are we allowed to use deleted data?</p>\n<p>For instance, may I upload the deleted part to a Kaggle dataset or elsewhere and make it open-sourced? According to the rules, as I understand them correctly, this should be allowed.</p>",
          "rawMarkdown": "That's an interesting point. Are we allowed to use deleted data?\n\nFor instance, may I upload the deleted part to a Kaggle dataset or elsewhere and make it open-sourced? According to the rules, as I understand them correctly, this should be allowed.",
          "votes": 2,
          "replies": [
            {
              "id": 2294421,
              "postDate": "2023-06-10T04:26:49.183Z",
              "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> I would encourage you not to do so. I believe you'll find that this is material that you want associated with your account. As for the rules, please refer to <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle's community guidelines</a>. We made a deliberate exception for this competition after an internal review of broader issues of fairness to the Deaf community and the severity of the competition material. That exception would not apply to a community dataset, particularly one specifically posting material that we have taken down.</p>",
              "rawMarkdown": "@kolyaforrat I would encourage you not to do so. I believe you'll find that this is material that you want associated with your account. As for the rules, please refer to [Kaggle's community guidelines](https://www.kaggle.com/community-guidelines). We made a deliberate exception for this competition after an internal review of broader issues of fairness to the Deaf community and the severity of the competition material. That exception would not apply to a community dataset, particularly one specifically posting material that we have taken down.",
              "votes": 2
            },
            {
              "id": 2294681,
              "postDate": "2023-06-10T08:27:03.547Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2294682,
              "postDate": "2023-06-10T08:27:28.353Z",
              "content": "<p>I am confused, does this mean that we are forced to delete and re-download the dataset on our end?</p>",
              "rawMarkdown": "I am confused, does this mean that we are forced to delete and re-download the dataset on our end?"
            },
            {
              "id": 2294701,
              "postDate": "2023-06-10T08:42:59.473Z",
              "content": "<p>Apparently not, you are free to train with the old data set if you are lucky enough to have a copy of it.</p>",
              "rawMarkdown": "Apparently not, you are free to train with the old data set if you are lucky enough to have a copy of it."
            },
            {
              "id": 2295197,
              "postDate": "2023-06-10T17:03:22.003Z",
              "content": "<p>If you already have a copy of the data you can keep it. We aren't going to try to impose the burden of re-downloading a 0.2 TB dataset.</p>",
              "rawMarkdown": "If you already have a copy of the data you can keep it. We aren't going to try to impose the burden of re-downloading a 0.2 TB dataset.",
              "votes": 2
            },
            {
              "id": 2295208,
              "postDate": "2023-06-10T17:16:29.813Z",
              "content": "<p>I checked the deleted data. It's only 200 rows. And yeah, you are right, I really don't want to upload it from my profile 😅</p>",
              "rawMarkdown": "I checked the deleted data. It's only 200 rows. And yeah, you are right, I really don't want to upload it from my profile 😅",
              "votes": 1
            }
          ]
        },
        {
          "id": 2294417,
          "postDate": "2023-06-10T04:14:45.983Z",
          "content": "<p>The test set is unchanged.</p>\n<p>Given that the train set is quite a healthy size and the <a href=\"https://www.kaggle.com/competitions/asl-signs/leaderboard?\" target=\"_blank\">leaderboard scores we saw on the previous competition</a> I would be very surprised if this change affects who wins. I would encourage anyone who does think that amount of data will be material to invest some time in creating or cleaning their own additional training data.</p>",
          "rawMarkdown": "The test set is unchanged.\n\nGiven that the train set is quite a healthy size and the [leaderboard scores we saw on the previous competition](https://www.kaggle.com/competitions/asl-signs/leaderboard?) I would be very surprised if this change affects who wins. I would encourage anyone who does think that amount of data will be material to invest some time in creating or cleaning their own additional training data."
        }
      ]
    },
    {
      "id": 2294531,
      "postDate": "2023-06-10T06:20:24.523Z",
      "content": "<p>Did this affect train, supplemental or both?  </p>",
      "rawMarkdown": "Did this affect train, supplemental or both?  ",
      "replies": [
        {
          "id": 2295196,
          "postDate": "2023-06-10T17:01:45.847Z",
          "content": "<p>Just train.</p>",
          "rawMarkdown": "Just train.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2294303,
      "postDate": "2023-06-10T00:17:23.803Z",
      "content": "<p>meet with 404 error when I try to download all data</p>",
      "rawMarkdown": "meet with 404 error when I try to download all data",
      "replies": [
        {
          "id": 2294407,
          "postDate": "2023-06-10T04:00:32.767Z",
          "content": "<p>Thanks for flagging this. Given the time of day, it's probably easiest if I just kick off a fresh upload. The whole process will take a few hours.</p>",
          "rawMarkdown": "Thanks for flagging this. Given the time of day, it's probably easiest if I just kick off a fresh upload. The whole process will take a few hours.",
          "replies": [
            {
              "id": 2294780,
              "postDate": "2023-06-10T10:08:34.293Z",
              "content": "<p>OK, Thanks</p>",
              "rawMarkdown": "OK, Thanks"
            }
          ]
        }
      ]
    },
    {
      "id": 2295065,
      "postDate": "2023-06-10T15:01:50.633Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2294567,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "2023-06-10T06:52:23.537000",
      "content": "<p>Great to see that Kaggle took appropriate actions in addressing the concerns. </p>\n<p>While the dataset change may introduce some fairness concerns, I personally think it would be great for users to voluntarily make transition to the updated dataset just to ensure a level playing field.😁</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2294357,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2023-06-10T02:39:44.293000",
      "content": "<p>Did you also changed the test set accordingly, so that they come from the same distribution? If not, it seems to me that it's better to use the old training set…even 0.1% can make a difference between winning and losing solutions. And if so, I hope you will allow us to download and save the old set before reapplying the change so that we have the same chance at winning as those who already did so. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2294369,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2023-06-10T02:44:09.473000",
          "content": "<p>On second thought, even if you did change the test set, It would probably still be better to use the old data set since it has more data…this abrupt change looks like a potentially unfair advantage for those that would keep training with the old set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2294377,
          "author_name": "Kolya Forrat",
          "author_url": "",
          "post_date": "2023-06-10T02:47:14.750000",
          "content": "<p>That's an interesting point. Are we allowed to use deleted data?</p>\n<p>For instance, may I upload the deleted part to a Kaggle dataset or elsewhere and make it open-sourced? According to the rules, as I understand them correctly, this should be allowed.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2294421,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-06-10T04:26:49.183000",
              "content": "<p><a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> I would encourage you not to do so. I believe you'll find that this is material that you want associated with your account. As for the rules, please refer to <a href=\"https://www.kaggle.com/community-guidelines\" target=\"_blank\">Kaggle's community guidelines</a>. We made a deliberate exception for this competition after an internal review of broader issues of fairness to the Deaf community and the severity of the competition material. That exception would not apply to a community dataset, particularly one specifically posting material that we have taken down.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2294681,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-10T08:27:03.547000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2294682,
              "author_name": "Roman Yakunin",
              "author_url": "",
              "post_date": "2023-06-10T08:27:28.353000",
              "content": "<p>I am confused, does this mean that we are forced to delete and re-download the dataset on our end?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2294701,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-06-10T08:42:59.473000",
              "content": "<p>Apparently not, you are free to train with the old data set if you are lucky enough to have a copy of it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2295197,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2023-06-10T17:03:22.003000",
              "content": "<p>If you already have a copy of the data you can keep it. We aren't going to try to impose the burden of re-downloading a 0.2 TB dataset.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2295208,
              "author_name": "Kolya Forrat",
              "author_url": "",
              "post_date": "2023-06-10T17:16:29.813000",
              "content": "<p>I checked the deleted data. It's only 200 rows. And yeah, you are right, I really don't want to upload it from my profile 😅</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2294417,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-06-10T04:14:45.983000",
          "content": "<p>The test set is unchanged.</p>\n<p>Given that the train set is quite a healthy size and the <a href=\"https://www.kaggle.com/competitions/asl-signs/leaderboard?\" target=\"_blank\">leaderboard scores we saw on the previous competition</a> I would be very surprised if this change affects who wins. I would encourage anyone who does think that amount of data will be material to invest some time in creating or cleaning their own additional training data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2294531,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-06-10T06:20:24.523000",
      "content": "<p>Did this affect train, supplemental or both?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2295196,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-06-10T17:01:45.847000",
          "content": "<p>Just train.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2294303,
      "author_name": "Xiang Huang",
      "author_url": "",
      "post_date": "2023-06-10T00:17:23.803000",
      "content": "<p>meet with 404 error when I try to download all data</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2294407,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2023-06-10T04:00:32.767000",
          "content": "<p>Thanks for flagging this. Given the time of day, it's probably easiest if I just kick off a fresh upload. The whole process will take a few hours.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2294780,
              "author_name": "Xiang Huang",
              "author_url": "",
              "post_date": "2023-06-10T10:08:34.293000",
              "content": "<p>OK, Thanks</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2295065,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-10T15:01:50.633000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2294158": "We intentionally included adult-themed content in our training dataset to ensure the Deaf and Hard of Hearing community can engage with technology on equal footing with other adults. Out of an abundance of caution, we scrubbed the dataset and removed phrases that were excessively inappropriate, implied illegal content, or wrongly (and randomly) associated names with questionable content. This update accounts for approximately 0.1% of the original train set. We've also shuffled the order of the data to ensure the preview of the training set meets our policy. \n\nThanks for your understanding and happy modeling!",
    "2294567": "Great to see that Kaggle took appropriate actions in addressing the concerns. \n\nWhile the dataset change may introduce some fairness concerns, I personally think it would be great for users to voluntarily make transition to the updated dataset just to ensure a level playing field.😁",
    "2294357": "Did you also changed the test set accordingly, so that they come from the same distribution? If not, it seems to me that it's better to use the old training set...even 0.1% can make a difference between winning and losing solutions. And if so, I hope you will allow us to download and save the old set before reapplying the change so that we have the same chance at winning as those who already did so. ",
    "2294531": "Did this affect train, supplemental or both?  ",
    "2294303": "meet with 404 error when I try to download all data",
    "2295065": ""
  }
}