{
  "id": 409951,
  "title": "Rule questions: additional data and submissions",
  "url": "/competitions/asl-fingerspelling/discussion/409951",
  "author_name": "flg",
  "post_date": "2023-05-13T10:16:48.449000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks for running this competition! This looks really interesting but also quite challenging, so I had some questions before starting:</p>\n<h3>Submissions</h3>\n<p>When making a submission it tells me my submission needs to be a notebook, when trying to submit a notebook it tells me, my submission must be a submission.zip. This part isn't entirely clear to me: </p>\n<ol>\n<li>I put my model (and potentially an inference_args.json) inside the zip file, but where do I put the zip, into /kaggle/working?</li>\n<li>What's the structure inside the zip - what does my model have to be called? \"model.tflite\"?</li>\n<li>Is the code in the notebook irrelevant (only the zip matters)?</li>\n<li>Just to be sure: the notebooks submitted are only for inference. Training can be done with arbitrary hardware offline?</li>\n<li>What accelerators are used for inference? Does the one I choose when submitting matter?</li>\n</ol>\n<h3>Data usage</h3>\n<ol>\n<li>Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?</li>\n<li>Is it allowed to manually alter, augment, relabel the training data?</li>\n<li>If I record myself doing ASL (poorly) - I have to share that too if I use it for training?</li>\n</ol>\n<p>Thanks a lot and looking forward to this!</p>",
  "messages": [
    {
      "id": 2258095,
      "postDate": "2023-05-13T21:45:07.507Z",
      "content": "<p>There are successful submissions in the previous competition here<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs</a><br>\nThese will show you how to make submit</p>",
      "rawMarkdown": "There are successful submissions in the previous competition here\nhttps://www.kaggle.com/competitions/asl-signs\nThese will show you how to make submit",
      "votes": 1
    },
    {
      "id": 2257721,
      "postDate": "2023-05-13T15:57:09.510Z",
      "content": "<pre><code>Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n</code></pre>\n<p>From my previous experience, the data must be publicly available (i.e. free) somewhere on the Internet, you don't necessarily need to mention it here, I suppose. As a fact, mentioning external data links are not mentioned in the official rules:</p>\n<pre><code>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n</code></pre>\n<pre><code>Is it allowed to manually alter, augment, relabel the training data?\n</code></pre>\n<p>You can do almost everything that you want with data, at least everything from this list. Augmentations are basic techniques to regularize and make the model more robust to different data. If you mean alter as preprocessing, I can't imagine how you are going to train your models without basic normalization and resizing (e.g. for CV). Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. </p>\n<pre><code>If I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n</code></pre>\n<p>Nope, you can create data by yourself, this is essentially synthetic data from another perspective view… So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\". </p>",
      "rawMarkdown": "```\nDo ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n```\n\nFrom my previous experience, the data must be publicly available (i.e. free) somewhere on the Internet, you don't necessarily need to mention it here, I suppose. As a fact, mentioning external data links are not mentioned in the official rules:\n\n```\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n```\n\n```\nIs it allowed to manually alter, augment, relabel the training data?\n```\n\nYou can do almost everything that you want with data, at least everything from this list. Augmentations are basic techniques to regularize and make the model more robust to different data. If you mean alter as preprocessing, I can't imagine how you are going to train your models without basic normalization and resizing (e.g. for CV). Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. \n\n```\nIf I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n```\n\nNope, you can create data by yourself, this is essentially synthetic data from another perspective view... So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\". ",
      "votes": 1,
      "replies": [
        {
          "id": 2259944,
          "postDate": "2023-05-15T10:37:10.827Z",
          "content": "<blockquote>\n  <p>Nope, you can create data by yourself, this is essentially synthetic data from another perspective view… So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\".</p>\n</blockquote>\n<p>In the previous competition, <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989\" target=\"_blank\">I got the opposite answer from the kaggle staff</a></p>\n<blockquote>\n  <p>Mykola: I treat this as that you are not required to create a discussion with a dataset link you are using. It is rather a restriction of using datasets that are only accessible to PhD students or smth like that. It would be perfect if someone from kaggle stuff clarify this.</p>\n  <p>Another interesting question - using your own dataset. If I collect my own ASL dataset - is it treated as an external dataset? Do you need to publish it somewhere in order to satisfy ensure the External Data is publicly available?</p>\n  <p>Ashley Chow: Thanks <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a>! Yes, data other than the Competition Data is considered as “External Data”, and thus it would need to be \"publicly available and equally accessible to use by all participants of the Competition\". Thanks!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989\" target=\"_blank\">Link</a></p>",
          "rawMarkdown": "> Nope, you can create data by yourself, this is essentially synthetic data from another perspective view… So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\".\n\nIn the previous competition, [I got the opposite answer from the kaggle staff](https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989)\n\n> Mykola: I treat this as that you are not required to create a discussion with a dataset link you are using. It is rather a restriction of using datasets that are only accessible to PhD students or smth like that. It would be perfect if someone from kaggle stuff clarify this.\n\n> Another interesting question - using your own dataset. If I collect my own ASL dataset - is it treated as an external dataset? Do you need to publish it somewhere in order to satisfy ensure the External Data is publicly available?\n\n> Ashley Chow: Thanks @meowmeowmeowmeowmeow! Yes, data other than the Competition Data is considered as “External Data”, and thus it would need to be \"publicly available and equally accessible to use by all participants of the Competition\". Thanks!\n\n[Link](https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989)",
          "votes": 2,
          "replies": [
            {
              "id": 2260288,
              "postDate": "2023-05-15T15:27:05.730Z",
              "content": "<p>Thanks a lot! Some of the other questions in that thread about \"publicly available\" were exactly what went through my head too. Super helpful!</p>",
              "rawMarkdown": "Thanks a lot! Some of the other questions in that thread about \"publicly available\" were exactly what went through my head too. Super helpful!",
              "votes": 1
            }
          ]
        },
        {
          "id": 2260290,
          "postDate": "2023-05-15T15:29:39.203Z",
          "content": "<p>Thanks for the comment!</p>\n<p>\"You can do almost everything that you want with data, at least everything from this list. […]. Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. \"<br>\nI was mostly worried about manually labeling, which is sometimes forbidden. The rules for example say \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\" But that seems to be a bit vague in that hand labeling might refer to test data or not…</p>",
          "rawMarkdown": "Thanks for the comment!\n\n\"You can do almost everything that you want with data, at least everything from this list. [...]. Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. \"\nI was mostly worried about manually labeling, which is sometimes forbidden. The rules for example say \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\" But that seems to be a bit vague in that hand labeling might refer to test data or not...",
          "votes": 1,
          "replies": [
            {
              "id": 2261517,
              "postDate": "2023-05-16T11:36:00.640Z",
              "content": "<p>Because you submit a model (<code>.tflite</code> file) you do not actually have access to the test set so it's impossible to manually label it.</p>",
              "rawMarkdown": "Because you submit a model (`.tflite` file) you do not actually have access to the test set so it's impossible to manually label it."
            }
          ]
        }
      ]
    },
    {
      "id": 2257395,
      "postDate": "2023-05-13T10:16:48.450Z",
      "content": "<p>Thanks for running this competition! This looks really interesting but also quite challenging, so I had some questions before starting:</p>\n<h3>Submissions</h3>\n<p>When making a submission it tells me my submission needs to be a notebook, when trying to submit a notebook it tells me, my submission must be a submission.zip. This part isn't entirely clear to me: </p>\n<ol>\n<li>I put my model (and potentially an inference_args.json) inside the zip file, but where do I put the zip, into /kaggle/working?</li>\n<li>What's the structure inside the zip - what does my model have to be called? \"model.tflite\"?</li>\n<li>Is the code in the notebook irrelevant (only the zip matters)?</li>\n<li>Just to be sure: the notebooks submitted are only for inference. Training can be done with arbitrary hardware offline?</li>\n<li>What accelerators are used for inference? Does the one I choose when submitting matter?</li>\n</ol>\n<h3>Data usage</h3>\n<ol>\n<li>Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?</li>\n<li>Is it allowed to manually alter, augment, relabel the training data?</li>\n<li>If I record myself doing ASL (poorly) - I have to share that too if I use it for training?</li>\n</ol>\n<p>Thanks a lot and looking forward to this!</p>",
      "rawMarkdown": "Thanks for running this competition! This looks really interesting but also quite challenging, so I had some questions before starting:\n\n### Submissions\n\nWhen making a submission it tells me my submission needs to be a notebook, when trying to submit a notebook it tells me, my submission must be a submission.zip. This part isn't entirely clear to me: \n\n1. I put my model (and potentially an inference_args.json) inside the zip file, but where do I put the zip, into /kaggle/working?\n2. What's the structure inside the zip - what does my model have to be called? \"model.tflite\"?\n3. Is the code in the notebook irrelevant (only the zip matters)?\n4. Just to be sure: the notebooks submitted are only for inference. Training can be done with arbitrary hardware offline?\n5. What accelerators are used for inference? Does the one I choose when submitting matter?\n\n### Data usage\n\n6. Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n7. Is it allowed to manually alter, augment, relabel the training data?\n8. If I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n\nThanks a lot and looking forward to this!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2258095,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-05-13T21:45:07.507000",
      "content": "<p>There are successful submissions in the previous competition here<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs</a><br>\nThese will show you how to make submit</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2257721,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2023-05-13T15:57:09.510000",
      "content": "<pre><code>Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n</code></pre>\n<p>From my previous experience, the data must be publicly available (i.e. free) somewhere on the Internet, you don't necessarily need to mention it here, I suppose. As a fact, mentioning external data links are not mentioned in the official rules:</p>\n<pre><code>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n</code></pre>\n<pre><code>Is it allowed to manually alter, augment, relabel the training data?\n</code></pre>\n<p>You can do almost everything that you want with data, at least everything from this list. Augmentations are basic techniques to regularize and make the model more robust to different data. If you mean alter as preprocessing, I can't imagine how you are going to train your models without basic normalization and resizing (e.g. for CV). Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. </p>\n<pre><code>If I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n</code></pre>\n<p>Nope, you can create data by yourself, this is essentially synthetic data from another perspective view… So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\". </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2259944,
          "author_name": "Mykola",
          "author_url": "",
          "post_date": "2023-05-15T10:37:10.827000",
          "content": "<blockquote>\n  <p>Nope, you can create data by yourself, this is essentially synthetic data from another perspective view… So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\".</p>\n</blockquote>\n<p>In the previous competition, <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989\" target=\"_blank\">I got the opposite answer from the kaggle staff</a></p>\n<blockquote>\n  <p>Mykola: I treat this as that you are not required to create a discussion with a dataset link you are using. It is rather a restriction of using datasets that are only accessible to PhD students or smth like that. It would be perfect if someone from kaggle stuff clarify this.</p>\n  <p>Another interesting question - using your own dataset. If I collect my own ASL dataset - is it treated as an external dataset? Do you need to publish it somewhere in order to satisfy ensure the External Data is publicly available?</p>\n  <p>Ashley Chow: Thanks <a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a>! Yes, data other than the Competition Data is considered as “External Data”, and thus it would need to be \"publicly available and equally accessible to use by all participants of the Competition\". Thanks!</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/399822#2211989\" target=\"_blank\">Link</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2260288,
              "author_name": "flg",
              "author_url": "",
              "post_date": "2023-05-15T15:27:05.730000",
              "content": "<p>Thanks a lot! Some of the other questions in that thread about \"publicly available\" were exactly what went through my head too. Super helpful!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2260290,
          "author_name": "flg",
          "author_url": "",
          "post_date": "2023-05-15T15:29:39.203000",
          "content": "<p>Thanks for the comment!</p>\n<p>\"You can do almost everything that you want with data, at least everything from this list. […]. Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. \"<br>\nI was mostly worried about manually labeling, which is sometimes forbidden. The rules for example say \"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\" But that seems to be a bit vague in that hand labeling might refer to test data or not…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2261517,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-05-16T11:36:00.640000",
              "content": "<p>Because you submit a model (<code>.tflite</code> file) you do not actually have access to the test set so it's impossible to manually label it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2258095": "There are successful submissions in the previous competition here\nhttps://www.kaggle.com/competitions/asl-signs\nThese will show you how to make submit",
    "2257721": "```\nDo ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n```\n\nFrom my previous experience, the data must be publicly available (i.e. free) somewhere on the Internet, you don't necessarily need to mention it here, I suppose. As a fact, mentioning external data links are not mentioned in the official rules:\n\n```\nC. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions. However, you will ensure the External Data is publicly available and equally accessible to use by all participants of the Competition for purposes of the competition at no cost to the other participants. The ability to use External Data under this Section 7.C (External Data) does not limit your other obligations under these Competition Rules, including but not limited to Section 11 (Winners Obligations).\n```\n\n```\nIs it allowed to manually alter, augment, relabel the training data?\n```\n\nYou can do almost everything that you want with data, at least everything from this list. Augmentations are basic techniques to regularize and make the model more robust to different data. If you mean alter as preprocessing, I can't imagine how you are going to train your models without basic normalization and resizing (e.g. for CV). Relabeling was one of the key techniques in some competitions (e.g. chaii) and it is still a good technique almost everywhere. \n\n```\nIf I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n```\n\nNope, you can create data by yourself, this is essentially synthetic data from another perspective view... So, I am sure there won't be a problem with that because it is your work (data). My team had fun adding our data to the Feedback Prize dataset saying \"These samples will increase model generalization ability by a lot))\". ",
    "2257395": "Thanks for running this competition! This looks really interesting but also quite challenging, so I had some questions before starting:\n\n### Submissions\n\nWhen making a submission it tells me my submission needs to be a notebook, when trying to submit a notebook it tells me, my submission must be a submission.zip. This part isn't entirely clear to me: \n\n1. I put my model (and potentially an inference_args.json) inside the zip file, but where do I put the zip, into /kaggle/working?\n2. What's the structure inside the zip - what does my model have to be called? \"model.tflite\"?\n3. Is the code in the notebook irrelevant (only the zip matters)?\n4. Just to be sure: the notebooks submitted are only for inference. Training can be done with arbitrary hardware offline?\n5. What accelerators are used for inference? Does the one I choose when submitting matter?\n\n### Data usage\n\n6. Do ppl have to share what 3rd party data they use (e.g. in a Discussion thread)?\n7. Is it allowed to manually alter, augment, relabel the training data?\n8. If I record myself doing ASL (poorly) - I have to share that too if I use it for training?\n\nThanks a lot and looking forward to this!"
  }
}