{
  "id": 420709,
  "title": "Framerate of trainingdata",
  "url": "/competitions/asl-fingerspelling/discussion/420709",
  "author_name": "MaxUhl98",
  "post_date": "2023-07-02T07:00:15.616000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Were the videos from which the trainingdata is created captured at the same framerate and if yes does anybody know the fps ? </p>",
  "messages": [
    {
      "id": 2326479,
      "postDate": "2023-07-02T07:00:15.617Z",
      "content": "<p>Were the videos from which the trainingdata is created captured at the same framerate and if yes does anybody know the fps ? </p>",
      "rawMarkdown": "Were the videos from which the trainingdata is created captured at the same framerate and if yes does anybody know the fps ? ",
      "votes": 4
    },
    {
      "id": 2331451,
      "postDate": "2023-07-05T14:54:58.910Z",
      "content": "<p>I've also wondered this, so I'll comment in hopes of getting a confirmed answer. It looks like the average might be around 18 characters and 150 frames. Elsewhere I read that ASL signers typically sign at about 60 words per minute. Average word length is about 5 characters in English. So the signers are signing 3-4 words in 150 frames at 60 words per minute. That's 3-4 seconds of recording in 150 frames, which would be a frame rate of 37.5-50 fps. I'd bet the rate of fingerspelling in these videos is slowed down a little compared with that average, it's fingerspelling addresses and unusual words. 50 wpm? At 50 wpm, we're at 35 fps. 45 wpm, 31 fps. My gut says that these are all recorded at the same frame rate, and it's likely 30 fps… default video recording frame rate on an iPhone anyway… Maybe this isn't important for building a model though.</p>",
      "rawMarkdown": "I've also wondered this, so I'll comment in hopes of getting a confirmed answer. It looks like the average might be around 18 characters and 150 frames. Elsewhere I read that ASL signers typically sign at about 60 words per minute. Average word length is about 5 characters in English. So the signers are signing 3-4 words in 150 frames at 60 words per minute. That's 3-4 seconds of recording in 150 frames, which would be a frame rate of 37.5-50 fps. I'd bet the rate of fingerspelling in these videos is slowed down a little compared with that average, it's fingerspelling addresses and unusual words. 50 wpm? At 50 wpm, we're at 35 fps. 45 wpm, 31 fps. My gut says that these are all recorded at the same frame rate, and it's likely 30 fps... default video recording frame rate on an iPhone anyway... Maybe this isn't important for building a model though.",
      "votes": 1,
      "replies": [
        {
          "id": 2331824,
          "postDate": "2023-07-05T19:56:29.343Z",
          "content": "<p>Thank you! <br>\nAnd yes this is not important for building the model but it is (in my opinion, i am still a beginner) important for processing the trainingdata to feed the model</p>",
          "rawMarkdown": "Thank you! \nAnd yes this is not important for building the model but it is (in my opinion, i am still a beginner) important for processing the trainingdata to feed the model",
          "replies": [
            {
              "id": 2332134,
              "postDate": "2023-07-06T02:29:41.390Z",
              "content": "<p>I'm also a total beginner. Before I really tried to figure out how to build a model on this I wanted to just visualize the fingerspelling and understand more about that side of it, so I built code to make a little gif from a sequence. I had to guess at the framerate, but the only difference it made was how easily I could interpret some of the letters being spelled. Then I saw someone else had done this in a different way: <a href=\"https://www.kaggle.com/code/embeddedravi/eda-visualize-asl-parquet-data-using-matplotlib\" target=\"_blank\">https://www.kaggle.com/code/embeddedravi/eda-visualize-asl-parquet-data-using-matplotlib</a> And interestingly, if you look there, they chose 50 fps. I have not figured out how to build a useful model yet, but my thinking is that you use the frames and don't have to know the rate. That said, if it was recorded at 240 fps, then you could surely cut the cpu workload a lot by ignoring a lot of frames. I will say that anecdotally, skipping frame by frame through the image stack used to make a gif, it didn't seem like much excess (like seeing an e held for 4 frames or something), so my default will be to use all frames.</p>",
              "rawMarkdown": "I'm also a total beginner. Before I really tried to figure out how to build a model on this I wanted to just visualize the fingerspelling and understand more about that side of it, so I built code to make a little gif from a sequence. I had to guess at the framerate, but the only difference it made was how easily I could interpret some of the letters being spelled. Then I saw someone else had done this in a different way: https://www.kaggle.com/code/embeddedravi/eda-visualize-asl-parquet-data-using-matplotlib And interestingly, if you look there, they chose 50 fps. I have not figured out how to build a useful model yet, but my thinking is that you use the frames and don't have to know the rate. That said, if it was recorded at 240 fps, then you could surely cut the cpu workload a lot by ignoring a lot of frames. I will say that anecdotally, skipping frame by frame through the image stack used to make a gif, it didn't seem like much excess (like seeing an e held for 4 frames or something), so my default will be to use all frames."
            },
            {
              "id": 2332218,
              "postDate": "2023-07-06T04:34:25.203Z",
              "content": "<p>I am mostly focussing on preprocessing right now and in that i can use the framerate to calculate which frames i consider to be legit, for example there is data with one frame for multiple letters, which is impossible to detect or fingerspell. The formula i use to check if the amout of frames/character rate is achievable is framerate/5 =30/5=6<br>\n(fingerspellers can spell 60 words per minute, thus 1 word per second, average word lenght in english is 4.7 letters which i round up to five, according to <a href=\"http://fingerspell.sierra-charlie.com/en/\" target=\"_blank\">http://fingerspell.sierra-charlie.com/en/</a> experienced fingerspellers spell 2 characters per second, so maybe the double of that is is the max possible speed which would give us 4 chars per second and 30/4 = 7.5 frames/character) I use this information to delete all data which has a lower frames/character rate then 6, which should give me a cleaner dataset</p>",
              "rawMarkdown": "I am mostly focussing on preprocessing right now and in that i can use the framerate to calculate which frames i consider to be legit, for example there is data with one frame for multiple letters, which is impossible to detect or fingerspell. The formula i use to check if the amout of frames/character rate is achievable is framerate/5 =30/5=6\n(fingerspellers can spell 60 words per minute, thus 1 word per second, average word lenght in english is 4.7 letters which i round up to five, according to http://fingerspell.sierra-charlie.com/en/ experienced fingerspellers spell 2 characters per second, so maybe the double of that is is the max possible speed which would give us 4 chars per second and 30/4 = 7.5 frames/character) I use this information to delete all data which has a lower frames/character rate then 6, which should give me a cleaner dataset"
            },
            {
              "id": 2332964,
              "postDate": "2023-07-06T14:59:48.400Z",
              "content": "<p>I agree, something to get rid of the very-few-frame samples needs to be imposed. To me, your approach seems potentially over-complicated relative to just saying anything less than… 10? frames is not real data. You may look at <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/419909\" target=\"_blank\">this Q&amp;A</a> if you haven't and see the distribution of frame lengths, and the notebook that created those is linked there. There's a lot in that notebook. But it looks to me like there's a pretty high number of sub-8 frame entries. What you suggested makes sense in theory, but it is somewhat important what the frame rate is, or you'd set the boundary differently, right? what if it is 50 fps? My other thinking is just based on gifs that I made- I wrote it so that when it drew the hand, if one of the points was nan, it would skip drawing that segment. However, my output didn't have any hands just missing a section, and something like 122 frames for the whole phrase turned into 65 frames of actual output, but I could see the entire phrase in that 65 frames. Which is just to say, check carefully as you use that approach, I think. Again, not that I know what I'm doing :-)  </p>",
              "rawMarkdown": "I agree, something to get rid of the very-few-frame samples needs to be imposed. To me, your approach seems potentially over-complicated relative to just saying anything less than... 10? frames is not real data. You may look at [this Q&A](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/419909) if you haven't and see the distribution of frame lengths, and the notebook that created those is linked there. There's a lot in that notebook. But it looks to me like there's a pretty high number of sub-8 frame entries. What you suggested makes sense in theory, but it is somewhat important what the frame rate is, or you'd set the boundary differently, right? what if it is 50 fps? My other thinking is just based on gifs that I made- I wrote it so that when it drew the hand, if one of the points was nan, it would skip drawing that segment. However, my output didn't have any hands just missing a section, and something like 122 frames for the whole phrase turned into 65 frames of actual output, but I could see the entire phrase in that 65 frames. Which is just to say, check carefully as you use that approach, I think. Again, not that I know what I'm doing :-)  "
            },
            {
              "id": 2332997,
              "postDate": "2023-07-06T15:17:33.550Z",
              "content": "<p>In my approach the framerate determines the minimum frames per character an entry has to have for me to not delete it (Note: There are a lot of entries that do not spell out the entire phrase given by train.csv so i try to delete the phrases where it is unlikely for the participant to spell that amount of characters in the amount of time which i can calculate via framerate) <br>\nSo in my approch the framerate determines the threshold of frames needed per phrase length for a dataseries to qualify for training. For example assuming 30 fps my threshold would be 6 frames per character in the phrase (spaces count too), with 50 fps my threshold would be 10 characters per frame so the amout of data i cut out will be larger. (i think 10 is definietly too high, since the mean of frames/char is around 9)</p>",
              "rawMarkdown": "In my approach the framerate determines the minimum frames per character an entry has to have for me to not delete it (Note: There are a lot of entries that do not spell out the entire phrase given by train.csv so i try to delete the phrases where it is unlikely for the participant to spell that amount of characters in the amount of time which i can calculate via framerate) \nSo in my approch the framerate determines the threshold of frames needed per phrase length for a dataseries to qualify for training. For example assuming 30 fps my threshold would be 6 frames per character in the phrase (spaces count too), with 50 fps my threshold would be 10 characters per frame so the amout of data i cut out will be larger. (i think 10 is definietly too high, since the mean of frames/char is around 9)"
            },
            {
              "id": 2364029,
              "postDate": "2023-07-29T05:00:28.057Z",
              "content": "<p>I'm assuming the data is recorded at 10FPS - referring to my <a href=\"https://www.kaggle.com/code/nadigshreekanth/data-visualization-using-mediapipe-apis\" target=\"_blank\">data visualization experiments</a>, 10FPS looks \"natural\" to me.</p>\n<blockquote>\n  <p>For example, this can be very useful in graphs that combine a slower GPU inference path (eg, at 10 FPS) with a faster GPU rendering path (eg, at 30 FPS)<br>\n  From <a href=\"https://developers.google.com/mediapipe/framework/framework_concepts/gpu\" target=\"_blank\">https://developers.google.com/mediapipe/framework/framework_concepts/gpu</a></p>\n</blockquote>",
              "rawMarkdown": "I'm assuming the data is recorded at 10FPS - referring to my [data visualization experiments](https://www.kaggle.com/code/nadigshreekanth/data-visualization-using-mediapipe-apis), 10FPS looks \"natural\" to me.\n>For example, this can be very useful in graphs that combine a slower GPU inference path (eg, at 10 FPS) with a faster GPU rendering path (eg, at 30 FPS)\nFrom https://developers.google.com/mediapipe/framework/framework_concepts/gpu"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2331451,
      "author_name": "Kyle Proffitt",
      "author_url": "",
      "post_date": "2023-07-05T14:54:58.910000",
      "content": "<p>I've also wondered this, so I'll comment in hopes of getting a confirmed answer. It looks like the average might be around 18 characters and 150 frames. Elsewhere I read that ASL signers typically sign at about 60 words per minute. Average word length is about 5 characters in English. So the signers are signing 3-4 words in 150 frames at 60 words per minute. That's 3-4 seconds of recording in 150 frames, which would be a frame rate of 37.5-50 fps. I'd bet the rate of fingerspelling in these videos is slowed down a little compared with that average, it's fingerspelling addresses and unusual words. 50 wpm? At 50 wpm, we're at 35 fps. 45 wpm, 31 fps. My gut says that these are all recorded at the same frame rate, and it's likely 30 fps… default video recording frame rate on an iPhone anyway… Maybe this isn't important for building a model though.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2331824,
          "author_name": "MaxUhl98",
          "author_url": "",
          "post_date": "2023-07-05T19:56:29.343000",
          "content": "<p>Thank you! <br>\nAnd yes this is not important for building the model but it is (in my opinion, i am still a beginner) important for processing the trainingdata to feed the model</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2332134,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-07-06T02:29:41.390000",
              "content": "<p>I'm also a total beginner. Before I really tried to figure out how to build a model on this I wanted to just visualize the fingerspelling and understand more about that side of it, so I built code to make a little gif from a sequence. I had to guess at the framerate, but the only difference it made was how easily I could interpret some of the letters being spelled. Then I saw someone else had done this in a different way: <a href=\"https://www.kaggle.com/code/embeddedravi/eda-visualize-asl-parquet-data-using-matplotlib\" target=\"_blank\">https://www.kaggle.com/code/embeddedravi/eda-visualize-asl-parquet-data-using-matplotlib</a> And interestingly, if you look there, they chose 50 fps. I have not figured out how to build a useful model yet, but my thinking is that you use the frames and don't have to know the rate. That said, if it was recorded at 240 fps, then you could surely cut the cpu workload a lot by ignoring a lot of frames. I will say that anecdotally, skipping frame by frame through the image stack used to make a gif, it didn't seem like much excess (like seeing an e held for 4 frames or something), so my default will be to use all frames.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2332218,
              "author_name": "MaxUhl98",
              "author_url": "",
              "post_date": "2023-07-06T04:34:25.203000",
              "content": "<p>I am mostly focussing on preprocessing right now and in that i can use the framerate to calculate which frames i consider to be legit, for example there is data with one frame for multiple letters, which is impossible to detect or fingerspell. The formula i use to check if the amout of frames/character rate is achievable is framerate/5 =30/5=6<br>\n(fingerspellers can spell 60 words per minute, thus 1 word per second, average word lenght in english is 4.7 letters which i round up to five, according to <a href=\"http://fingerspell.sierra-charlie.com/en/\" target=\"_blank\">http://fingerspell.sierra-charlie.com/en/</a> experienced fingerspellers spell 2 characters per second, so maybe the double of that is is the max possible speed which would give us 4 chars per second and 30/4 = 7.5 frames/character) I use this information to delete all data which has a lower frames/character rate then 6, which should give me a cleaner dataset</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2332964,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-07-06T14:59:48.400000",
              "content": "<p>I agree, something to get rid of the very-few-frame samples needs to be imposed. To me, your approach seems potentially over-complicated relative to just saying anything less than… 10? frames is not real data. You may look at <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/419909\" target=\"_blank\">this Q&amp;A</a> if you haven't and see the distribution of frame lengths, and the notebook that created those is linked there. There's a lot in that notebook. But it looks to me like there's a pretty high number of sub-8 frame entries. What you suggested makes sense in theory, but it is somewhat important what the frame rate is, or you'd set the boundary differently, right? what if it is 50 fps? My other thinking is just based on gifs that I made- I wrote it so that when it drew the hand, if one of the points was nan, it would skip drawing that segment. However, my output didn't have any hands just missing a section, and something like 122 frames for the whole phrase turned into 65 frames of actual output, but I could see the entire phrase in that 65 frames. Which is just to say, check carefully as you use that approach, I think. Again, not that I know what I'm doing :-)  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2332997,
              "author_name": "MaxUhl98",
              "author_url": "",
              "post_date": "2023-07-06T15:17:33.550000",
              "content": "<p>In my approach the framerate determines the minimum frames per character an entry has to have for me to not delete it (Note: There are a lot of entries that do not spell out the entire phrase given by train.csv so i try to delete the phrases where it is unlikely for the participant to spell that amount of characters in the amount of time which i can calculate via framerate) <br>\nSo in my approch the framerate determines the threshold of frames needed per phrase length for a dataseries to qualify for training. For example assuming 30 fps my threshold would be 6 frames per character in the phrase (spaces count too), with 50 fps my threshold would be 10 characters per frame so the amout of data i cut out will be larger. (i think 10 is definietly too high, since the mean of frames/char is around 9)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2364029,
              "author_name": "sknadig",
              "author_url": "",
              "post_date": "2023-07-29T05:00:28.057000",
              "content": "<p>I'm assuming the data is recorded at 10FPS - referring to my <a href=\"https://www.kaggle.com/code/nadigshreekanth/data-visualization-using-mediapipe-apis\" target=\"_blank\">data visualization experiments</a>, 10FPS looks \"natural\" to me.</p>\n<blockquote>\n  <p>For example, this can be very useful in graphs that combine a slower GPU inference path (eg, at 10 FPS) with a faster GPU rendering path (eg, at 30 FPS)<br>\n  From <a href=\"https://developers.google.com/mediapipe/framework/framework_concepts/gpu\" target=\"_blank\">https://developers.google.com/mediapipe/framework/framework_concepts/gpu</a></p>\n</blockquote>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2326479": "Were the videos from which the trainingdata is created captured at the same framerate and if yes does anybody know the fps ? ",
    "2331451": "I've also wondered this, so I'll comment in hopes of getting a confirmed answer. It looks like the average might be around 18 characters and 150 frames. Elsewhere I read that ASL signers typically sign at about 60 words per minute. Average word length is about 5 characters in English. So the signers are signing 3-4 words in 150 frames at 60 words per minute. That's 3-4 seconds of recording in 150 frames, which would be a frame rate of 37.5-50 fps. I'd bet the rate of fingerspelling in these videos is slowed down a little compared with that average, it's fingerspelling addresses and unusual words. 50 wpm? At 50 wpm, we're at 35 fps. 45 wpm, 31 fps. My gut says that these are all recorded at the same frame rate, and it's likely 30 fps... default video recording frame rate on an iPhone anyway... Maybe this isn't important for building a model though."
  }
}