{
  "id": 411060,
  "title": "CV Leaderboard",
  "url": "/competitions/asl-fingerspelling/discussion/411060",
  "author_name": "Mark Wijkhuizen",
  "post_date": "2023-05-17T14:26:05.667000",
  "votes": 27,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hello Fellow Kagglers,</p>\n<p>Since no successful submissions are made so far I wanted to share my CV performance.<br>\nThe model is a transformer ender/decoder and training is done using next token prediction with categorical cross entropy loss for 50 epochs.</p>\n<p>For validation 10% is left out and stratification happens based on participant id.</p>\n<p>Accuracy is based on predicting 30 characters, including the padding tokens, with a teacher forcing approach. This means the ground truth of the first N tokens is used to predict token N+1, instead of the predicted first N tokens. Including the padding tokens makes the accuracy artificially high, as ~40% of the tokens are padding tokens. With a teaching forcing approach the model quickly learns to predict a padding token when the previous token is a padding token.</p>\n<p>There are 7878 validation samples and inference took 25:36.</p>\n<table>\n<thead>\n<tr>\n<th>accuracy</th>\n<th>top 5 accuracy</th>\n<th>Levenshtein distance</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7920</td>\n<td>0.9301</td>\n<td>16.8379</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F26640df5234559febf751a20bbe3cd9a%2Fld_train_hist.png?generation=1684587094810240&amp;alt=media\" alt=\"\"></p>\n<p>As discussed in the comments, it is already challenging to fit the training data. The following table shows the <strong>training</strong> predictions without teacher forcing, meaning predicting token N is based on the predicted N-1 tokens, and not the true N-1 tokens.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb614852dccfa4803e6e9f3aede0dcb6e%2Fld_examples_train.png?generation=1684587073605828&amp;alt=media\" alt=\"\"></p>\n<p>Here are the validation Levenshtein distance histogram and some example predictions on the <strong>validation</strong> set.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F908a99172795e7213e12b5b39234f394%2Fld_val_hist.png?generation=1684587926310655&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F059d9402e3be2983e0abe85a28170681%2Fld_examples_val.png?generation=1684587937407662&amp;alt=media\" alt=\"\"></p>\n<p>Highly interested in your approach and CV</p>",
  "messages": [
    {
      "id": 2263359,
      "postDate": "2023-05-17T14:26:05.667Z",
      "content": "<p>Hello Fellow Kagglers,</p>\n<p>Since no successful submissions are made so far I wanted to share my CV performance.<br>\nThe model is a transformer ender/decoder and training is done using next token prediction with categorical cross entropy loss for 50 epochs.</p>\n<p>For validation 10% is left out and stratification happens based on participant id.</p>\n<p>Accuracy is based on predicting 30 characters, including the padding tokens, with a teacher forcing approach. This means the ground truth of the first N tokens is used to predict token N+1, instead of the predicted first N tokens. Including the padding tokens makes the accuracy artificially high, as ~40% of the tokens are padding tokens. With a teaching forcing approach the model quickly learns to predict a padding token when the previous token is a padding token.</p>\n<p>There are 7878 validation samples and inference took 25:36.</p>\n<table>\n<thead>\n<tr>\n<th>accuracy</th>\n<th>top 5 accuracy</th>\n<th>Levenshtein distance</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.7920</td>\n<td>0.9301</td>\n<td>16.8379</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F26640df5234559febf751a20bbe3cd9a%2Fld_train_hist.png?generation=1684587094810240&amp;alt=media\" alt=\"\"></p>\n<p>As discussed in the comments, it is already challenging to fit the training data. The following table shows the <strong>training</strong> predictions without teacher forcing, meaning predicting token N is based on the predicted N-1 tokens, and not the true N-1 tokens.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb614852dccfa4803e6e9f3aede0dcb6e%2Fld_examples_train.png?generation=1684587073605828&amp;alt=media\" alt=\"\"></p>\n<p>Here are the validation Levenshtein distance histogram and some example predictions on the <strong>validation</strong> set.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F908a99172795e7213e12b5b39234f394%2Fld_val_hist.png?generation=1684587926310655&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F059d9402e3be2983e0abe85a28170681%2Fld_examples_val.png?generation=1684587937407662&amp;alt=media\" alt=\"\"></p>\n<p>Highly interested in your approach and CV</p>",
      "rawMarkdown": "Hello Fellow Kagglers,\n\nSince no successful submissions are made so far I wanted to share my CV performance.\nThe model is a transformer ender/decoder and training is done using next token prediction with categorical cross entropy loss for 50 epochs.\n\nFor validation 10% is left out and stratification happens based on participant id.\n\nAccuracy is based on predicting 30 characters, including the padding tokens, with a teacher forcing approach. This means the ground truth of the first N tokens is used to predict token N+1, instead of the predicted first N tokens. Including the padding tokens makes the accuracy artificially high, as ~40% of the tokens are padding tokens. With a teaching forcing approach the model quickly learns to predict a padding token when the previous token is a padding token.\n\nThere are 7878 validation samples and inference took 25:36.\n\n| accuracy | top 5 accuracy | Levenshtein distance |\n| --- | --- | --- |\n| 0.7920 | 0.9301 | 16.8379 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F26640df5234559febf751a20bbe3cd9a%2Fld_train_hist.png?generation=1684587094810240&alt=media)\n\nAs discussed in the comments, it is already challenging to fit the training data. The following table shows the **training** predictions without teacher forcing, meaning predicting token N is based on the predicted N-1 tokens, and not the true N-1 tokens.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb614852dccfa4803e6e9f3aede0dcb6e%2Fld_examples_train.png?generation=1684587073605828&alt=media)\n\nHere are the validation Levenshtein distance histogram and some example predictions on the **validation** set.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F908a99172795e7213e12b5b39234f394%2Fld_val_hist.png?generation=1684587926310655&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F059d9402e3be2983e0abe85a28170681%2Fld_examples_val.png?generation=1684587937407662&alt=media)\n\nHighly interested in your approach and CV",
      "votes": 26
    },
    {
      "id": 2266313,
      "postDate": "2023-05-19T22:30:11.077Z",
      "content": "<p>Why hasn't anyone made a submission to LB yet? I have been watching this comp (I have not entered) and I find this strange.</p>",
      "rawMarkdown": "Why hasn't anyone made a submission to LB yet? I have been watching this comp (I have not entered) and I find this strange.",
      "votes": 6,
      "replies": [
        {
          "id": 2266319,
          "postDate": "2023-05-19T22:40:04.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2266183\" target=\"_blank\">This</a> was posted few hours ago. I am hoping we'll have some submissions. Although I did ASL, I don't find this one straight fwd yet. </p>",
          "rawMarkdown": "[This](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2266183) was posted few hours ago. I am hoping we'll have some submissions. Although I did ASL, I don't find this one straight fwd yet. ",
          "votes": 1
        },
        {
          "id": 2266584,
          "postDate": "2023-05-20T07:42:39.413Z",
          "content": "<p>There appear to be some issues with the submission process. Several people (including me) have already written models that work, but none of them pass the submission.</p>\n<p>Because you just get a generic error message when the submission fails, it's not clear what the issue is. It's probably something to do with the expected input or output format.</p>",
          "rawMarkdown": "There appear to be some issues with the submission process. Several people (including me) have already written models that work, but none of them pass the submission.\n\nBecause you just get a generic error message when the submission fails, it's not clear what the issue is. It's probably something to do with the expected input or output format.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2265095,
      "postDate": "2023-05-19T01:22:34.300Z",
      "content": "<p>Perhaps the edit distance is normalized and calculated by single parquet file every time, finally LB is the average of all test files.</p>",
      "rawMarkdown": "Perhaps the edit distance is normalized and calculated by single parquet file every time, finally LB is the average of all test files.",
      "votes": 1,
      "replies": [
        {
          "id": 2265527,
          "postDate": "2023-05-19T10:06:45.470Z",
          "content": "<p>Well, since we don't have any LB info thus far it's hard to check anything …</p>\n<p>This is yet another thing I would like the organisers to clarify: I have found several variations of Levenshtein Distance so it would be useful to know exactly which function is used, e.g. to confirm that is is indeed the tf.edit_distance or a different implementation …</p>",
          "rawMarkdown": "Well, since we don't have any LB info thus far it's hard to check anything ...\n\nThis is yet another thing I would like the organisers to clarify: I have found several variations of Levenshtein Distance so it would be useful to know exactly which function is used, e.g. to confirm that is is indeed the tf.edit_distance or a different implementation ...",
          "votes": 1,
          "replies": [
            {
              "id": 2266444,
              "postDate": "2023-05-20T04:40:12.857Z",
              "content": "<p>Correct me if I am wrong. From the evaluation page Levenshtein Distance = (N - D) / N where N is number of characters of true label and D is number of edits. For best case D=0, Levenshtein Distance equals 1. Worst case comes when predicted label length are far more longer than true label, it not only needs to replace wrong characters but also needs to remove extra characters, the result could be a negative value.</p>",
              "rawMarkdown": "Correct me if I am wrong. From the evaluation page Levenshtein Distance = (N - D) / N where N is number of characters of true label and D is number of edits. For best case D=0, Levenshtein Distance equals 1. Worst case comes when predicted label length are far more longer than true label, it not only needs to replace wrong characters but also needs to remove extra characters, the result could be a negative value."
            },
            {
              "id": 2266581,
              "postDate": "2023-05-20T07:41:13.817Z",
              "content": "<p>I interpreted it differently <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> </p>\n<p>First we need to compute the Levenshtein distance D. Then we take the number of characters of the ground truth N, and then we can normalize the distance and invert it to become more of a similarity measure. Then it is indeed smaller than or equal to 1 and can be negative.</p>",
              "rawMarkdown": "I interpreted it differently @lonnieqin \n\nFirst we need to compute the Levenshtein distance D. Then we take the number of characters of the ground truth N, and then we can normalize the distance and invert it to become more of a similarity measure. Then it is indeed smaller than or equal to 1 and can be negative.",
              "votes": 1
            },
            {
              "id": 2266593,
              "postDate": "2023-05-20T07:52:27.773Z",
              "content": "<p>Yes, the metric range from -n ~ 1.</p>",
              "rawMarkdown": "Yes, the metric range from -n ~ 1."
            }
          ]
        }
      ]
    },
    {
      "id": 2377110,
      "postDate": "2023-08-07T01:05:31.403Z",
      "content": "<p>Hi! Thank you for your post!<br>\nI've not implemented cross validation, but I'm noticing a large gap between my validation score and leaderboard scores.<br>\nAbout 0.15 difference. I'm not sure why this is happening at all.</p>\n<p>I was wondering if anyone has had a similar experience?</p>",
      "rawMarkdown": "Hi! Thank you for your post!\nI've not implemented cross validation, but I'm noticing a large gap between my validation score and leaderboard scores.\nAbout 0.15 difference. I'm not sure why this is happening at all.\n\nI was wondering if anyone has had a similar experience?"
    },
    {
      "id": 2358257,
      "postDate": "2023-07-25T11:53:03.013Z",
      "content": "<p>Thnx for sharing, mate</p>",
      "rawMarkdown": "Thnx for sharing, mate"
    },
    {
      "id": 2266878,
      "postDate": "2023-05-20T12:44:27.670Z",
      "content": "<p>Regarding the challenge in question, in general terms, what would be the objective? We should break the video into parts, character by character, and then try to find out which character matches a character in the json file. Would it be this? Or am I getting it all wrong.</p>",
      "rawMarkdown": "Regarding the challenge in question, in general terms, what would be the objective? We should break the video into parts, character by character, and then try to find out which character matches a character in the json file. Would it be this? Or am I getting it all wrong."
    },
    {
      "id": 2266720,
      "postDate": "2023-05-20T09:57:13.207Z",
      "content": "<p>With regard to your approach and the competition expectation of what will be in test - had the impression the phrase entries were randomly generated (not considering supplemental data here).  For the train data there is very little repetition of phrases even for the same participant id.   So was thinking test may not have any of the same exact phrases as in train but not sure of that.</p>\n<p>Learning that a character + or  - is followed by a number as a pattern could be useful - not sure how easy to do regex in TF! But if patterns can be detected then perhaps infill where uncertain might help the total levenshtein distance. So if a pattern like a phone number is recognised as a +nnn-nnn-nnn-nn getting the - in the right places even if not predicted with certainty would be better than omitting them.  Would be good to have some examples of how the metric will score here.</p>\n<p>Initially am considering the characters which are static, require no movement of fingers to determine and a shifting window of a certain size to predict the most likely in that small window of frames.  Since the train data has a lot of NaNs have not yet determined if these are good indicators of breaks in characters or not, but eliminating them to reduce data size. Early days, this is hard.</p>",
      "rawMarkdown": "With regard to your approach and the competition expectation of what will be in test - had the impression the phrase entries were randomly generated (not considering supplemental data here).  For the train data there is very little repetition of phrases even for the same participant id.   So was thinking test may not have any of the same exact phrases as in train but not sure of that.\n \nLearning that a character + or  - is followed by a number as a pattern could be useful - not sure how easy to do regex in TF! But if patterns can be detected then perhaps infill where uncertain might help the total levenshtein distance. So if a pattern like a phone number is recognised as a +nnn-nnn-nnn-nn getting the - in the right places even if not predicted with certainty would be better than omitting them.  Would be good to have some examples of how the metric will score here.\n\nInitially am considering the characters which are static, require no movement of fingers to determine and a shifting window of a certain size to predict the most likely in that small window of frames.  Since the train data has a lot of NaNs have not yet determined if these are good indicators of breaks in characters or not, but eliminating them to reduce data size. Early days, this is hard.\n",
      "replies": [
        {
          "id": 2266734,
          "postDate": "2023-05-20T10:13:06.943Z",
          "content": "<p>Thanks for sharing your thoughts.</p>\n<p>The training data description mentions:</p>\n<pre><code>The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls.\n</code></pre>\n<p>There are thus only 3 \"phrase type classes\" which should all be relatively easy to automatically label. These phrase type labels could be used as an additional training objective and might help improve the performance.</p>",
          "rawMarkdown": "Thanks for sharing your thoughts.\n\nThe training data description mentions:\n\n```\nThe train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls.\n```\n\nThere are thus only 3 \"phrase type classes\" which should all be relatively easy to automatically label. These phrase type labels could be used as an additional training objective and might help improve the performance.",
          "votes": 3,
          "replies": [
            {
              "id": 2266747,
              "postDate": "2023-05-20T10:39:44.430Z",
              "content": "<p>That's a good idea.  There are some randoms that are just numbers or letters, short phrases so maybe a 4th component class or just leave out of training if too few?  If you find it is useful, let us know and likewise will report back any findings.<br>\nThanks!</p>",
              "rawMarkdown": "That's a good idea.  There are some randoms that are just numbers or letters, short phrases so maybe a 4th component class or just leave out of training if too few?  If you find it is useful, let us know and likewise will report back any findings.\nThanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2264879,
      "postDate": "2023-05-18T20:03:52.270Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>! Could you please share (text or code) of how you are computing Accuracy (text vs text, tokens before eos vs shifted tokens)? </p>",
      "rawMarkdown": "Hello, @markwijkhuizen! Could you please share (text or code) of how you are computing Accuracy (text vs text, tokens before eos vs shifted tokens)? ",
      "replies": [
        {
          "id": 2264896,
          "postDate": "2023-05-18T20:20:39.720Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a>, added a description in the original post. Please let me know if any details remain unclear.</p>",
          "rawMarkdown": "Hello @vad13irt, added a description in the original post. Please let me know if any details remain unclear.",
          "votes": 1,
          "replies": [
            {
              "id": 2264913,
              "postDate": "2023-05-18T20:46:27.993Z",
              "content": "<p>Thank you for detailed description. Maybe be better just ignoring padding tokens with setting <code>ignore_index=pad_token_id</code>. My accuracy is ~35% (tokens before eos vs shifted tokens) and levenshetein distance is ~17 with only LSTM layers (note, not bidirectional). Although, when I am trying to see text predictions they are very terrible. Have you notice  something similar? And yes Leveshtein ~16-17 is awful for such short texts (length ~30-40). What is your opinion about it? </p>",
              "rawMarkdown": "Thank you for detailed description. Maybe be better just ignoring padding tokens with setting `ignore_index=pad_token_id`. My accuracy is ~35% (tokens before eos vs shifted tokens) and levenshetein distance is ~17 with only LSTM layers (note, not bidirectional). Although, when I am trying to see text predictions they are very terrible. Have you notice  something similar? And yes Leveshtein ~16-17 is awful for such short texts (length ~30-40). What is your opinion about it? "
            },
            {
              "id": 2265375,
              "postDate": "2023-05-19T08:05:07.543Z",
              "content": "<p>Ignoring the padding token does not seem to be an option in the native Tensorflow Accuracy metric, but I will write a custom metric to get comparable and fair results.</p>\n<p>The mean phrase length is ~17.8, a Levenshtein distance of ~17 is therefore roughly random guessing.<br>\nEven on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.</p>\n<p>In the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.</p>\n<p>Even with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.</p>\n<p>In addition to this, there are zero successful submissions.</p>\n<p>I am truly curious what we are missing.</p>\n<p>When this competition started I thought the community would quickly come up with interesting solutions due to the similarity with the previous Isolated Sign Language competition, but I could not have been more wrong.</p>",
              "rawMarkdown": "Ignoring the padding token does not seem to be an option in the native Tensorflow Accuracy metric, but I will write a custom metric to get comparable and fair results.\n\nThe mean phrase length is ~17.8, a Levenshtein distance of ~17 is therefore roughly random guessing.\nEven on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.\n\nIn the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.\n\nEven with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.\n\nIn addition to this, there are zero successful submissions.\n\nI am truly curious what we are missing.\n\nWhen this competition started I thought the community would quickly come up with interesting solutions due to the similarity with the previous Isolated Sign Language competition, but I could not have been more wrong.",
              "votes": 1
            },
            {
              "id": 2265512,
              "postDate": "2023-05-19T09:50:00.210Z",
              "content": "<pre><code>Even on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.\n\nIn the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.\n\nEven with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.\n</code></pre>\n<p>The same.<br>\nMountain, …, springboard. </p>\n<p><a href=\"https://ibb.co/Vv15nHr\"><img src=\"https://i.ibb.co/K9PS4NR/Screenshot-6.png\" alt=\"Screenshot-6\"></a></p>",
              "rawMarkdown": "```\nEven on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.\n\nIn the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.\n\nEven with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.\n```\n\nThe same.\nMountain, ..., springboard. \n\n<a href=\"https://ibb.co/Vv15nHr\"><img src=\"https://i.ibb.co/K9PS4NR/Screenshot-6.png\" alt=\"Screenshot-6\" border=\"0\"></a>\n\n"
            },
            {
              "id": 2265568,
              "postDate": "2023-05-19T10:47:35.483Z",
              "content": "<p>Thanks for sharing your training history.<br>\nIs the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?</p>\n<p>P.S. some training predictions are added in the original post.</p>",
              "rawMarkdown": "Thanks for sharing your training history.\nIs the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?\n\nP.S. some training predictions are added in the original post.",
              "votes": 1
            },
            {
              "id": 2265644,
              "postDate": "2023-05-19T12:17:54.977Z",
              "content": "<blockquote>\n  <p>P.S. some training predictions are added in the original post.</p>\n</blockquote>\n<p>Thank you for sharing. Your predictions are better than mine… </p>\n<p><a href=\"https://ibb.co/85XTmtd\"><img src=\"https://i.ibb.co/bB7q6D2/Screenshot-7.png\" alt=\"Screenshot-7\"></a></p>\n<blockquote>\n  <p>Is the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?</p>\n</blockquote>\n<p>Here is my training step and forward codes. I am using PyTorch Lightning and TorchMetrics</p>\n<pre><code> ():\n         encoder_hidden_state  :\n            inputs = self.(inputs)\n            encoder_outputs, encoder_hidden_state = self.encoder(inputs)\n\n        embedding_outputs = self.embedding(sequences)\n        decoder_outputs, decoder_hidden_state = self.decoder(embedding_outputs, encoder_hidden_state)\n        decoder_outputs = self.head(decoder_outputs)\n\n         return_encoder_hidden_state:\n             decoder_outputs, encoder_hidden_state\n\n         decoder_outputs\n\n     ():\n        inputs = batch[]\n        labels = batch[]\n        texts = batch[]\n\n        outputs = self(inputs=inputs, sequences=labels[:, :-])\n\n        \n        loss_input = torch.reshape(outputs, shape=(-, self.tokenizer.vocab_size))\n        loss_target = torch.reshape(labels[:, :], shape=(-, ))\n        loss = F.cross_entropy(\n            =loss_input, \n            target=loss_target, \n            reduction=, \n            weight=, \n            ignore_index=self.tokenizer.pad_token_id,\n            label_smoothing=,\n        )\n\n        \n        predictions = torch.argmax(outputs, dim=-).detach().cpu()\n        predictions = torch.reshape(predictions, shape=(-, ))\n        labels = labels[:, :].detach().cpu()\n        labels = torch.reshape(labels, shape=(-, ))\n\n        accuracy = metrics.classification.accuracy(\n            preds=predictions, \n            target=labels, \n            num_classes=self.tokenizer.vocab_size,\n            ignore_index=self.tokenizer.pad_token_id,\n            task=,\n        )\n\n        logs = {\n            : loss,\n            : accuracy,\n        }\n\n        self.log_dict(logs, on_step=, on_epoch=, prog_bar=)\n\n         loss\n\n ():        \n        predicted_texts = []\n           inputs:\n             = .unsqueeze(dim=)\n\n            sequence = [self.tokenizer.bos_token_id]\n\n            encoder_hidden_state = \n             i  (self.max_length):\n                input_sequence = torch.tensor(sequence, dtype=torch.long).unsqueeze(dim=).to(self.device)\n                outputs, encoder_hidden_state = self(\n                    sequences=input_sequence, \n                    inputs=, \n                    encoder_hidden_state=encoder_hidden_state,\n                    return_encoder_hidden_state=,\n                )\n\n                outputs = outputs.detach().cpu().squeeze(dim=)\n                predictions = torch.argmax(outputs, dim=-)\n                token = predictions[-].item()\n                sequence.append(token)\n\n                 token == self.tokenizer.eos_token_id:\n                    \n\n            text = self.tokenizer.decode(sequence, ignore_special_tokens=)\n            predicted_texts.append(text)\n\n        predicted_texts = np.array(predicted_texts)\n\n         predicted_texts\n</code></pre>\n<p>I am a beginner in training seq2seq models, so I am struggling now with predictions… If you see a mistake, please let me now. </p>",
              "rawMarkdown": "> P.S. some training predictions are added in the original post.\n\nThank you for sharing. Your predictions are better than mine... \n\n<a href=\"https://ibb.co/85XTmtd\"><img src=\"https://i.ibb.co/bB7q6D2/Screenshot-7.png\" alt=\"Screenshot-7\" border=\"0\"></a>\n\n> Is the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?\n\nHere is my training step and forward codes. I am using PyTorch Lightning and TorchMetrics\n\n```py\ndef forward(self, sequences, inputs=None, encoder_hidden_state=None, return_encoder_hidden_state=False):\n        if encoder_hidden_state is None:\n            inputs = self.input(inputs)\n            encoder_outputs, encoder_hidden_state = self.encoder(inputs)\n        \n        embedding_outputs = self.embedding(sequences)\n        decoder_outputs, decoder_hidden_state = self.decoder(embedding_outputs, encoder_hidden_state)\n        decoder_outputs = self.head(decoder_outputs)\n        \n        if return_encoder_hidden_state:\n            return decoder_outputs, encoder_hidden_state\n        \n        return decoder_outputs\n    \n    def training_step(self, batch, batch_index):\n        inputs = batch[\"inputs\"]\n        labels = batch[\"labels\"]\n        texts = batch[\"texts\"]\n        \n        outputs = self(inputs=inputs, sequences=labels[:, :-1])\n        \n        # loss\n        loss_input = torch.reshape(outputs, shape=(-1, self.tokenizer.vocab_size))\n        loss_target = torch.reshape(labels[:, 1:], shape=(-1, ))\n        loss = F.cross_entropy(\n            input=loss_input, \n            target=loss_target, \n            reduction=\"mean\", \n            weight=None, \n            ignore_index=self.tokenizer.pad_token_id,\n            label_smoothing=0.0,\n        )\n        \n        # metrics\n        predictions = torch.argmax(outputs, dim=-1).detach().cpu()\n        predictions = torch.reshape(predictions, shape=(-1, ))\n        labels = labels[:, 1:].detach().cpu()\n        labels = torch.reshape(labels, shape=(-1, ))\n        \n        accuracy = metrics.classification.accuracy(\n            preds=predictions, \n            target=labels, \n            num_classes=self.tokenizer.vocab_size,\n            ignore_index=self.tokenizer.pad_token_id,\n            task=\"multiclass\",\n        )\n        \n        logs = {\n            \"train/loss\": loss,\n            \"train/accuracy\": accuracy,\n        }\n        \n        self.log_dict(logs, on_step=True, on_epoch=True, prog_bar=True)\n        \n        return loss\n\ndef generate(self, inputs):        \n        predicted_texts = []\n        for input in inputs:\n            input = input.unsqueeze(dim=0)\n            \n            sequence = [self.tokenizer.bos_token_id]\n            \n            encoder_hidden_state = None\n            for i in range(self.max_length):\n                input_sequence = torch.tensor(sequence, dtype=torch.long).unsqueeze(dim=0).to(self.device)\n                outputs, encoder_hidden_state = self(\n                    sequences=input_sequence, \n                    inputs=input, \n                    encoder_hidden_state=encoder_hidden_state,\n                    return_encoder_hidden_state=True,\n                )\n                \n                outputs = outputs.detach().cpu().squeeze(dim=0)\n                predictions = torch.argmax(outputs, dim=-1)\n                token = predictions[-1].item()\n                sequence.append(token)\n                \n                if token == self.tokenizer.eos_token_id:\n                    break\n            \n            text = self.tokenizer.decode(sequence, ignore_special_tokens=True)\n            predicted_texts.append(text)\n            \n        predicted_texts = np.array(predicted_texts)\n            \n        return predicted_texts\n```\n\nI am a beginner in training seq2seq models, so I am struggling now with predictions... If you see a mistake, please let me now. "
            },
            {
              "id": 2266890,
              "postDate": "2023-05-20T12:53:58.370Z",
              "content": "<p>Your inference pipeline seems correct to me.</p>\n<p>See the updated post for my new training predictions. With some modifications to my dataset, model and training method the model seems to fit the training dataset.</p>",
              "rawMarkdown": "Your inference pipeline seems correct to me.\n\nSee the updated post for my new training predictions. With some modifications to my dataset, model and training method the model seems to fit the training dataset.",
              "votes": 1
            },
            {
              "id": 2267212,
              "postDate": "2023-05-20T17:16:25.433Z",
              "content": "<p>Now my predictions look slightly better with Transformer Encoder-Decoder models rather than LSTMs. Maybe I will figure out this problem, but now I am going to understand how correctly apply Teacher Forcing with Transformers or I will rewrite again my code to train token by token… </p>\n<p>This competition is going hard…</p>",
              "rawMarkdown": "Now my predictions look slightly better with Transformer Encoder-Decoder models rather than LSTMs. Maybe I will figure out this problem, but now I am going to understand how correctly apply Teacher Forcing with Transformers or I will rewrite again my code to train token by token... \n\nThis competition is going hard..."
            }
          ]
        }
      ]
    },
    {
      "id": 2263607,
      "postDate": "2023-05-17T18:21:56.237Z",
      "content": "<p>I will publish my own results with a baseline, but I am stuck, my model doesn't want to overfit a few samples like a couple of days ago. It just learns some tokens (i.e classes/chars) and outputs almost constant predictions :( </p>\n<p>My baseline is very simple, just \"old school\" LSTMs and CV split as yours. </p>",
      "rawMarkdown": "I will publish my own results with a baseline, but I am stuck, my model doesn't want to overfit a few samples like a couple of days ago. It just learns some tokens (i.e classes/chars) and outputs almost constant predictions :( \n\nMy baseline is very simple, just \"old school\" LSTMs and CV split as yours. "
    },
    {
      "id": 2263387,
      "postDate": "2023-05-17T14:45:43.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2266313,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-05-19T22:30:11.077000",
      "content": "<p>Why hasn't anyone made a submission to LB yet? I have been watching this comp (I have not entered) and I find this strange.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2266319,
          "author_name": "RB",
          "author_url": "",
          "post_date": "2023-05-19T22:40:04.043000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/410459#2266183\" target=\"_blank\">This</a> was posted few hours ago. I am hoping we'll have some submissions. Although I did ASL, I don't find this one straight fwd yet. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2266584,
          "author_name": "Mathieu De Coster",
          "author_url": "",
          "post_date": "2023-05-20T07:42:39.413000",
          "content": "<p>There appear to be some issues with the submission process. Several people (including me) have already written models that work, but none of them pass the submission.</p>\n<p>Because you just get a generic error message when the submission fails, it's not clear what the issue is. It's probably something to do with the expected input or output format.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2265095,
      "author_name": "Lonnie",
      "author_url": "",
      "post_date": "2023-05-19T01:22:34.300000",
      "content": "<p>Perhaps the edit distance is normalized and calculated by single parquet file every time, finally LB is the average of all test files.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2265527,
          "author_name": "Wondering Alice",
          "author_url": "",
          "post_date": "2023-05-19T10:06:45.470000",
          "content": "<p>Well, since we don't have any LB info thus far it's hard to check anything …</p>\n<p>This is yet another thing I would like the organisers to clarify: I have found several variations of Levenshtein Distance so it would be useful to know exactly which function is used, e.g. to confirm that is is indeed the tf.edit_distance or a different implementation …</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2266444,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-20T04:40:12.857000",
              "content": "<p>Correct me if I am wrong. From the evaluation page Levenshtein Distance = (N - D) / N where N is number of characters of true label and D is number of edits. For best case D=0, Levenshtein Distance equals 1. Worst case comes when predicted label length are far more longer than true label, it not only needs to replace wrong characters but also needs to remove extra characters, the result could be a negative value.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2266581,
              "author_name": "Mathieu De Coster",
              "author_url": "",
              "post_date": "2023-05-20T07:41:13.817000",
              "content": "<p>I interpreted it differently <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> </p>\n<p>First we need to compute the Levenshtein distance D. Then we take the number of characters of the ground truth N, and then we can normalize the distance and invert it to become more of a similarity measure. Then it is indeed smaller than or equal to 1 and can be negative.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2266593,
              "author_name": "Lonnie",
              "author_url": "",
              "post_date": "2023-05-20T07:52:27.773000",
              "content": "<p>Yes, the metric range from -n ~ 1.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2377110,
      "author_name": "Aaryam Sharma",
      "author_url": "",
      "post_date": "2023-08-07T01:05:31.403000",
      "content": "<p>Hi! Thank you for your post!<br>\nI've not implemented cross validation, but I'm noticing a large gap between my validation score and leaderboard scores.<br>\nAbout 0.15 difference. I'm not sure why this is happening at all.</p>\n<p>I was wondering if anyone has had a similar experience?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2358257,
      "author_name": "Yaroslav Petrov",
      "author_url": "",
      "post_date": "2023-07-25T11:53:03.013000",
      "content": "<p>Thnx for sharing, mate</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2266878,
      "author_name": "Regis Vargas",
      "author_url": "",
      "post_date": "2023-05-20T12:44:27.670000",
      "content": "<p>Regarding the challenge in question, in general terms, what would be the objective? We should break the video into parts, character by character, and then try to find out which character matches a character in the json file. Would it be this? Or am I getting it all wrong.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2266720,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2023-05-20T09:57:13.207000",
      "content": "<p>With regard to your approach and the competition expectation of what will be in test - had the impression the phrase entries were randomly generated (not considering supplemental data here).  For the train data there is very little repetition of phrases even for the same participant id.   So was thinking test may not have any of the same exact phrases as in train but not sure of that.</p>\n<p>Learning that a character + or  - is followed by a number as a pattern could be useful - not sure how easy to do regex in TF! But if patterns can be detected then perhaps infill where uncertain might help the total levenshtein distance. So if a pattern like a phone number is recognised as a +nnn-nnn-nnn-nn getting the - in the right places even if not predicted with certainty would be better than omitting them.  Would be good to have some examples of how the metric will score here.</p>\n<p>Initially am considering the characters which are static, require no movement of fingers to determine and a shifting window of a certain size to predict the most likely in that small window of frames.  Since the train data has a lot of NaNs have not yet determined if these are good indicators of breaks in characters or not, but eliminating them to reduce data size. Early days, this is hard.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2266734,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-05-20T10:13:06.943000",
          "content": "<p>Thanks for sharing your thoughts.</p>\n<p>The training data description mentions:</p>\n<pre><code>The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls.\n</code></pre>\n<p>There are thus only 3 \"phrase type classes\" which should all be relatively easy to automatically label. These phrase type labels could be used as an additional training objective and might help improve the performance.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2266747,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2023-05-20T10:39:44.430000",
              "content": "<p>That's a good idea.  There are some randoms that are just numbers or letters, short phrases so maybe a 4th component class or just leave out of training if too few?  If you find it is useful, let us know and likewise will report back any findings.<br>\nThanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2264879,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2023-05-18T20:03:52.270000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>! Could you please share (text or code) of how you are computing Accuracy (text vs text, tokens before eos vs shifted tokens)? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2264896,
          "author_name": "Mark Wijkhuizen",
          "author_url": "",
          "post_date": "2023-05-18T20:20:39.720000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a>, added a description in the original post. Please let me know if any details remain unclear.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2264913,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-18T20:46:27.993000",
              "content": "<p>Thank you for detailed description. Maybe be better just ignoring padding tokens with setting <code>ignore_index=pad_token_id</code>. My accuracy is ~35% (tokens before eos vs shifted tokens) and levenshetein distance is ~17 with only LSTM layers (note, not bidirectional). Although, when I am trying to see text predictions they are very terrible. Have you notice  something similar? And yes Leveshtein ~16-17 is awful for such short texts (length ~30-40). What is your opinion about it? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265375,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-05-19T08:05:07.543000",
              "content": "<p>Ignoring the padding token does not seem to be an option in the native Tensorflow Accuracy metric, but I will write a custom metric to get comparable and fair results.</p>\n<p>The mean phrase length is ~17.8, a Levenshtein distance of ~17 is therefore roughly random guessing.<br>\nEven on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.</p>\n<p>In the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.</p>\n<p>Even with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.</p>\n<p>In addition to this, there are zero successful submissions.</p>\n<p>I am truly curious what we are missing.</p>\n<p>When this competition started I thought the community would quickly come up with interesting solutions due to the similarity with the previous Isolated Sign Language competition, but I could not have been more wrong.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265512,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-19T09:50:00.210000",
              "content": "<pre><code>Even on the training set I achieve a Levenshtein distance of ~17 and the predicted phrases are nowhere close to the ground truth.\n\nIn the Isolated Sign Language competition you would near perfectly fit the training set, but in this competition that already seems challenging.\n\nEven with a large model with close to the maximum allowed 10M parameters, without any augmentations, I do not seem to be able to fit the training dataset.\n</code></pre>\n<p>The same.<br>\nMountain, …, springboard. </p>\n<p><a href=\"https://ibb.co/Vv15nHr\"><img src=\"https://i.ibb.co/K9PS4NR/Screenshot-6.png\" alt=\"Screenshot-6\"></a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2265568,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-05-19T10:47:35.483000",
              "content": "<p>Thanks for sharing your training history.<br>\nIs the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?</p>\n<p>P.S. some training predictions are added in the original post.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2265644,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-19T12:17:54.977000",
              "content": "<blockquote>\n  <p>P.S. some training predictions are added in the original post.</p>\n</blockquote>\n<p>Thank you for sharing. Your predictions are better than mine… </p>\n<p><a href=\"https://ibb.co/85XTmtd\"><img src=\"https://i.ibb.co/bB7q6D2/Screenshot-7.png\" alt=\"Screenshot-7\"></a></p>\n<blockquote>\n  <p>Is the Levenshtein distance without teacher forcing and the accuracy with teacher forcing?</p>\n</blockquote>\n<p>Here is my training step and forward codes. I am using PyTorch Lightning and TorchMetrics</p>\n<pre><code> ():\n         encoder_hidden_state  :\n            inputs = self.(inputs)\n            encoder_outputs, encoder_hidden_state = self.encoder(inputs)\n\n        embedding_outputs = self.embedding(sequences)\n        decoder_outputs, decoder_hidden_state = self.decoder(embedding_outputs, encoder_hidden_state)\n        decoder_outputs = self.head(decoder_outputs)\n\n         return_encoder_hidden_state:\n             decoder_outputs, encoder_hidden_state\n\n         decoder_outputs\n\n     ():\n        inputs = batch[]\n        labels = batch[]\n        texts = batch[]\n\n        outputs = self(inputs=inputs, sequences=labels[:, :-])\n\n        \n        loss_input = torch.reshape(outputs, shape=(-, self.tokenizer.vocab_size))\n        loss_target = torch.reshape(labels[:, :], shape=(-, ))\n        loss = F.cross_entropy(\n            =loss_input, \n            target=loss_target, \n            reduction=, \n            weight=, \n            ignore_index=self.tokenizer.pad_token_id,\n            label_smoothing=,\n        )\n\n        \n        predictions = torch.argmax(outputs, dim=-).detach().cpu()\n        predictions = torch.reshape(predictions, shape=(-, ))\n        labels = labels[:, :].detach().cpu()\n        labels = torch.reshape(labels, shape=(-, ))\n\n        accuracy = metrics.classification.accuracy(\n            preds=predictions, \n            target=labels, \n            num_classes=self.tokenizer.vocab_size,\n            ignore_index=self.tokenizer.pad_token_id,\n            task=,\n        )\n\n        logs = {\n            : loss,\n            : accuracy,\n        }\n\n        self.log_dict(logs, on_step=, on_epoch=, prog_bar=)\n\n         loss\n\n ():        \n        predicted_texts = []\n           inputs:\n             = .unsqueeze(dim=)\n\n            sequence = [self.tokenizer.bos_token_id]\n\n            encoder_hidden_state = \n             i  (self.max_length):\n                input_sequence = torch.tensor(sequence, dtype=torch.long).unsqueeze(dim=).to(self.device)\n                outputs, encoder_hidden_state = self(\n                    sequences=input_sequence, \n                    inputs=, \n                    encoder_hidden_state=encoder_hidden_state,\n                    return_encoder_hidden_state=,\n                )\n\n                outputs = outputs.detach().cpu().squeeze(dim=)\n                predictions = torch.argmax(outputs, dim=-)\n                token = predictions[-].item()\n                sequence.append(token)\n\n                 token == self.tokenizer.eos_token_id:\n                    \n\n            text = self.tokenizer.decode(sequence, ignore_special_tokens=)\n            predicted_texts.append(text)\n\n        predicted_texts = np.array(predicted_texts)\n\n         predicted_texts\n</code></pre>\n<p>I am a beginner in training seq2seq models, so I am struggling now with predictions… If you see a mistake, please let me now. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2266890,
              "author_name": "Mark Wijkhuizen",
              "author_url": "",
              "post_date": "2023-05-20T12:53:58.370000",
              "content": "<p>Your inference pipeline seems correct to me.</p>\n<p>See the updated post for my new training predictions. With some modifications to my dataset, model and training method the model seems to fit the training dataset.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2267212,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-20T17:16:25.433000",
              "content": "<p>Now my predictions look slightly better with Transformer Encoder-Decoder models rather than LSTMs. Maybe I will figure out this problem, but now I am going to understand how correctly apply Teacher Forcing with Transformers or I will rewrite again my code to train token by token… </p>\n<p>This competition is going hard…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2263607,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2023-05-17T18:21:56.237000",
      "content": "<p>I will publish my own results with a baseline, but I am stuck, my model doesn't want to overfit a few samples like a couple of days ago. It just learns some tokens (i.e classes/chars) and outputs almost constant predictions :( </p>\n<p>My baseline is very simple, just \"old school\" LSTMs and CV split as yours. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2263387,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-17T14:45:43.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2263359": "Hello Fellow Kagglers,\n\nSince no successful submissions are made so far I wanted to share my CV performance.\nThe model is a transformer ender/decoder and training is done using next token prediction with categorical cross entropy loss for 50 epochs.\n\nFor validation 10% is left out and stratification happens based on participant id.\n\nAccuracy is based on predicting 30 characters, including the padding tokens, with a teacher forcing approach. This means the ground truth of the first N tokens is used to predict token N+1, instead of the predicted first N tokens. Including the padding tokens makes the accuracy artificially high, as ~40% of the tokens are padding tokens. With a teaching forcing approach the model quickly learns to predict a padding token when the previous token is a padding token.\n\nThere are 7878 validation samples and inference took 25:36.\n\n| accuracy | top 5 accuracy | Levenshtein distance |\n| --- | --- | --- |\n| 0.7920 | 0.9301 | 16.8379 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F26640df5234559febf751a20bbe3cd9a%2Fld_train_hist.png?generation=1684587094810240&alt=media)\n\nAs discussed in the comments, it is already challenging to fit the training data. The following table shows the **training** predictions without teacher forcing, meaning predicting token N is based on the predicted N-1 tokens, and not the true N-1 tokens.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2Fb614852dccfa4803e6e9f3aede0dcb6e%2Fld_examples_train.png?generation=1684587073605828&alt=media)\n\nHere are the validation Levenshtein distance histogram and some example predictions on the **validation** set.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F908a99172795e7213e12b5b39234f394%2Fld_val_hist.png?generation=1684587926310655&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4433335%2F059d9402e3be2983e0abe85a28170681%2Fld_examples_val.png?generation=1684587937407662&alt=media)\n\nHighly interested in your approach and CV",
    "2266313": "Why hasn't anyone made a submission to LB yet? I have been watching this comp (I have not entered) and I find this strange.",
    "2265095": "Perhaps the edit distance is normalized and calculated by single parquet file every time, finally LB is the average of all test files.",
    "2377110": "Hi! Thank you for your post!\nI've not implemented cross validation, but I'm noticing a large gap between my validation score and leaderboard scores.\nAbout 0.15 difference. I'm not sure why this is happening at all.\n\nI was wondering if anyone has had a similar experience?",
    "2358257": "Thnx for sharing, mate",
    "2266878": "Regarding the challenge in question, in general terms, what would be the objective? We should break the video into parts, character by character, and then try to find out which character matches a character in the json file. Would it be this? Or am I getting it all wrong.",
    "2266720": "With regard to your approach and the competition expectation of what will be in test - had the impression the phrase entries were randomly generated (not considering supplemental data here).  For the train data there is very little repetition of phrases even for the same participant id.   So was thinking test may not have any of the same exact phrases as in train but not sure of that.\n \nLearning that a character + or  - is followed by a number as a pattern could be useful - not sure how easy to do regex in TF! But if patterns can be detected then perhaps infill where uncertain might help the total levenshtein distance. So if a pattern like a phone number is recognised as a +nnn-nnn-nnn-nn getting the - in the right places even if not predicted with certainty would be better than omitting them.  Would be good to have some examples of how the metric will score here.\n\nInitially am considering the characters which are static, require no movement of fingers to determine and a shifting window of a certain size to predict the most likely in that small window of frames.  Since the train data has a lot of NaNs have not yet determined if these are good indicators of breaks in characters or not, but eliminating them to reduce data size. Early days, this is hard.\n",
    "2264879": "Hello, @markwijkhuizen! Could you please share (text or code) of how you are computing Accuracy (text vs text, tokens before eos vs shifted tokens)? ",
    "2263607": "I will publish my own results with a baseline, but I am stuck, my model doesn't want to overfit a few samples like a couple of days ago. It just learns some tokens (i.e classes/chars) and outputs almost constant predictions :( \n\nMy baseline is very simple, just \"old school\" LSTMs and CV split as yours. ",
    "2263387": ""
  }
}