{
  "id": 434388,
  "title": "🥈27th Place Solution🥈",
  "url": "/competitions/asl-fingerspelling/discussion/434388",
  "author_name": "siwooyong",
  "post_date": "2023-08-25T03:53:36.294000",
  "votes": 18,
  "comment_count": 10,
  "views": 0,
  "content": "<h1>TLDR</h1>\n<p>This competition is very similar to speech recognition. Therefore, I attempted new ideas focusing on techniques used in speech recognition. Among them, the most effective ones were <strong>conformer architecture</strong>,  <strong>interCTC technique</strong>, <strong>strong augmentation</strong> and <strong>nomasking</strong>.<br>\n<br></p>\n<h1>Data Processing</h1>\n<p>Out of the provided 543 landmarks, I utilized 42 hand data, 33 pose data, and 40 lip data. Input features were constructed by concatenating xy coordinates (drop z) and motion (xy[1:] - xy[:-1]).</p>\n<ul>\n<li><p>preprocess</p>\n<pre><code>def pre:\n  x = x\n  x = tf.where(tf.math.is, tf.zeros, x)\n  n_frames = tf.shape(x)\n\n  lhand = tf.transpose(tf.reshape(x, (n_frames, , args.n_hand_landmarks)), (, , ))\n  rhand = tf.transpose(tf.reshape(x, (n_frames, , args.n_hand_landmarks)), (, , ))\n  pose = tf.transpose(tf.reshape(x, (n_frames, , args.n_pose_landmarks)), (, , ))\n  face = tf.transpose(tf.reshape(x, (n_frames, , args.n_face_landmarks)), (, , ))\n\n  x = tf.concat(, axis=)\n\n  x = x\n  return x\n\ndef decode:\n  schema = {\n      'coordinates': tf.io.,\n      'phrase': tf.io.\n  }\n  x = tf.io.parse\n\n  coordinates = tf.reshape(tf.sparse., (-, args.input_dim))\n  phrase = tf.sparse.\n\n   augment:\n    coordinates, phrase = augment\n\n  dx = tf.cond(tf.shape(coordinates)&gt;,lambda:tf.pad(coordinates - coordinates, ,]),lambda:tf.zeros)\n  coordinates = tf.concat(, axis=-)\n\n  return coordinates, phrase\n</code></pre></li>\n</ul>\n<p><br></p>\n<h1>Augmentation</h1>\n<p>Three types of augmentations were applied, each of which significantly influenced CV.</p>\n<ul>\n<li><p>flip hand(CV improvement : ~0.003)</p>\n<pre><code>\ndef flip_hand(video):\n  video = tf.reshape(video, shape=(-1, args.n_landmarks, 2))\n  hands = video[:, :int(2 * args.n_hand_landmarks)]\n  other = video[:, int(2 * args.n_hand_landmarks):]\n\n  lhand = hands[:, :args.n_hand_landmarks]\n  rhand = hands[:, args.n_hand_landmarks:]\n\n  lhand_x, rhand_x = lhand[:, :, 0], rhand[:, :, 0]\n  lhand_x = tf.negative(lhand_x) + 2 * tf.reduce_mean(lhand_x, =1, =)\n  rhand_x = tf.negative(rhand_x) + 2 * tf.reduce_mean(rhand_x, =1, =)\n\n  lhand = tf.concat([tf.expand_dims(lhand_x, =-1), lhand[:, :, 1:]], =-1)\n  rhand = tf.concat([tf.expand_dims(rhand_x, =-1), rhand[:, :, 1:]], =-1)\n\n  flipped_hands = tf.concat([rhand, lhand, other], =1)\n  flipped_hands = tf.reshape(flipped_hands, shape=(-1, args.input_dim))\n  return flipped_hands\n</code></pre></li>\n<li><p>flip video and phrase(CV improvement : ~0.005)</p>\n<pre><code>\n ():\n  x = x[]\n  y = y[]\n   x, y\n</code></pre></li>\n<li><p>concat 2 videos and 2 phrases(CV improvement : ~0.01)</p>\n<pre><code>#   videos   phrases\ndef cat_augment(inputs, inputs2):\n  x, y = inputs\n  x2, y2 = inputs2\n\n  x_shape = tf.shape(x)\n  x2_shape = tf.shape(x2)\n\n  should_concat = tf..uniform(()) &lt; \n\n  x_condition = should_concat &amp; (x_shape[] + x2_shape[] &lt; .max_frame)\n\n  x = tf.cond(x_condition, : tf.([x, x2], axis=), : x)\n  y = tf.cond(x_condition, : tf.([y, y2], axis=), : y)\n\n   x, y\n</code></pre></li>\n</ul>\n<p><br></p>\n<h1>Model</h1>\n<p>The model was constructed utilizing the architectures proposed in Conformer.  At first, I employed the Conformer code from <a href=\"https://github.com/TensorSpeech/TensorFlowASR\" target=\"_blank\">https://github.com/TensorSpeech/TensorFlowASR</a>. However, by making specific modifications to the code utilized by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> , I successfully implemented the Conformer architecture. This implementation led to an enhancement of 0.008 in the LB score. This improvement is likely due to the removal of residual blocks between Conformer blocks, allowing for the integration of more layers and resulting in improved robustness against overfitting.  Furthermore, the incorporation of the interCTC(<a href=\"https://arxiv.org/abs/2102.03216\" target=\"_blank\">paper</a>) contributed to an LB score(~0.005).</p>\n<p>Furthermore, as the model's output was computed without masking and a maximum length of 384 was employed for the CTC loss \"input-length' variable, using a larger kernel size(=31) effectively conveyed frame information across the 384-length range, resulting in improved performance.</p>\n<ul>\n<li><p>The overall structure of the model is as follows.</p>\n<pre><code>def get:\n  NUM_CLASSES=\n  PAD=\n  inp = tf.keras.)\n  x = inp\n  x = tf.keras.layers.(x)\n  x = tf.keras.layers.(x)\n\n  xs = \n   i  range(num_layers):\n      x = (x)\n      x = (x)\n      xs.append(x)\n\n  classifier = tf.keras.layers.\n\n  x1 = tf.keras.layers.(xs)\n  x1 = classifier(x1)\n\n  x2 = tf.keras.layers.(xs)\n  x2 = classifier(x2)\n  return x1, x2\n</code></pre>\n<p><br></p></li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.1 </li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 300(last 100 epoch for only train data(not suppl))</li>\n<li>Learning Rate : 1e-3 </li>\n<li>Loss Function : CTCLoss</li>\n</ul>\n<p><br></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>masking </li>\n<li>adding distance features</li>\n<li>mlm pretraining <br>\n<br></li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition\" target=\"_blank\">https://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition</a><br>\n<br></p>",
  "messages": [
    {
      "id": 2407365,
      "postDate": "2023-08-25T03:53:36.293Z",
      "content": "<h1>TLDR</h1>\n<p>This competition is very similar to speech recognition. Therefore, I attempted new ideas focusing on techniques used in speech recognition. Among them, the most effective ones were <strong>conformer architecture</strong>,  <strong>interCTC technique</strong>, <strong>strong augmentation</strong> and <strong>nomasking</strong>.<br>\n<br></p>\n<h1>Data Processing</h1>\n<p>Out of the provided 543 landmarks, I utilized 42 hand data, 33 pose data, and 40 lip data. Input features were constructed by concatenating xy coordinates (drop z) and motion (xy[1:] - xy[:-1]).</p>\n<ul>\n<li><p>preprocess</p>\n<pre><code>def pre:\n  x = x\n  x = tf.where(tf.math.is, tf.zeros, x)\n  n_frames = tf.shape(x)\n\n  lhand = tf.transpose(tf.reshape(x, (n_frames, , args.n_hand_landmarks)), (, , ))\n  rhand = tf.transpose(tf.reshape(x, (n_frames, , args.n_hand_landmarks)), (, , ))\n  pose = tf.transpose(tf.reshape(x, (n_frames, , args.n_pose_landmarks)), (, , ))\n  face = tf.transpose(tf.reshape(x, (n_frames, , args.n_face_landmarks)), (, , ))\n\n  x = tf.concat(, axis=)\n\n  x = x\n  return x\n\ndef decode:\n  schema = {\n      'coordinates': tf.io.,\n      'phrase': tf.io.\n  }\n  x = tf.io.parse\n\n  coordinates = tf.reshape(tf.sparse., (-, args.input_dim))\n  phrase = tf.sparse.\n\n   augment:\n    coordinates, phrase = augment\n\n  dx = tf.cond(tf.shape(coordinates)&gt;,lambda:tf.pad(coordinates - coordinates, ,]),lambda:tf.zeros)\n  coordinates = tf.concat(, axis=-)\n\n  return coordinates, phrase\n</code></pre></li>\n</ul>\n<p><br></p>\n<h1>Augmentation</h1>\n<p>Three types of augmentations were applied, each of which significantly influenced CV.</p>\n<ul>\n<li><p>flip hand(CV improvement : ~0.003)</p>\n<pre><code>\ndef flip_hand(video):\n  video = tf.reshape(video, shape=(-1, args.n_landmarks, 2))\n  hands = video[:, :int(2 * args.n_hand_landmarks)]\n  other = video[:, int(2 * args.n_hand_landmarks):]\n\n  lhand = hands[:, :args.n_hand_landmarks]\n  rhand = hands[:, args.n_hand_landmarks:]\n\n  lhand_x, rhand_x = lhand[:, :, 0], rhand[:, :, 0]\n  lhand_x = tf.negative(lhand_x) + 2 * tf.reduce_mean(lhand_x, =1, =)\n  rhand_x = tf.negative(rhand_x) + 2 * tf.reduce_mean(rhand_x, =1, =)\n\n  lhand = tf.concat([tf.expand_dims(lhand_x, =-1), lhand[:, :, 1:]], =-1)\n  rhand = tf.concat([tf.expand_dims(rhand_x, =-1), rhand[:, :, 1:]], =-1)\n\n  flipped_hands = tf.concat([rhand, lhand, other], =1)\n  flipped_hands = tf.reshape(flipped_hands, shape=(-1, args.input_dim))\n  return flipped_hands\n</code></pre></li>\n<li><p>flip video and phrase(CV improvement : ~0.005)</p>\n<pre><code>\n ():\n  x = x[]\n  y = y[]\n   x, y\n</code></pre></li>\n<li><p>concat 2 videos and 2 phrases(CV improvement : ~0.01)</p>\n<pre><code>#   videos   phrases\ndef cat_augment(inputs, inputs2):\n  x, y = inputs\n  x2, y2 = inputs2\n\n  x_shape = tf.shape(x)\n  x2_shape = tf.shape(x2)\n\n  should_concat = tf..uniform(()) &lt; \n\n  x_condition = should_concat &amp; (x_shape[] + x2_shape[] &lt; .max_frame)\n\n  x = tf.cond(x_condition, : tf.([x, x2], axis=), : x)\n  y = tf.cond(x_condition, : tf.([y, y2], axis=), : y)\n\n   x, y\n</code></pre></li>\n</ul>\n<p><br></p>\n<h1>Model</h1>\n<p>The model was constructed utilizing the architectures proposed in Conformer.  At first, I employed the Conformer code from <a href=\"https://github.com/TensorSpeech/TensorFlowASR\" target=\"_blank\">https://github.com/TensorSpeech/TensorFlowASR</a>. However, by making specific modifications to the code utilized by <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> , I successfully implemented the Conformer architecture. This implementation led to an enhancement of 0.008 in the LB score. This improvement is likely due to the removal of residual blocks between Conformer blocks, allowing for the integration of more layers and resulting in improved robustness against overfitting.  Furthermore, the incorporation of the interCTC(<a href=\"https://arxiv.org/abs/2102.03216\" target=\"_blank\">paper</a>) contributed to an LB score(~0.005).</p>\n<p>Furthermore, as the model's output was computed without masking and a maximum length of 384 was employed for the CTC loss \"input-length' variable, using a larger kernel size(=31) effectively conveyed frame information across the 384-length range, resulting in improved performance.</p>\n<ul>\n<li><p>The overall structure of the model is as follows.</p>\n<pre><code>def get:\n  NUM_CLASSES=\n  PAD=\n  inp = tf.keras.)\n  x = inp\n  x = tf.keras.layers.(x)\n  x = tf.keras.layers.(x)\n\n  xs = \n   i  range(num_layers):\n      x = (x)\n      x = (x)\n      xs.append(x)\n\n  classifier = tf.keras.layers.\n\n  x1 = tf.keras.layers.(xs)\n  x1 = classifier(x1)\n\n  x2 = tf.keras.layers.(xs)\n  x2 = classifier(x2)\n  return x1, x2\n</code></pre>\n<p><br></p></li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>Scheduler : lr_warmup_cosine_decay </li>\n<li>Warmup Ratio : 0.1 </li>\n<li>Optimizer : AdamW </li>\n<li>Weight Decay : 0.01</li>\n<li>Epoch : 300(last 100 epoch for only train data(not suppl))</li>\n<li>Learning Rate : 1e-3 </li>\n<li>Loss Function : CTCLoss</li>\n</ul>\n<p><br></p>\n<h1>Didn't Work</h1>\n<ul>\n<li>masking </li>\n<li>adding distance features</li>\n<li>mlm pretraining <br>\n<br></li>\n</ul>\n<h1>Code</h1>\n<p><a href=\"https://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition\" target=\"_blank\">https://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition</a><br>\n<br></p>",
      "rawMarkdown": "# TLDR\nThis competition is very similar to speech recognition. Therefore, I attempted new ideas focusing on techniques used in speech recognition. Among them, the most effective ones were **conformer architecture**,  **interCTC technique**, **strong augmentation** and **nomasking**.\n<br>\n\n# Data Processing\nOut of the provided 543 landmarks, I utilized 42 hand data, 33 pose data, and 40 lip data. Input features were constructed by concatenating xy coordinates (drop z) and motion (xy[1:] - xy[:-1]).\n\n * preprocess\n\n        def pre_process0(x):\n          x = x[:args.max_frame]\n          x = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n          n_frames = tf.shape(x)[0]\n\n          lhand = tf.transpose(tf.reshape(x[:, 0:63], (n_frames, 3, args.n_hand_landmarks)), (0, 2, 1))\n          rhand = tf.transpose(tf.reshape(x[:, 63:126], (n_frames, 3, args.n_hand_landmarks)), (0, 2, 1))\n          pose = tf.transpose(tf.reshape(x[:, 126:225], (n_frames, 3, args.n_pose_landmarks)), (0, 2, 1))\n          face = tf.transpose(tf.reshape(x[:, 225:345], (n_frames, 3, args.n_face_landmarks)), (0, 2, 1))\n\n          x = tf.concat([\n              lhand,\n              rhand,\n              pose,\n              face\n              ], axis=1)\n\n          x = x[:, :, :2]\n          return x\n\n        def decode_fn(record_bytes, augment=False):\n          schema = {\n              'coordinates': tf.io.VarLenFeature(tf.float32),\n              'phrase': tf.io.VarLenFeature(tf.int64)\n          }\n          x = tf.io.parse_single_example(record_bytes, schema)\n\n          coordinates = tf.reshape(tf.sparse.to_dense(x[\"coordinates\"]), (-1, args.input_dim))\n          phrase = tf.sparse.to_dense(x[\"phrase\"])\n\n          if augment:\n            coordinates, phrase = augment_fn(coordinates, phrase)\n\n          dx = tf.cond(tf.shape(coordinates)[0]>1,lambda:tf.pad(coordinates[1:] - coordinates[:-1], [[0,1],[0,0]]),lambda:tf.zeros_like(coordinates))\n          coordinates = tf.concat([coordinates, dx], axis=-1)\n\n          return coordinates, phrase\n\n<br>\n\n# Augmentation\nThree types of augmentations were applied, each of which significantly influenced CV.\n\n* flip hand(CV improvement : ~0.003)\n        \n        # flip hand\n        def flip_hand(video):\n          video = tf.reshape(video, shape=(-1, args.n_landmarks, 2))\n          hands = video[:, :int(2 * args.n_hand_landmarks)]\n          other = video[:, int(2 * args.n_hand_landmarks):]\n\n          lhand = hands[:, :args.n_hand_landmarks]\n          rhand = hands[:, args.n_hand_landmarks:]\n\n          lhand_x, rhand_x = lhand[:, :, 0], rhand[:, :, 0]\n          lhand_x = tf.negative(lhand_x) + 2 * tf.reduce_mean(lhand_x, axis=1, keepdims=True)\n          rhand_x = tf.negative(rhand_x) + 2 * tf.reduce_mean(rhand_x, axis=1, keepdims=True)\n\n          lhand = tf.concat([tf.expand_dims(lhand_x, axis=-1), lhand[:, :, 1:]], axis=-1)\n          rhand = tf.concat([tf.expand_dims(rhand_x, axis=-1), rhand[:, :, 1:]], axis=-1)\n\n          flipped_hands = tf.concat([rhand, lhand, other], axis=1)\n          flipped_hands = tf.reshape(flipped_hands, shape=(-1, args.input_dim))\n          return flipped_hands\n\n* flip video and phrase(CV improvement : ~0.005)\n          \n        # flip video and phrase\n        def reverse_frames(x, y):\n          x = x[::-1]\n          y = y[::-1]\n          return x, y\n\n* concat 2 videos and 2 phrases(CV improvement : ~0.01)\n        \n        # concat 2 videos and 2 phrases\n        def cat_augment(inputs, inputs2):\n          x, y = inputs\n          x2, y2 = inputs2\n\n          x_shape = tf.shape(x)\n          x2_shape = tf.shape(x2)\n\n          should_concat = tf.random.uniform(()) < 0.5\n\n          x_condition = should_concat & (x_shape[0] + x2_shape[0] < args.max_frame)\n\n          x = tf.cond(x_condition, lambda: tf.concat([x, x2], axis=0), lambda: x)\n          y = tf.cond(x_condition, lambda: tf.concat([y, y2], axis=0), lambda: y)\n\n          return x, y\n\n<br>\n\n# Model\nThe model was constructed utilizing the architectures proposed in Conformer.  At first, I employed the Conformer code from https://github.com/TensorSpeech/TensorFlowASR. However, by making specific modifications to the code utilized by @hoyso48 , I successfully implemented the Conformer architecture. This implementation led to an enhancement of 0.008 in the LB score. This improvement is likely due to the removal of residual blocks between Conformer blocks, allowing for the integration of more layers and resulting in improved robustness against overfitting.  Furthermore, the incorporation of the interCTC([paper](https://arxiv.org/abs/2102.03216)) contributed to an LB score(~0.005).\n\n\nFurthermore, as the model's output was computed without masking and a maximum length of 384 was employed for the CTC loss \"input-length' variable, using a larger kernel size(=31) effectively conveyed frame information across the 384-length range, resulting in improved performance.\n\n* The overall structure of the model is as follows.\n\n        def get_model(max_len=384, dim=160,ksize=31,drop_rate=0.1, num_layers=16):\n          NUM_CLASSES=63\n          PAD=0\n          inp = tf.keras.Input((None,2*230))\n          x = inp\n          x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n          x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n          xs = []\n          for i in range(num_layers):\n              x = TransformerBlock(dim,expand=2)(x)\n              x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n              xs.append(x)\n\n          classifier = tf.keras.layers.Dense(NUM_CLASSES,name='classifier')\n\n          x1 = tf.keras.layers.Dropout(0.2)(xs[-8])\n          x1 = classifier(x1)\n\n          x2 = tf.keras.layers.Dropout(0.2)(xs[-1])\n          x2 = classifier(x2)\n          return x1, x2\n<br>\n\n# Training\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.1 \n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 300(last 100 epoch for only train data(not suppl))\n* Learning Rate : 1e-3 \n* Loss Function : CTCLoss\n\n<br>\n\n# Didn't Work\n* masking \n* adding distance features\n* mlm pretraining \n<br>\n\n# Code\nhttps://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition\n<br>",
      "votes": 18
    },
    {
      "id": 2409082,
      "postDate": "2023-08-26T04:02:28.623Z",
      "content": "<p>Great Work Man!</p>",
      "rawMarkdown": "Great Work Man!",
      "votes": 1
    },
    {
      "id": 2407604,
      "postDate": "2023-08-25T06:42:41.787Z",
      "content": "<p>Congratulations. Thanks for sharing the details. I observe with 50 epochs, the notebook was consuming around 9 hrs. Can you advise how you could increase the number of epochs to 300 while keeping the notebook time &lt;9hrs.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details. I observe with 50 epochs, the notebook was consuming around 9 hrs. Can you advise how you could increase the number of epochs to 300 while keeping the notebook time <9hrs.",
      "votes": 1,
      "replies": [
        {
          "id": 2407831,
          "postDate": "2023-08-25T09:20:17.590Z",
          "content": "<p>I trained the model using TPUs on colab, and furthermore, I was able to achieve faster training speeds by using the CTC loss on TPU by <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>",
          "rawMarkdown": "I trained the model using TPUs on colab, and furthermore, I was able to achieve faster training speeds by using the CTC loss on TPU by @shlomoron "
        }
      ]
    },
    {
      "id": 2407440,
      "postDate": "2023-08-25T05:02:21.433Z",
      "content": "<p>Thanks for sharing, 'concat 2 videos and 2 phrases' how much does this help?</p>",
      "rawMarkdown": "Thanks for sharing, 'concat 2 videos and 2 phrases' how much does this help?",
      "votes": 1,
      "replies": [
        {
          "id": 2407804,
          "postDate": "2023-08-25T09:11:40.333Z",
          "content": "<p>Concatenating 2 random samples from the \"train\" and \"suppl\" data (with an augmentation probability of 1.0) resulted in an LB score improvement of 0.009~0.01.</p>",
          "rawMarkdown": "Concatenating 2 random samples from the \"train\" and \"suppl\" data (with an augmentation probability of 1.0) resulted in an LB score improvement of 0.009~0.01.",
          "votes": 1,
          "replies": [
            {
              "id": 2407875,
              "postDate": "2023-08-25T09:42:14.720Z",
              "content": "<p>Cool, thanks!</p>",
              "rawMarkdown": "Cool, thanks!",
              "votes": 1
            },
            {
              "id": 2408444,
              "postDate": "2023-08-25T16:23:20.347Z",
              "content": "<p>Thanks for sharing. I tried something similar without improvement. I noticed that you used separate batches, whereas I concatenated a batch with a random sorting of itself. I wonder if that's what made the difference or the fact that I trained for a small number of epochs (I already noticed that some of the things that didn't work for me worked for others, probably because I tested with a small model for only 50 epochs). Any thoughts?</p>",
              "rawMarkdown": "Thanks for sharing. I tried something similar without improvement. I noticed that you used separate batches, whereas I concatenated a batch with a random sorting of itself. I wonder if that's what made the difference or the fact that I trained for a small number of epochs (I already noticed that some of the things that didn't work for me worked for others, probably because I tested with a small model for only 50 epochs). Any thoughts?",
              "votes": 1
            },
            {
              "id": 2408945,
              "postDate": "2023-08-25T23:56:31.110Z",
              "content": "<p>Concat augmentation I implemented showed notable performance improvements when utilized between the training and supplementary data, as well as with an increased number of training epochs.</p>",
              "rawMarkdown": "Concat augmentation I implemented showed notable performance improvements when utilized between the training and supplementary data, as well as with an increased number of training epochs.",
              "votes": 1
            },
            {
              "id": 2408975,
              "postDate": "2023-08-26T01:17:11.483Z",
              "content": "<p>I found 1st solution use similar cut mix method.</p>",
              "rawMarkdown": "I found 1st solution use similar cut mix method."
            },
            {
              "id": 2409028,
              "postDate": "2023-08-26T02:55:37.800Z",
              "content": "<p>That's right! The first solution suggested \"cutting 2 sequences and their corresponding phrases by that percentage and mixing them.\" Given my belief that combining data after cutting wouldn't enhance performance, I'm really surprised to see an increase of 0.005 points!</p>",
              "rawMarkdown": "That's right! The first solution suggested \"cutting 2 sequences and their corresponding phrases by that percentage and mixing them.\" Given my belief that combining data after cutting wouldn't enhance performance, I'm really surprised to see an increase of 0.005 points!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2409082,
      "author_name": "Mystic Shadow",
      "author_url": "",
      "post_date": "2023-08-26T04:02:28.623000",
      "content": "<p>Great Work Man!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2407604,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T06:42:41.787000",
      "content": "<p>Congratulations. Thanks for sharing the details. I observe with 50 epochs, the notebook was consuming around 9 hrs. Can you advise how you could increase the number of epochs to 300 while keeping the notebook time &lt;9hrs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407831,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-08-25T09:20:17.590000",
          "content": "<p>I trained the model using TPUs on colab, and furthermore, I was able to achieve faster training speeds by using the CTC loss on TPU by <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2407440,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2023-08-25T05:02:21.433000",
      "content": "<p>Thanks for sharing, 'concat 2 videos and 2 phrases' how much does this help?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2407804,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-08-25T09:11:40.333000",
          "content": "<p>Concatenating 2 random samples from the \"train\" and \"suppl\" data (with an augmentation probability of 1.0) resulted in an LB score improvement of 0.009~0.01.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2407875,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-25T09:42:14.720000",
              "content": "<p>Cool, thanks!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2408444,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "2023-08-25T16:23:20.347000",
              "content": "<p>Thanks for sharing. I tried something similar without improvement. I noticed that you used separate batches, whereas I concatenated a batch with a random sorting of itself. I wonder if that's what made the difference or the fact that I trained for a small number of epochs (I already noticed that some of the things that didn't work for me worked for others, probably because I tested with a small model for only 50 epochs). Any thoughts?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2408945,
              "author_name": "siwooyong",
              "author_url": "",
              "post_date": "2023-08-25T23:56:31.110000",
              "content": "<p>Concat augmentation I implemented showed notable performance improvements when utilized between the training and supplementary data, as well as with an increased number of training epochs.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2408975,
              "author_name": "gezi",
              "author_url": "",
              "post_date": "2023-08-26T01:17:11.483000",
              "content": "<p>I found 1st solution use similar cut mix method.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2409028,
              "author_name": "siwooyong",
              "author_url": "",
              "post_date": "2023-08-26T02:55:37.800000",
              "content": "<p>That's right! The first solution suggested \"cutting 2 sequences and their corresponding phrases by that percentage and mixing them.\" Given my belief that combining data after cutting wouldn't enhance performance, I'm really surprised to see an increase of 0.005 points!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2407365": "# TLDR\nThis competition is very similar to speech recognition. Therefore, I attempted new ideas focusing on techniques used in speech recognition. Among them, the most effective ones were **conformer architecture**,  **interCTC technique**, **strong augmentation** and **nomasking**.\n<br>\n\n# Data Processing\nOut of the provided 543 landmarks, I utilized 42 hand data, 33 pose data, and 40 lip data. Input features were constructed by concatenating xy coordinates (drop z) and motion (xy[1:] - xy[:-1]).\n\n * preprocess\n\n        def pre_process0(x):\n          x = x[:args.max_frame]\n          x = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n          n_frames = tf.shape(x)[0]\n\n          lhand = tf.transpose(tf.reshape(x[:, 0:63], (n_frames, 3, args.n_hand_landmarks)), (0, 2, 1))\n          rhand = tf.transpose(tf.reshape(x[:, 63:126], (n_frames, 3, args.n_hand_landmarks)), (0, 2, 1))\n          pose = tf.transpose(tf.reshape(x[:, 126:225], (n_frames, 3, args.n_pose_landmarks)), (0, 2, 1))\n          face = tf.transpose(tf.reshape(x[:, 225:345], (n_frames, 3, args.n_face_landmarks)), (0, 2, 1))\n\n          x = tf.concat([\n              lhand,\n              rhand,\n              pose,\n              face\n              ], axis=1)\n\n          x = x[:, :, :2]\n          return x\n\n        def decode_fn(record_bytes, augment=False):\n          schema = {\n              'coordinates': tf.io.VarLenFeature(tf.float32),\n              'phrase': tf.io.VarLenFeature(tf.int64)\n          }\n          x = tf.io.parse_single_example(record_bytes, schema)\n\n          coordinates = tf.reshape(tf.sparse.to_dense(x[\"coordinates\"]), (-1, args.input_dim))\n          phrase = tf.sparse.to_dense(x[\"phrase\"])\n\n          if augment:\n            coordinates, phrase = augment_fn(coordinates, phrase)\n\n          dx = tf.cond(tf.shape(coordinates)[0]>1,lambda:tf.pad(coordinates[1:] - coordinates[:-1], [[0,1],[0,0]]),lambda:tf.zeros_like(coordinates))\n          coordinates = tf.concat([coordinates, dx], axis=-1)\n\n          return coordinates, phrase\n\n<br>\n\n# Augmentation\nThree types of augmentations were applied, each of which significantly influenced CV.\n\n* flip hand(CV improvement : ~0.003)\n        \n        # flip hand\n        def flip_hand(video):\n          video = tf.reshape(video, shape=(-1, args.n_landmarks, 2))\n          hands = video[:, :int(2 * args.n_hand_landmarks)]\n          other = video[:, int(2 * args.n_hand_landmarks):]\n\n          lhand = hands[:, :args.n_hand_landmarks]\n          rhand = hands[:, args.n_hand_landmarks:]\n\n          lhand_x, rhand_x = lhand[:, :, 0], rhand[:, :, 0]\n          lhand_x = tf.negative(lhand_x) + 2 * tf.reduce_mean(lhand_x, axis=1, keepdims=True)\n          rhand_x = tf.negative(rhand_x) + 2 * tf.reduce_mean(rhand_x, axis=1, keepdims=True)\n\n          lhand = tf.concat([tf.expand_dims(lhand_x, axis=-1), lhand[:, :, 1:]], axis=-1)\n          rhand = tf.concat([tf.expand_dims(rhand_x, axis=-1), rhand[:, :, 1:]], axis=-1)\n\n          flipped_hands = tf.concat([rhand, lhand, other], axis=1)\n          flipped_hands = tf.reshape(flipped_hands, shape=(-1, args.input_dim))\n          return flipped_hands\n\n* flip video and phrase(CV improvement : ~0.005)\n          \n        # flip video and phrase\n        def reverse_frames(x, y):\n          x = x[::-1]\n          y = y[::-1]\n          return x, y\n\n* concat 2 videos and 2 phrases(CV improvement : ~0.01)\n        \n        # concat 2 videos and 2 phrases\n        def cat_augment(inputs, inputs2):\n          x, y = inputs\n          x2, y2 = inputs2\n\n          x_shape = tf.shape(x)\n          x2_shape = tf.shape(x2)\n\n          should_concat = tf.random.uniform(()) < 0.5\n\n          x_condition = should_concat & (x_shape[0] + x2_shape[0] < args.max_frame)\n\n          x = tf.cond(x_condition, lambda: tf.concat([x, x2], axis=0), lambda: x)\n          y = tf.cond(x_condition, lambda: tf.concat([y, y2], axis=0), lambda: y)\n\n          return x, y\n\n<br>\n\n# Model\nThe model was constructed utilizing the architectures proposed in Conformer.  At first, I employed the Conformer code from https://github.com/TensorSpeech/TensorFlowASR. However, by making specific modifications to the code utilized by @hoyso48 , I successfully implemented the Conformer architecture. This implementation led to an enhancement of 0.008 in the LB score. This improvement is likely due to the removal of residual blocks between Conformer blocks, allowing for the integration of more layers and resulting in improved robustness against overfitting.  Furthermore, the incorporation of the interCTC([paper](https://arxiv.org/abs/2102.03216)) contributed to an LB score(~0.005).\n\n\nFurthermore, as the model's output was computed without masking and a maximum length of 384 was employed for the CTC loss \"input-length' variable, using a larger kernel size(=31) effectively conveyed frame information across the 384-length range, resulting in improved performance.\n\n* The overall structure of the model is as follows.\n\n        def get_model(max_len=384, dim=160,ksize=31,drop_rate=0.1, num_layers=16):\n          NUM_CLASSES=63\n          PAD=0\n          inp = tf.keras.Input((None,2*230))\n          x = inp\n          x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n          x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n          xs = []\n          for i in range(num_layers):\n              x = TransformerBlock(dim,expand=2)(x)\n              x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n              xs.append(x)\n\n          classifier = tf.keras.layers.Dense(NUM_CLASSES,name='classifier')\n\n          x1 = tf.keras.layers.Dropout(0.2)(xs[-8])\n          x1 = classifier(x1)\n\n          x2 = tf.keras.layers.Dropout(0.2)(xs[-1])\n          x2 = classifier(x2)\n          return x1, x2\n<br>\n\n# Training\n\n* Scheduler : lr_warmup_cosine_decay \n* Warmup Ratio : 0.1 \n* Optimizer : AdamW \n* Weight Decay : 0.01\n* Epoch : 300(last 100 epoch for only train data(not suppl))\n* Learning Rate : 1e-3 \n* Loss Function : CTCLoss\n\n<br>\n\n# Didn't Work\n* masking \n* adding distance features\n* mlm pretraining \n<br>\n\n# Code\nhttps://github.com/siwooyong/Google-American-Sign-Language-Fingerspelling-Recognition\n<br>",
    "2409082": "Great Work Man!",
    "2407604": "Congratulations. Thanks for sharing the details. I observe with 50 epochs, the notebook was consuming around 9 hrs. Can you advise how you could increase the number of epochs to 300 while keeping the notebook time <9hrs.",
    "2407440": "Thanks for sharing, 'concat 2 videos and 2 phrases' how much does this help?"
  }
}