{
  "id": 434380,
  "title": "95th Bronze solutions: Focal loss + KD + Supplemental data",
  "url": "/competitions/asl-fingerspelling/discussion/434380",
  "author_name": "Chakkrit Termritthikun",
  "post_date": "2023-08-25T03:01:50.021000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<h2><strong>95th Bronze Solutions (public score 0.714, private score 0.671)</strong></h2>\n<p>I express my gratitude to Kaggle and Google for organizing this competition. Participating in this competition has been an invaluable learning experience for me.</p>\n<p>One significant achievement was improving the public score from Version 17 of Saidineshpola's notebook, which initially achieved a public score of 0.689. The improvements were achieved by implementing the following modifications:</p>\n<h1><strong>Model &amp; Config:</strong></h1>\n<ul>\n<li>Set model dimensions (Dim) to 192.</li>\n<li>Set 16 CNN-Transformer blocks.</li>\n<li>Incorporated a dropout rate of 0.4.</li>\n<li>Utilized 8 attention heads and an expansion factor of 8 in the Transformer blocks.</li>\n<li>The total number of parameters amounted to 19,255,980.</li>\n<li>Employed tf.float16 for model conversion.</li>\n<li>Configured training with 100 epochs, with 10 warm-up epochs.</li>\n<li>Set the maximum learning rate (LR_MAX) to 1e-3.</li>\n<li>Set the weight decay (WD_RATIO) to 1e-4.</li>\n</ul>\n<h1><strong>CTC loss:</strong></h1>\n<p>I utilized the CTC loss function from the <code>pip install tf-seq2seq-losses</code> library, which provided comparable results to TensorFlow's built-in function. This implementation offered a speed enhancement of around 25% on V100 GPUs. Additionally, I applied the Focal loss to the CTC loss function.</p>\n<pre><code> tf_seq2seq_losses  classic_ctc_loss \n ():\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-)\n    logit_length = tf.ones(tf.shape(logits)[], dtype=tf.int32) * tf.shape(logits)[]\n    ctc_loss = classic_ctc_loss(\n            labels=tf.cast(labels, dtype=tf.int32),\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n    )\n    p= tf.exp(-ctc_loss)\n    alpha = \n    gamma = \n    focal_ctc_loss= tf.multiply(tf.multiply(alpha,tf.((-p),gamma)),ctc_loss) \n    loss = tf.reduce_mean(focal_ctc_loss)\n     loss\n</code></pre>\n<h1><strong>Data augmentation:</strong></h1>\n<p>No data augmentation techniques were applied.</p>\n<h1><strong>Additional data:</strong></h1>\n<p>Supplemental data was integrated into the training process. I followed <a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset/notebook\" target=\"_blank\">ROHITH INGILELA's notebook</a> to generate supplemental TFrecord files. The resources used are available <a href=\"https://www.kaggle.com/datasets/kongpasom/aslfr-sup-tfrecord\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/datasets/kongpasom/aslfr-sup\" target=\"_blank\">here</a> .</p>\n<h1><strong>Training Approach:</strong></h1>\n<ol>\n<li>Trained the model using supplemental data for 100 epochs.</li>\n<li>Trained the model with training data and the pre-trained weights derived from the - supplemental data for an additional 100 epochs (resulting in a public score of 0.713).</li>\n<li>Executed a fine-tuning process using knowledge distillation, refining the model with the best pre-trained weights through the AdamW optimizer, resulting in a modest score improvement of +0.001.</li>\n</ol>\n<h1><strong>Model Architecture:</strong></h1>\n<p>Here's a summarized version of the model architecture:</p>\n<pre><code> ():\n    inp = tf.keras.Input(INPUT_SHAPE)\n    x = tf.keras.layers.Masking(mask_value=)(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=, name=)(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[], dim), dtype=x.dtype)\n    x = x + pe\n    x = tf.keras.layers.BatchNormalization(name=)(x)\n\n    num_blocks = \n    drop_rate  = \n     i  (num_blocks):\n        x = Conv1DBlock(dim, , drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  , drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  , drop_rate=drop_rate)(x)\n        x = TransformerBlock(dim=dim, num_heads=, expand=, attn_dropout=, drop_rate=)(x)\n\n    x = tf.keras.layers.Dense(dim*,activation=,name=)(x)\n    x = tf.keras.layers.Dropout()(x)\n    x = tf.keras.layers.Dense((char_to_num), name=)(x)\n    model = tf.keras.Model(inp, x)\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=)\n\n     model\n</code></pre>\n<h1><strong>Knowledge Distillation Training :</strong></h1>\n<p>A training step using knowledge distillation was employed to refine the model:</p>\n<pre><code>loss_fn = CTCLoss\nkld_loss_fn = tf.keras.losses.KLDivergence()\n\n\n ():\n    temperature = \n    alpha = \n\n    teacher_pred = model_teacher(x, training=)\n\n     tf.GradientTape()  tape:\n        logits = model(x, training=)\n        loss_value = loss_fn(y, logits)\n\n        kld_loss = kld_loss_fn(\n            tf.nn.softmax(teacher_pred / temperature, axis=),\n            tf.nn.softmax(logits / temperature, axis=),\n        ) * temperature**\n\n        loss = alpha * loss_value + ( - alpha) * kld_loss\n\n    grads = tape.gradient(loss, model.trainable_weights)\n    model.optimizer.apply_gradients((grads, model.trainable_weights))\n\n     loss, loss_value, kld_loss\n</code></pre>\n<h1><strong>Reference (</strong>If I have seen further, it is by standing on the shoulders of giants.<strong>) :</strong></h1>\n<p><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA's notebook</a><br>\n<a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">RSAIDINESHPOLA's notebook</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset\" target=\"_blank\">ROHITH INGILELA's pre-process dataset notebook</a></p>",
  "messages": [
    {
      "id": 2407320,
      "postDate": "2023-08-25T03:01:50.020Z",
      "content": "<h2><strong>95th Bronze Solutions (public score 0.714, private score 0.671)</strong></h2>\n<p>I express my gratitude to Kaggle and Google for organizing this competition. Participating in this competition has been an invaluable learning experience for me.</p>\n<p>One significant achievement was improving the public score from Version 17 of Saidineshpola's notebook, which initially achieved a public score of 0.689. The improvements were achieved by implementing the following modifications:</p>\n<h1><strong>Model &amp; Config:</strong></h1>\n<ul>\n<li>Set model dimensions (Dim) to 192.</li>\n<li>Set 16 CNN-Transformer blocks.</li>\n<li>Incorporated a dropout rate of 0.4.</li>\n<li>Utilized 8 attention heads and an expansion factor of 8 in the Transformer blocks.</li>\n<li>The total number of parameters amounted to 19,255,980.</li>\n<li>Employed tf.float16 for model conversion.</li>\n<li>Configured training with 100 epochs, with 10 warm-up epochs.</li>\n<li>Set the maximum learning rate (LR_MAX) to 1e-3.</li>\n<li>Set the weight decay (WD_RATIO) to 1e-4.</li>\n</ul>\n<h1><strong>CTC loss:</strong></h1>\n<p>I utilized the CTC loss function from the <code>pip install tf-seq2seq-losses</code> library, which provided comparable results to TensorFlow's built-in function. This implementation offered a speed enhancement of around 25% on V100 GPUs. Additionally, I applied the Focal loss to the CTC loss function.</p>\n<pre><code> tf_seq2seq_losses  classic_ctc_loss \n ():\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-)\n    logit_length = tf.ones(tf.shape(logits)[], dtype=tf.int32) * tf.shape(logits)[]\n    ctc_loss = classic_ctc_loss(\n            labels=tf.cast(labels, dtype=tf.int32),\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n    )\n    p= tf.exp(-ctc_loss)\n    alpha = \n    gamma = \n    focal_ctc_loss= tf.multiply(tf.multiply(alpha,tf.((-p),gamma)),ctc_loss) \n    loss = tf.reduce_mean(focal_ctc_loss)\n     loss\n</code></pre>\n<h1><strong>Data augmentation:</strong></h1>\n<p>No data augmentation techniques were applied.</p>\n<h1><strong>Additional data:</strong></h1>\n<p>Supplemental data was integrated into the training process. I followed <a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset/notebook\" target=\"_blank\">ROHITH INGILELA's notebook</a> to generate supplemental TFrecord files. The resources used are available <a href=\"https://www.kaggle.com/datasets/kongpasom/aslfr-sup-tfrecord\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/datasets/kongpasom/aslfr-sup\" target=\"_blank\">here</a> .</p>\n<h1><strong>Training Approach:</strong></h1>\n<ol>\n<li>Trained the model using supplemental data for 100 epochs.</li>\n<li>Trained the model with training data and the pre-trained weights derived from the - supplemental data for an additional 100 epochs (resulting in a public score of 0.713).</li>\n<li>Executed a fine-tuning process using knowledge distillation, refining the model with the best pre-trained weights through the AdamW optimizer, resulting in a modest score improvement of +0.001.</li>\n</ol>\n<h1><strong>Model Architecture:</strong></h1>\n<p>Here's a summarized version of the model architecture:</p>\n<pre><code> ():\n    inp = tf.keras.Input(INPUT_SHAPE)\n    x = tf.keras.layers.Masking(mask_value=)(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=, name=)(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[], dim), dtype=x.dtype)\n    x = x + pe\n    x = tf.keras.layers.BatchNormalization(name=)(x)\n\n    num_blocks = \n    drop_rate  = \n     i  (num_blocks):\n        x = Conv1DBlock(dim, , drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  , drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  , drop_rate=drop_rate)(x)\n        x = TransformerBlock(dim=dim, num_heads=, expand=, attn_dropout=, drop_rate=)(x)\n\n    x = tf.keras.layers.Dense(dim*,activation=,name=)(x)\n    x = tf.keras.layers.Dropout()(x)\n    x = tf.keras.layers.Dense((char_to_num), name=)(x)\n    model = tf.keras.Model(inp, x)\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=)\n\n     model\n</code></pre>\n<h1><strong>Knowledge Distillation Training :</strong></h1>\n<p>A training step using knowledge distillation was employed to refine the model:</p>\n<pre><code>loss_fn = CTCLoss\nkld_loss_fn = tf.keras.losses.KLDivergence()\n\n\n ():\n    temperature = \n    alpha = \n\n    teacher_pred = model_teacher(x, training=)\n\n     tf.GradientTape()  tape:\n        logits = model(x, training=)\n        loss_value = loss_fn(y, logits)\n\n        kld_loss = kld_loss_fn(\n            tf.nn.softmax(teacher_pred / temperature, axis=),\n            tf.nn.softmax(logits / temperature, axis=),\n        ) * temperature**\n\n        loss = alpha * loss_value + ( - alpha) * kld_loss\n\n    grads = tape.gradient(loss, model.trainable_weights)\n    model.optimizer.apply_gradients((grads, model.trainable_weights))\n\n     loss, loss_value, kld_loss\n</code></pre>\n<h1><strong>Reference (</strong>If I have seen further, it is by standing on the shoulders of giants.<strong>) :</strong></h1>\n<p><a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">ROHITH INGILELA's notebook</a><br>\n<a href=\"https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">RSAIDINESHPOLA's notebook</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset\" target=\"_blank\">ROHITH INGILELA's pre-process dataset notebook</a></p>",
      "rawMarkdown": "## **95th Bronze Solutions (public score 0.714, private score 0.671)**\n\nI express my gratitude to Kaggle and Google for organizing this competition. Participating in this competition has been an invaluable learning experience for me.\n\nOne significant achievement was improving the public score from Version 17 of Saidineshpola's notebook, which initially achieved a public score of 0.689. The improvements were achieved by implementing the following modifications:\n\n# **Model & Config:**\n- Set model dimensions (Dim) to 192.\n- Set 16 CNN-Transformer blocks.\n- Incorporated a dropout rate of 0.4.\n- Utilized 8 attention heads and an expansion factor of 8 in the Transformer blocks.\n- The total number of parameters amounted to 19,255,980.\n- Employed tf.float16 for model conversion.\n- Configured training with 100 epochs, with 10 warm-up epochs.\n- Set the maximum learning rate (LR_MAX) to 1e-3.\n- Set the weight decay (WD_RATIO) to 1e-4.\n\n# **CTC loss:**\nI utilized the CTC loss function from the `pip install tf-seq2seq-losses` library, which provided comparable results to TensorFlow's built-in function. This implementation offered a speed enhancement of around 25% on V100 GPUs. Additionally, I applied the Focal loss to the CTC loss function.\n\n```python\nfrom tf_seq2seq_losses import classic_ctc_loss \ndef CTCLoss(labels, logits):\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-1)\n    logit_length = tf.ones(tf.shape(logits)[0], dtype=tf.int32) * tf.shape(logits)[1]\n    ctc_loss = classic_ctc_loss(\n            labels=tf.cast(labels, dtype=tf.int32),\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n    )\n    p= tf.exp(-ctc_loss)\n    alpha = 0.25\n    gamma = 2.0\n    focal_ctc_loss= tf.multiply(tf.multiply(alpha,tf.pow((1-p),gamma)),ctc_loss) #((alpha)*((1-p)**gamma)*(ctc_loss))\n    loss = tf.reduce_mean(focal_ctc_loss)\n    return loss\n\n```\n\n# **Data augmentation:**\nNo data augmentation techniques were applied.\n\n# **Additional data:**\n\nSupplemental data was integrated into the training process. I followed [ROHITH INGILELA's notebook](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset/notebook) to generate supplemental TFrecord files. The resources used are available [here](https://www.kaggle.com/datasets/kongpasom/aslfr-sup-tfrecord) and [here](https://www.kaggle.com/datasets/kongpasom/aslfr-sup) .\n\n# **Training Approach:**\n\n1. Trained the model using supplemental data for 100 epochs.\n2. Trained the model with training data and the pre-trained weights derived from the - supplemental data for an additional 100 epochs (resulting in a public score of 0.713).\n3. Executed a fine-tuning process using knowledge distillation, refining the model with the best pre-trained weights through the AdamW optimizer, resulting in a modest score improvement of +0.001.\n\n# **Model Architecture:**\n\nHere's a summarized version of the model architecture:\n\n```python\n\ndef get_model(dim = 192):\n    inp = tf.keras.Input(INPUT_SHAPE)\n    x = tf.keras.layers.Masking(mask_value=0.0)(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=False, name='stem_conv')(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[0], dim), dtype=x.dtype)\n    x = x + pe\n    x = tf.keras.layers.BatchNormalization(name='stem_bn')(x)\n\n    num_blocks = 16\n    drop_rate  = 0.4\n    for i in range(num_blocks):\n        x = Conv1DBlock(dim, 11, drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  5, drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  3, drop_rate=drop_rate)(x)\n        x = TransformerBlock(dim=dim, num_heads=8, expand=8, attn_dropout=0.4, drop_rate=0.4)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation='relu',name='top_conv')(x)\n    x = tf.keras.layers.Dropout(0.4)(x)\n    x = tf.keras.layers.Dense(len(char_to_num), name='classifier')(x)\n    model = tf.keras.Model(inp, x)\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n    return model\n```\n\n# **Knowledge Distillation Training :**\nA training step using knowledge distillation was employed to refine the model:\n\n```python\n\nloss_fn = CTCLoss\nkld_loss_fn = tf.keras.losses.KLDivergence()\n\n@tf.function\ndef train_step(x, y):\n    temperature = 20\n    alpha = 0.5\n\n    teacher_pred = model_teacher(x, training=False)\n\n    with tf.GradientTape() as tape:\n        logits = model(x, training=True)\n        loss_value = loss_fn(y, logits)\n\n        kld_loss = kld_loss_fn(\n            tf.nn.softmax(teacher_pred / temperature, axis=1),\n            tf.nn.softmax(logits / temperature, axis=1),\n        ) * temperature**2\n\n        loss = alpha * loss_value + (1 - alpha) * kld_loss\n\n    grads = tape.gradient(loss, model.trainable_weights)\n    model.optimizer.apply_gradients(zip(grads, model.trainable_weights))\n\n    return loss, loss_value, kld_loss\n\n```\n\n# **Reference (**If I have seen further, it is by standing on the shoulders of giants.**) :**\n[ROHITH INGILELA's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place)\n[RSAIDINESHPOLA's notebook](https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place)\n[ROHITH INGILELA's pre-process dataset notebook](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset)",
      "votes": 3
    },
    {
      "id": 2408314,
      "postDate": "2023-08-25T15:05:14.757Z",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution.<br>\nWith 16 Blocks, how did you manage to limit the notebook time to &lt;9hrs?</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the details of your solution.\nWith 16 Blocks, how did you manage to limit the notebook time to <9hrs?",
      "replies": [
        {
          "id": 2408563,
          "postDate": "2023-08-25T17:44:20.210Z",
          "content": "<p>Thank you, Kumar. Utilizing TPU would result in quicker performance compared to a GPU.</p>",
          "rawMarkdown": "Thank you, Kumar. Utilizing TPU would result in quicker performance compared to a GPU."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2408314,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-08-25T15:05:14.757000",
      "content": "<p>Congratulations. Thanks for sharing the details of your solution.<br>\nWith 16 Blocks, how did you manage to limit the notebook time to &lt;9hrs?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2408563,
          "author_name": "Chakkrit Termritthikun",
          "author_url": "",
          "post_date": "2023-08-25T17:44:20.210000",
          "content": "<p>Thank you, Kumar. Utilizing TPU would result in quicker performance compared to a GPU.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2407320": "## **95th Bronze Solutions (public score 0.714, private score 0.671)**\n\nI express my gratitude to Kaggle and Google for organizing this competition. Participating in this competition has been an invaluable learning experience for me.\n\nOne significant achievement was improving the public score from Version 17 of Saidineshpola's notebook, which initially achieved a public score of 0.689. The improvements were achieved by implementing the following modifications:\n\n# **Model & Config:**\n- Set model dimensions (Dim) to 192.\n- Set 16 CNN-Transformer blocks.\n- Incorporated a dropout rate of 0.4.\n- Utilized 8 attention heads and an expansion factor of 8 in the Transformer blocks.\n- The total number of parameters amounted to 19,255,980.\n- Employed tf.float16 for model conversion.\n- Configured training with 100 epochs, with 10 warm-up epochs.\n- Set the maximum learning rate (LR_MAX) to 1e-3.\n- Set the weight decay (WD_RATIO) to 1e-4.\n\n# **CTC loss:**\nI utilized the CTC loss function from the `pip install tf-seq2seq-losses` library, which provided comparable results to TensorFlow's built-in function. This implementation offered a speed enhancement of around 25% on V100 GPUs. Additionally, I applied the Focal loss to the CTC loss function.\n\n```python\nfrom tf_seq2seq_losses import classic_ctc_loss \ndef CTCLoss(labels, logits):\n    label_length = tf.reduce_sum(tf.cast(labels != pad_token_idx, tf.int32), axis=-1)\n    logit_length = tf.ones(tf.shape(logits)[0], dtype=tf.int32) * tf.shape(logits)[1]\n    ctc_loss = classic_ctc_loss(\n            labels=tf.cast(labels, dtype=tf.int32),\n            logits=logits,\n            label_length=label_length,\n            logit_length=logit_length,\n            blank_index=pad_token_idx,\n    )\n    p= tf.exp(-ctc_loss)\n    alpha = 0.25\n    gamma = 2.0\n    focal_ctc_loss= tf.multiply(tf.multiply(alpha,tf.pow((1-p),gamma)),ctc_loss) #((alpha)*((1-p)**gamma)*(ctc_loss))\n    loss = tf.reduce_mean(focal_ctc_loss)\n    return loss\n\n```\n\n# **Data augmentation:**\nNo data augmentation techniques were applied.\n\n# **Additional data:**\n\nSupplemental data was integrated into the training process. I followed [ROHITH INGILELA's notebook](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset/notebook) to generate supplemental TFrecord files. The resources used are available [here](https://www.kaggle.com/datasets/kongpasom/aslfr-sup-tfrecord) and [here](https://www.kaggle.com/datasets/kongpasom/aslfr-sup) .\n\n# **Training Approach:**\n\n1. Trained the model using supplemental data for 100 epochs.\n2. Trained the model with training data and the pre-trained weights derived from the - supplemental data for an additional 100 epochs (resulting in a public score of 0.713).\n3. Executed a fine-tuning process using knowledge distillation, refining the model with the best pre-trained weights through the AdamW optimizer, resulting in a modest score improvement of +0.001.\n\n# **Model Architecture:**\n\nHere's a summarized version of the model architecture:\n\n```python\n\ndef get_model(dim = 192):\n    inp = tf.keras.Input(INPUT_SHAPE)\n    x = tf.keras.layers.Masking(mask_value=0.0)(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=False, name='stem_conv')(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[0], dim), dtype=x.dtype)\n    x = x + pe\n    x = tf.keras.layers.BatchNormalization(name='stem_bn')(x)\n\n    num_blocks = 16\n    drop_rate  = 0.4\n    for i in range(num_blocks):\n        x = Conv1DBlock(dim, 11, drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  5, drop_rate=drop_rate)(x)\n        x = Conv1DBlock(dim,  3, drop_rate=drop_rate)(x)\n        x = TransformerBlock(dim=dim, num_heads=8, expand=8, attn_dropout=0.4, drop_rate=0.4)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation='relu',name='top_conv')(x)\n    x = tf.keras.layers.Dropout(0.4)(x)\n    x = tf.keras.layers.Dense(len(char_to_num), name='classifier')(x)\n    model = tf.keras.Model(inp, x)\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold=4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period=5)\n\n    return model\n```\n\n# **Knowledge Distillation Training :**\nA training step using knowledge distillation was employed to refine the model:\n\n```python\n\nloss_fn = CTCLoss\nkld_loss_fn = tf.keras.losses.KLDivergence()\n\n@tf.function\ndef train_step(x, y):\n    temperature = 20\n    alpha = 0.5\n\n    teacher_pred = model_teacher(x, training=False)\n\n    with tf.GradientTape() as tape:\n        logits = model(x, training=True)\n        loss_value = loss_fn(y, logits)\n\n        kld_loss = kld_loss_fn(\n            tf.nn.softmax(teacher_pred / temperature, axis=1),\n            tf.nn.softmax(logits / temperature, axis=1),\n        ) * temperature**2\n\n        loss = alpha * loss_value + (1 - alpha) * kld_loss\n\n    grads = tape.gradient(loss, model.trainable_weights)\n    model.optimizer.apply_gradients(zip(grads, model.trainable_weights))\n\n    return loss, loss_value, kld_loss\n\n```\n\n# **Reference (**If I have seen further, it is by standing on the shoulders of giants.**) :**\n[ROHITH INGILELA's notebook](https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place)\n[RSAIDINESHPOLA's notebook](https://www.kaggle.com/code/saidineshpola/aslfr-ctc-based-on-prev-comp-1st-place)\n[ROHITH INGILELA's pre-process dataset notebook](https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset)",
    "2408314": "Congratulations. Thanks for sharing the details of your solution.\nWith 16 Blocks, how did you manage to limit the notebook time to <9hrs?"
  }
}