{
  "id": 437123,
  "title": "449th Position Solution for Google ASLFR Competition - Adversarial Regularization, Quantization Aware Training, and KD Code for Tensorflow! ",
  "url": "/competitions/asl-fingerspelling/discussion/437123",
  "author_name": "Aaryam Sharma",
  "post_date": "2023-09-05T13:42:29.159000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Everyone! I'm just posting my solution here, in case it could help someone! I used a lot of techniques, but didn't have the computing power to push the boundaries required for a very high rank. Nevertheless, I learnt a lot of techniques, and am writing this solution here to help provide useful code!</p>\n<h1>Competition Context</h1>\n<p>The context of this solution is the Google - American Sign Language Fingerspelling Recognition Competition:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></p>\n<p>The data of this competition can be found here:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></p>\n<h1>My Approach</h1>\n<p>My approach was identical to most people, I referred to the really awesome notebook by Rohit Ingilela here: <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a></p>\n<p>I made several additions and changes, namely by adding Quantization Aware Training, AWP, and Knowledge distillation (not necessarily in the same model). I tested a lot of parameters as well and made a lot of models.</p>\n<p>My idea was to use Quantization to increase model size. I found that after 24 Million parameters, even though the model size was small enough, the performance became too poor to run within the 5 hours allocated for testing. I found that KD didn't improve my performance too much, so I didn't explore it a lot. I tried AWP only at the end, and even though it improved performance by around 0.003, I didn't get to test it a lot.</p>\n<p>Nevertheless, here is some code for each idea:</p>\n<p><strong>Quantization:</strong></p>\n<pre><code>!pip install -q tensorflow-model-optimization\n tensorflow_model_optimization  tfmot\n tensorflow.keras  layers\n\n\nx = tfmot.quantization.keras.quantize_annotate_layer(layers.Dense(channels, activation = activation, kernel_regularizer = regularizer))(inputs)\n\n\n\n\n\ntflitemodel_base = TFLiteModel(model)\ntflitemodel_base(frames)[].shape\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\n\n\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT] \n\n\n\n\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]\ntflite_model = keras_model_converter.convert()\n (, )  f:\n    f.write(tflite_model)\n</code></pre>\n<p><strong>Adversarial Regularization in Training</strong></p>\n<pre><code>! pip install -q neural-structured-learning\n\n neural_structured_learning  nsl\n\n\nadv_config = nsl.configs.make_adv_reg_config(\n    multiplier = ,\n    adv_step_size = ,\n    adv_grad_norm = \n)\n\n\n ():\n     {: image, : label}\n\ntrain_dataset = train_dataset.(convert_to_dictionaries)\nval_dataset = val_dataset.(convert_to_dictionaries)\n\n\n\n\n strategy.scope():\n    model = nsl.keras.AdversarialRegularization(\n        model,\n        label_keys = [],\n        adv_config = adv_config\n    )\n\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold = )\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period = )\n\n    model.(loss = loss, optimizer = optimizer)\n\n\n\n</code></pre>\n<p><strong>Knowledge Distillation</strong></p>\n<pre><code> (keras.Model):\n     ():\n        ().__init__()\n        self.student = student\n        self.teacher = teacher\n\n     (): \n        ().(optimizer=optimizer)\n        self.loss = CTCLoss\n        self.distiller_loss_fn = keras.losses.KLDivergence()\n        self.alpha = alpha\n        self.temp = temperature\n\n     ():\n        x, y = data\n        \n        is_batch_small = tf.math.less(tf.shape(x)[], train_batch_size)\n        is_batch_small = tf.math.reduce_all(is_batch_small)\n\n         ():\n             tf.constant(), tf.constant()\n\n         ():\n            teacher_pred = self.teacher(x, training = )\n             tf.GradientTape()  tape:\n                student_pred = self.student(x, training = )\n                student_loss = self.loss(y, student_pred)\n\n                distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(student_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                    ) * (self.temp ** )\n                loss = self.alpha * student_loss + ( - self.alpha) * distilled_loss\n\n            trainable_vars = self.student.trainable_variables\n            gradients = tape.gradient(loss, trainable_vars)\n            self.optimizer.apply_gradients((gradients, trainable_vars))\n            self.compiled_metrics.update_state(y, student_pred)\n\n             student_loss, distilled_loss\n\n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n\n        results = {m.name : m.result()  m  self.metrics}\n        results.update(\n            {: student_loss, : distilled_loss}\n        )\n\n         results\n\n     ():\n        ()\n        x, y = data\n        is_batch_small = tf.less(tf.shape(x)[],  val_batch_size)\n        is_batch_small = tf.reduce_all(is_batch_small)\n\n         ():\n             tf.constant(), tf.constant()\n\n         ():\n            y_pred = self.student(x, training = )\n            teacher_pred = self.teacher(x, training = )\n\n            student_loss = self.loss(y, y_pred)\n\n            distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(y_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                ) * (self.temp ** )\n\n            self.compiled_metrics.update_state(y, y_pred)\n             student_loss, distilled_loss\n\n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n\n        results = {m.name : m.result()  m  self.metrics}\n        results.update(\n            {: student_loss, : distilled_loss}\n        )\n\n         results\n\n\n</code></pre>\n<p>While I didn't get a super high rank, I did learn a lot and used a plethora of techniques! I wrote most things from scratch, to learn things myself, and referred to other's work only to get a better understanding of what they did and improve my own code as a result.</p>\n<p><em>Validation</em><br>\nQuite standard - kept a holdout set and computer Normalized Levenshtein distance on it after every epoch.<br>\nNo CV due to low computational resources.</p>\n<h1>Improvements</h1>\n<p>I am quite disappointed that the things I missed out were very trivial - such as not training for more than 300 epochs (I could only manage 100 to try enough models and variations out) and that keeping some missing hands frames was important.</p>\n<p>I also didn't use enough augmentations which was important, even though I used all the ones used by top submissions except cutout.</p>\n<p>I also read about Squeezeformer, but got demotivated because I saw high scoring public submissions without Squeezeformer :(</p>\n<p>To further improve, I'll correct all the points above.</p>\n<h1>Final Note</h1>\n<p>I am quite happy to have participated in this fun and awesome contest. I really learnt a lot, and experimented a lot. I am certain that what I did here will help me in future competitions, research and work.</p>\n<p>If you could give me any suggestions/feedback on my code above, please do so. It will help me improve and learn, and I will be very grateful for your help!</p>",
  "messages": [
    {
      "id": 2424785,
      "postDate": "2023-09-05T13:42:29.160Z",
      "content": "<p>Hi Everyone! I'm just posting my solution here, in case it could help someone! I used a lot of techniques, but didn't have the computing power to push the boundaries required for a very high rank. Nevertheless, I learnt a lot of techniques, and am writing this solution here to help provide useful code!</p>\n<h1>Competition Context</h1>\n<p>The context of this solution is the Google - American Sign Language Fingerspelling Recognition Competition:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></p>\n<p>The data of this competition can be found here:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></p>\n<h1>My Approach</h1>\n<p>My approach was identical to most people, I referred to the really awesome notebook by Rohit Ingilela here: <a href=\"https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place</a></p>\n<p>I made several additions and changes, namely by adding Quantization Aware Training, AWP, and Knowledge distillation (not necessarily in the same model). I tested a lot of parameters as well and made a lot of models.</p>\n<p>My idea was to use Quantization to increase model size. I found that after 24 Million parameters, even though the model size was small enough, the performance became too poor to run within the 5 hours allocated for testing. I found that KD didn't improve my performance too much, so I didn't explore it a lot. I tried AWP only at the end, and even though it improved performance by around 0.003, I didn't get to test it a lot.</p>\n<p>Nevertheless, here is some code for each idea:</p>\n<p><strong>Quantization:</strong></p>\n<pre><code>!pip install -q tensorflow-model-optimization\n tensorflow_model_optimization  tfmot\n tensorflow.keras  layers\n\n\nx = tfmot.quantization.keras.quantize_annotate_layer(layers.Dense(channels, activation = activation, kernel_regularizer = regularizer))(inputs)\n\n\n\n\n\ntflitemodel_base = TFLiteModel(model)\ntflitemodel_base(frames)[].shape\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\n\n\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT] \n\n\n\n\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]\ntflite_model = keras_model_converter.convert()\n (, )  f:\n    f.write(tflite_model)\n</code></pre>\n<p><strong>Adversarial Regularization in Training</strong></p>\n<pre><code>! pip install -q neural-structured-learning\n\n neural_structured_learning  nsl\n\n\nadv_config = nsl.configs.make_adv_reg_config(\n    multiplier = ,\n    adv_step_size = ,\n    adv_grad_norm = \n)\n\n\n ():\n     {: image, : label}\n\ntrain_dataset = train_dataset.(convert_to_dictionaries)\nval_dataset = val_dataset.(convert_to_dictionaries)\n\n\n\n\n strategy.scope():\n    model = nsl.keras.AdversarialRegularization(\n        model,\n        label_keys = [],\n        adv_config = adv_config\n    )\n\n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold = )\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period = )\n\n    model.(loss = loss, optimizer = optimizer)\n\n\n\n</code></pre>\n<p><strong>Knowledge Distillation</strong></p>\n<pre><code> (keras.Model):\n     ():\n        ().__init__()\n        self.student = student\n        self.teacher = teacher\n\n     (): \n        ().(optimizer=optimizer)\n        self.loss = CTCLoss\n        self.distiller_loss_fn = keras.losses.KLDivergence()\n        self.alpha = alpha\n        self.temp = temperature\n\n     ():\n        x, y = data\n        \n        is_batch_small = tf.math.less(tf.shape(x)[], train_batch_size)\n        is_batch_small = tf.math.reduce_all(is_batch_small)\n\n         ():\n             tf.constant(), tf.constant()\n\n         ():\n            teacher_pred = self.teacher(x, training = )\n             tf.GradientTape()  tape:\n                student_pred = self.student(x, training = )\n                student_loss = self.loss(y, student_pred)\n\n                distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(student_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                    ) * (self.temp ** )\n                loss = self.alpha * student_loss + ( - self.alpha) * distilled_loss\n\n            trainable_vars = self.student.trainable_variables\n            gradients = tape.gradient(loss, trainable_vars)\n            self.optimizer.apply_gradients((gradients, trainable_vars))\n            self.compiled_metrics.update_state(y, student_pred)\n\n             student_loss, distilled_loss\n\n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n\n        results = {m.name : m.result()  m  self.metrics}\n        results.update(\n            {: student_loss, : distilled_loss}\n        )\n\n         results\n\n     ():\n        ()\n        x, y = data\n        is_batch_small = tf.less(tf.shape(x)[],  val_batch_size)\n        is_batch_small = tf.reduce_all(is_batch_small)\n\n         ():\n             tf.constant(), tf.constant()\n\n         ():\n            y_pred = self.student(x, training = )\n            teacher_pred = self.teacher(x, training = )\n\n            student_loss = self.loss(y, y_pred)\n\n            distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(y_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                ) * (self.temp ** )\n\n            self.compiled_metrics.update_state(y, y_pred)\n             student_loss, distilled_loss\n\n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n\n        results = {m.name : m.result()  m  self.metrics}\n        results.update(\n            {: student_loss, : distilled_loss}\n        )\n\n         results\n\n\n</code></pre>\n<p>While I didn't get a super high rank, I did learn a lot and used a plethora of techniques! I wrote most things from scratch, to learn things myself, and referred to other's work only to get a better understanding of what they did and improve my own code as a result.</p>\n<p><em>Validation</em><br>\nQuite standard - kept a holdout set and computer Normalized Levenshtein distance on it after every epoch.<br>\nNo CV due to low computational resources.</p>\n<h1>Improvements</h1>\n<p>I am quite disappointed that the things I missed out were very trivial - such as not training for more than 300 epochs (I could only manage 100 to try enough models and variations out) and that keeping some missing hands frames was important.</p>\n<p>I also didn't use enough augmentations which was important, even though I used all the ones used by top submissions except cutout.</p>\n<p>I also read about Squeezeformer, but got demotivated because I saw high scoring public submissions without Squeezeformer :(</p>\n<p>To further improve, I'll correct all the points above.</p>\n<h1>Final Note</h1>\n<p>I am quite happy to have participated in this fun and awesome contest. I really learnt a lot, and experimented a lot. I am certain that what I did here will help me in future competitions, research and work.</p>\n<p>If you could give me any suggestions/feedback on my code above, please do so. It will help me improve and learn, and I will be very grateful for your help!</p>",
      "rawMarkdown": "Hi Everyone! I'm just posting my solution here, in case it could help someone! I used a lot of techniques, but didn't have the computing power to push the boundaries required for a very high rank. Nevertheless, I learnt a lot of techniques, and am writing this solution here to help provide useful code!\n\n#Competition Context\nThe context of this solution is the Google - American Sign Language Fingerspelling Recognition Competition:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/overview\n\nThe data of this competition can be found here:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/data\n\n#My Approach\nMy approach was identical to most people, I referred to the really awesome notebook by Rohit Ingilela here: https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\n\nI made several additions and changes, namely by adding Quantization Aware Training, AWP, and Knowledge distillation (not necessarily in the same model). I tested a lot of parameters as well and made a lot of models.\n\nMy idea was to use Quantization to increase model size. I found that after 24 Million parameters, even though the model size was small enough, the performance became too poor to run within the 5 hours allocated for testing. I found that KD didn't improve my performance too much, so I didn't explore it a lot. I tried AWP only at the end, and even though it improved performance by around 0.003, I didn't get to test it a lot.\n\nNevertheless, here is some code for each idea:\n\n**Quantization:**\n````python\n!pip install -q tensorflow-model-optimization\nimport tensorflow_model_optimization as tfmot\nfrom tensorflow.keras import layers\n\n# For any layer you want to quantize, wrap it as follows:\nx = tfmot.quantization.keras.quantize_annotate_layer(layers.Dense(channels, activation = activation, kernel_regularizer = regularizer))(inputs)\n\n\n\n# Train and compile as usual, but when converting to tflite, add the following:\n\ntflitemodel_base = TFLiteModel(model)\ntflitemodel_base(frames)[\"outputs\"].shape\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\n\n# Add the line below\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT] \n\n# If you want to convert to float16 and not 8-bit\n# keras_model_converter.target_spec.supported_types = [tf.float16]\n\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]#, tf.lite.OpsSet.SELECT_TF_OPS]\ntflite_model = keras_model_converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n````\n\n**Adversarial Regularization in Training**\n````python\n! pip install -q neural-structured-learning\n\nimport neural_structured_learning as nsl\n\n# Define the model as usual\nadv_config = nsl.configs.make_adv_reg_config(\n    multiplier = 0.2,\n    adv_step_size = 0.2,\n    adv_grad_norm = 'infinity'\n)\n\n# The input_1 means that is what your first layer takes as input as a name. You can check this in model summary. For me it's input_1\ndef convert_to_dictionaries(image, label):\n    return {'input_1': image, 'label': label}\n\ntrain_dataset = train_dataset.map(convert_to_dictionaries)\nval_dataset = val_dataset.map(convert_to_dictionaries)\n\n# This strategy scope is for TPU or any other tensorflow strategy (like distributing strategies). \n# You can remove it\n\nwith strategy.scope():\n    model = nsl.keras.AdversarialRegularization(\n        model,\n        label_keys = ['label'],\n        adv_config = adv_config\n    )\n    \n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold = 4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period = 5)\n\n    model.compile(loss = loss, optimizer = optimizer)\n\n\n# Train as usual with model.fit\n\n````\n\n**Knowledge Distillation**\n\n````python\n\nclass Distiller(keras.Model):\n    def __init__(self, student, teacher):\n        super().__init__()\n        self.student = student\n        self.teacher = teacher\n    \n    def compile(self, optimizer, alpha = 0.1, temperature = 10): \n        super().compile(optimizer=optimizer)\n        self.loss = CTCLoss\n        self.distiller_loss_fn = keras.losses.KLDivergence()\n        self.alpha = alpha\n        self.temp = temperature\n\n    def train_step(self, data):\n        x, y = data\n        # Need to add the following because tensorflow is weird sometimes :)\n        is_batch_small = tf.math.less(tf.shape(x)[0], train_batch_size)\n        is_batch_small = tf.math.reduce_all(is_batch_small)\n        \n        def pass_batch():\n            return tf.constant(0.0), tf.constant(0.0)\n        \n        def normal():\n            teacher_pred = self.teacher(x, training = False)\n            with tf.GradientTape() as tape:\n                student_pred = self.student(x, training = True)\n                student_loss = self.loss(y, student_pred)\n\n                distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(student_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                    ) * (self.temp ** 2)\n                loss = self.alpha * student_loss + (1 - self.alpha) * distilled_loss\n\n            trainable_vars = self.student.trainable_variables\n            gradients = tape.gradient(loss, trainable_vars)\n            self.optimizer.apply_gradients(zip(gradients, trainable_vars))\n            self.compiled_metrics.update_state(y, student_pred)\n\n            return student_loss, distilled_loss\n        \n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n        \n        results = {m.name : m.result() for m in self.metrics}\n        results.update(\n            {\"student_loss\": student_loss, \"distillation_loss\": distilled_loss}\n        )\n        \n        return results\n    \n    def test_step(self, data):\n        print(\"Val Step\")\n        x, y = data\n        is_batch_small = tf.less(tf.shape(x)[0],  val_batch_size)\n        is_batch_small = tf.reduce_all(is_batch_small)\n\n        def pass_batch():\n            return tf.constant(0.0), tf.constant(0.0)\n\n        def normal():\n            y_pred = self.student(x, training = False)\n            teacher_pred = self.teacher(x, training = False)\n            \n            student_loss = self.loss(y, y_pred)\n\n            distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(y_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                ) * (self.temp ** 2)\n            \n            self.compiled_metrics.update_state(y, y_pred)\n            return student_loss, distilled_loss\n        \n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n        \n        results = {m.name : m.result() for m in self.metrics}\n        results.update(\n            {\"student_loss\": student_loss, \"distillation_loss\": distilled_loss}\n        )\n        \n        return results\n\n# Initialize student and teacher models. Then initialize distiller. Use model.fit on the distiller\n````\n\nWhile I didn't get a super high rank, I did learn a lot and used a plethora of techniques! I wrote most things from scratch, to learn things myself, and referred to other's work only to get a better understanding of what they did and improve my own code as a result.\n\n*Validation*\nQuite standard - kept a holdout set and computer Normalized Levenshtein distance on it after every epoch.\nNo CV due to low computational resources.\n\n# Improvements\n\nI am quite disappointed that the things I missed out were very trivial - such as not training for more than 300 epochs (I could only manage 100 to try enough models and variations out) and that keeping some missing hands frames was important.\n\nI also didn't use enough augmentations which was important, even though I used all the ones used by top submissions except cutout.\n\nI also read about Squeezeformer, but got demotivated because I saw high scoring public submissions without Squeezeformer :(\n\nTo further improve, I'll correct all the points above.\n\n# Final Note\n\nI am quite happy to have participated in this fun and awesome contest. I really learnt a lot, and experimented a lot. I am certain that what I did here will help me in future competitions, research and work.\n\nIf you could give me any suggestions/feedback on my code above, please do so. It will help me improve and learn, and I will be very grateful for your help!",
      "votes": 1
    },
    {
      "id": 2426083,
      "postDate": "2023-09-06T11:53:18.583Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2427344,
          "postDate": "2023-09-07T07:08:17.167Z",
          "content": "<p>Thank you so much!</p>",
          "rawMarkdown": "Thank you so much!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2426083,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:53:18.583000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2427344,
          "author_name": "Aaryam Sharma",
          "author_url": "",
          "post_date": "2023-09-07T07:08:17.167000",
          "content": "<p>Thank you so much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2424785": "Hi Everyone! I'm just posting my solution here, in case it could help someone! I used a lot of techniques, but didn't have the computing power to push the boundaries required for a very high rank. Nevertheless, I learnt a lot of techniques, and am writing this solution here to help provide useful code!\n\n#Competition Context\nThe context of this solution is the Google - American Sign Language Fingerspelling Recognition Competition:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/overview\n\nThe data of this competition can be found here:\nhttps://www.kaggle.com/competitions/asl-fingerspelling/data\n\n#My Approach\nMy approach was identical to most people, I referred to the really awesome notebook by Rohit Ingilela here: https://www.kaggle.com/code/irohith/aslfr-ctc-based-on-prev-comp-1st-place\n\nI made several additions and changes, namely by adding Quantization Aware Training, AWP, and Knowledge distillation (not necessarily in the same model). I tested a lot of parameters as well and made a lot of models.\n\nMy idea was to use Quantization to increase model size. I found that after 24 Million parameters, even though the model size was small enough, the performance became too poor to run within the 5 hours allocated for testing. I found that KD didn't improve my performance too much, so I didn't explore it a lot. I tried AWP only at the end, and even though it improved performance by around 0.003, I didn't get to test it a lot.\n\nNevertheless, here is some code for each idea:\n\n**Quantization:**\n````python\n!pip install -q tensorflow-model-optimization\nimport tensorflow_model_optimization as tfmot\nfrom tensorflow.keras import layers\n\n# For any layer you want to quantize, wrap it as follows:\nx = tfmot.quantization.keras.quantize_annotate_layer(layers.Dense(channels, activation = activation, kernel_regularizer = regularizer))(inputs)\n\n\n\n# Train and compile as usual, but when converting to tflite, add the following:\n\ntflitemodel_base = TFLiteModel(model)\ntflitemodel_base(frames)[\"outputs\"].shape\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflitemodel_base)\n\n# Add the line below\nkeras_model_converter.optimizations = [tf.lite.Optimize.DEFAULT] \n\n# If you want to convert to float16 and not 8-bit\n# keras_model_converter.target_spec.supported_types = [tf.float16]\n\nkeras_model_converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS]#, tf.lite.OpsSet.SELECT_TF_OPS]\ntflite_model = keras_model_converter.convert()\nwith open('model.tflite', 'wb') as f:\n    f.write(tflite_model)\n````\n\n**Adversarial Regularization in Training**\n````python\n! pip install -q neural-structured-learning\n\nimport neural_structured_learning as nsl\n\n# Define the model as usual\nadv_config = nsl.configs.make_adv_reg_config(\n    multiplier = 0.2,\n    adv_step_size = 0.2,\n    adv_grad_norm = 'infinity'\n)\n\n# The input_1 means that is what your first layer takes as input as a name. You can check this in model summary. For me it's input_1\ndef convert_to_dictionaries(image, label):\n    return {'input_1': image, 'label': label}\n\ntrain_dataset = train_dataset.map(convert_to_dictionaries)\nval_dataset = val_dataset.map(convert_to_dictionaries)\n\n# This strategy scope is for TPU or any other tensorflow strategy (like distributing strategies). \n# You can remove it\n\nwith strategy.scope():\n    model = nsl.keras.AdversarialRegularization(\n        model,\n        label_keys = ['label'],\n        adv_config = adv_config\n    )\n    \n    loss = CTCLoss\n\n    optimizer = tfa.optimizers.RectifiedAdam(sma_threshold = 4)\n    optimizer = tfa.optimizers.Lookahead(optimizer, sync_period = 5)\n\n    model.compile(loss = loss, optimizer = optimizer)\n\n\n# Train as usual with model.fit\n\n````\n\n**Knowledge Distillation**\n\n````python\n\nclass Distiller(keras.Model):\n    def __init__(self, student, teacher):\n        super().__init__()\n        self.student = student\n        self.teacher = teacher\n    \n    def compile(self, optimizer, alpha = 0.1, temperature = 10): \n        super().compile(optimizer=optimizer)\n        self.loss = CTCLoss\n        self.distiller_loss_fn = keras.losses.KLDivergence()\n        self.alpha = alpha\n        self.temp = temperature\n\n    def train_step(self, data):\n        x, y = data\n        # Need to add the following because tensorflow is weird sometimes :)\n        is_batch_small = tf.math.less(tf.shape(x)[0], train_batch_size)\n        is_batch_small = tf.math.reduce_all(is_batch_small)\n        \n        def pass_batch():\n            return tf.constant(0.0), tf.constant(0.0)\n        \n        def normal():\n            teacher_pred = self.teacher(x, training = False)\n            with tf.GradientTape() as tape:\n                student_pred = self.student(x, training = True)\n                student_loss = self.loss(y, student_pred)\n\n                distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(student_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                    ) * (self.temp ** 2)\n                loss = self.alpha * student_loss + (1 - self.alpha) * distilled_loss\n\n            trainable_vars = self.student.trainable_variables\n            gradients = tape.gradient(loss, trainable_vars)\n            self.optimizer.apply_gradients(zip(gradients, trainable_vars))\n            self.compiled_metrics.update_state(y, student_pred)\n\n            return student_loss, distilled_loss\n        \n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n        \n        results = {m.name : m.result() for m in self.metrics}\n        results.update(\n            {\"student_loss\": student_loss, \"distillation_loss\": distilled_loss}\n        )\n        \n        return results\n    \n    def test_step(self, data):\n        print(\"Val Step\")\n        x, y = data\n        is_batch_small = tf.less(tf.shape(x)[0],  val_batch_size)\n        is_batch_small = tf.reduce_all(is_batch_small)\n\n        def pass_batch():\n            return tf.constant(0.0), tf.constant(0.0)\n\n        def normal():\n            y_pred = self.student(x, training = False)\n            teacher_pred = self.teacher(x, training = False)\n            \n            student_loss = self.loss(y, y_pred)\n\n            distilled_loss = self.distiller_loss_fn(\n                    tf.nn.softmax(y_pred / self.temp), \n                    tf.nn.softmax(teacher_pred / self.temp)\n                ) * (self.temp ** 2)\n            \n            self.compiled_metrics.update_state(y, y_pred)\n            return student_loss, distilled_loss\n        \n        student_loss, distilled_loss = tf.cond(is_batch_small, pass_batch, normal)\n\n        \n        results = {m.name : m.result() for m in self.metrics}\n        results.update(\n            {\"student_loss\": student_loss, \"distillation_loss\": distilled_loss}\n        )\n        \n        return results\n\n# Initialize student and teacher models. Then initialize distiller. Use model.fit on the distiller\n````\n\nWhile I didn't get a super high rank, I did learn a lot and used a plethora of techniques! I wrote most things from scratch, to learn things myself, and referred to other's work only to get a better understanding of what they did and improve my own code as a result.\n\n*Validation*\nQuite standard - kept a holdout set and computer Normalized Levenshtein distance on it after every epoch.\nNo CV due to low computational resources.\n\n# Improvements\n\nI am quite disappointed that the things I missed out were very trivial - such as not training for more than 300 epochs (I could only manage 100 to try enough models and variations out) and that keeping some missing hands frames was important.\n\nI also didn't use enough augmentations which was important, even though I used all the ones used by top submissions except cutout.\n\nI also read about Squeezeformer, but got demotivated because I saw high scoring public submissions without Squeezeformer :(\n\nTo further improve, I'll correct all the points above.\n\n# Final Note\n\nI am quite happy to have participated in this fun and awesome contest. I really learnt a lot, and experimented a lot. I am certain that what I did here will help me in future competitions, research and work.\n\nIf you could give me any suggestions/feedback on my code above, please do so. It will help me improve and learn, and I will be very grateful for your help!",
    "2426083": ""
  }
}