{
  "id": 436698,
  "title": "[429th Solution] Google - American Sign Language Fingerspelling Recognition - own, native)",
  "url": "/competitions/asl-fingerspelling/discussion/436698",
  "author_name": "Zaakcii Ru",
  "post_date": "2023-09-03T16:25:48.654000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello. I had the opportunity to describe my decision and the path by which I came to it) I decided not to deny myself the pleasure of taking this opportunity) Unfortunately, I no longer remember which options I chose for the final submission (I got confused in the list of submitted materials) and it probably doesn't matter considering, that this is a description of the solution that took 429th place. I did a lot of experiments, redid the n-th number of layers, tried to assemble something worthwhile out of them - it helped me become 429) Keras - interesting. Thanks to the organizers for a good time.</p>\n<h2><strong>Context</strong></h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></p>\n<h2><strong>Overview of the Approach</strong></h2>\n<p>In the beginning, I worked with two publicly available, well-known approaches (thanks to their authors). I was able to achieve cv 0.9+ with an encoder-decoder transformer by adding some LSTM or DepthwiseConv1D+ConvLSTM1D layers before the classifier, but all my attempts to convert and complete the tflite model ended in failure.</p>\n<p>The result of this work was a search and an attempt to create a custom LSTM layer. Somewhere on GitHub, I found a variant of a simple LSTM layer (AI Summer link below). Converting it to keras was the start. The created layer did not lead to the desired result, but laid the foundation for an interesting search process. Everything that was - corresponded in Keras. No less amusing was the combination of everything that was, with new finds.</p>\n<p>My experiments are based on the work of Rohith Ingilela's and MARK WIJKHUIZEN Loading, data processing, model output, minus my little additions are their activities.</p>\n<h2><strong>Details of the submission</strong></h2>\n<p>Once again, I note that I no longer remember which variants were selected by me in the end. I do not exclude that it could even be variants based on the above-mentioned works - pardom. Here are a couple of options for the layers that I used. As I understand it, something similar is often used in work on the classification of poses.</p>\n<p>Simple LSTM layer:</p>\n<pre><code> :\n    def :\n        super(Cell_AI_SUMMER, self).\n        self.input_length = input_length\n\n        self.linear_forget_w1 = tf.keras.layers.\n        self.linear_forget_r1 = tf.keras.layers.\n        self.sigmoid_forget = tf.keras.activations.sigmoid\n\n        self.linear_gate_w2 = tf.keras.layers.\n        self.linear_gate_r2 = tf.keras.layers.\n        self.sigmoid_gate = tf.keras.activations.sigmoid\n\n        self.linear_gate_w3 = tf.keras.layers.\n        self.linear_gate_r3 = tf.keras.layers.\n        self.activation_gate = tf.keras.activations.tanh\n\n        self.linear_gate_w4 = tf.keras.layers.\n        self.linear_gate_r4 = tf.keras.layers.\n        self.sigmoid_hidden_out = tf.keras.activations.sigmoid\n\n        self.drop_forget = tf.keras.layers.\n        self.activation_final = tf.keras.activations.tanh #gelu #tf.keras.activations.tanh\n\n    def forget(self, x, h):\n        x = self.linear\n        x = self.drop\n        h = self.linear\n        h = self.drop\n        return self.sigmoid\n\n    def input:\n        x_temp = self.linear\n        x_temp = self.drop\n        h_temp = self.linear\n        h_temp = self.drop\n        i = self.sigmoid\n        return i\n\n    def cell:\n        x = self.linear\n        h = self.linear\n        x = self.drop\n        h = self.drop\n\n        k = self.activation\n        g = ki\n\n        c = fc_prev\n        c_next = g + c\n        return c_next\n\n    def out:\n        x = self.linear\n        h = self.linear\n        return self.sigmoid\n\n    def call(self, x, tuple_in ):\n        (h, c_prev) = tuple_in\n        i = self.input\n        f = self.forget(x, h)\n        h = tf.cast(h, dtype=tf.float32)\n        c_prev = tf.cast(c_prev, dtype=tf.float32)\n        c_next = self.cell\n        o = self.out\n        h_next = oself.activation\n\n        return h_next, c_next\n</code></pre>\n<p><code>zero_core = 128</code></p>\n<pre><code> (tf.keras.Model):\n     ():\n        (Sequence2, self).__init__(name=)\n        self.input_length = input_length\n        self.rnn1 = Cell_AI_SUMMER(, input_length)\n        self.rnn2 = Cell_AI_SUMMER(, input_length)\n\n        self.add = tf.keras.layers.Add()\n        self.linear = tf.keras.activations.gelu\n\n     ():\n        outputs = []\n        h_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        c_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        h_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        c_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n\n        h_t, c_t = self.rnn1(, (h_t, c_t))\n        h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n        output = self.linear(h_t2)\n        outputs += [output]\n\n         i  (future):\n            h_t, c_t = self.rnn1(, (h_t, c_t))\n            h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n            output = self.linear(h_t2)\n            outputs += [output]\n\n        outputs = self.add(outputs)\n        outputs = self.add([outputs, ])\n\n         outputs\n</code></pre>\n<p>Simple Custom Layer:</p>\n<pre><code> :\n    def :\n        super(Convs, self).\n        self.input_length = input_length\n        self.ch1, self.ch2, self.ch3 = , ,  #, , \n        self.conv_0 = tf.keras.layers.\n        self.conv_1 = tf.keras.layers.\n        self.conv_2 = tf.keras.layers.\n        self.conv_3 = tf.keras.layers.\n        self.conv_4 = tf.keras.layers.\n        self.conv_5 = tf.keras.layers.\n        self.conv_6 = tf.keras.layers.\n        self.conv_7 = tf.keras.layers.\n        self.batch1 = tf.keras.layers.\n        self.batch2 = tf.keras.layers.\n        self.batch3 = tf.keras.layers.\n        self.batch4 = tf.keras.layers.\n        self.relu =  tf.keras.layers.\n        self.add = tf.keras.layers.\n        self.drop = tf.keras.layers.\n        self.max2 = tf.keras.layers.\n        self.uns = tf.keras.layers.\n        self.ave = tf.keras.layers.\n\n    def conv0(self, x):\n        x = self.conv\n        x = self.batch1(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv1(self, x):\n        x = self.conv\n        x = self.batch2(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.uns(x)\n\n        return x\n\n    def conv2(self, x):\n        x = self.conv\n        x = self.batch3(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv3(self, x):\n        x = self.conv\n        x = self.batch4(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n\n        return x\n\n    def call(self, input):\n        out = self.conv0(input)\n        out = self.conv1(out)\n        out = self.conv2(out)\n        out = self.conv3(out)\n        out = self.add()\n        return out\n</code></pre>\n<p>Model example (this is just one of the smaller options):</p>\n<pre><code>def get:\n    inp = tf.keras.\n\n    x = tf.keras.layers.(inp)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    pe = tf.cast(positional, dtype=x.dtype)\n    x = x + pe\n    x = (x)\n    x = (x)\n    x1 = (x)\n    x = (x)\n    x = tf.keras.layers.()\n    x = (x)\n    x = (x)\n\n    x = tf.keras.layers.(x)\n\n    model = tf.keras.\n\n    return model\n</code></pre>\n<h2><strong>Which led to an improvement.</strong></h2>\n<p>Augmentations at the initial stage. Subsequently, I abandoned them.<br>\nAlternating UpSampling1D and MaxPooling1D in custom layers.<br>\nReplacing Add with Average.<br>\nLeaving in the model a certain amount of Conv1DBlock, TransformerBlock from the 1st place solution of the last competition, by HOYSO48<br>\nWhat I didn't think to do)</p>\n<h2><strong>Sources</strong></h2>\n<p>My sincere respect to the authors of the following works:<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a><br>\n<a href=\"https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py\" target=\"_blank\">https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py</a><br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406978</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std</a></p>\n<p>Something like this. Thanks again for an interesting challenge. Have a good day.</p>",
  "messages": [
    {
      "id": 2421980,
      "postDate": "2023-09-03T16:25:48.653Z",
      "content": "<p>Hello. I had the opportunity to describe my decision and the path by which I came to it) I decided not to deny myself the pleasure of taking this opportunity) Unfortunately, I no longer remember which options I chose for the final submission (I got confused in the list of submitted materials) and it probably doesn't matter considering, that this is a description of the solution that took 429th place. I did a lot of experiments, redid the n-th number of layers, tried to assemble something worthwhile out of them - it helped me become 429) Keras - interesting. Thanks to the organizers for a good time.</p>\n<h2><strong>Context</strong></h2>\n<p>Business context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a><br>\nData context: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></p>\n<h2><strong>Overview of the Approach</strong></h2>\n<p>In the beginning, I worked with two publicly available, well-known approaches (thanks to their authors). I was able to achieve cv 0.9+ with an encoder-decoder transformer by adding some LSTM or DepthwiseConv1D+ConvLSTM1D layers before the classifier, but all my attempts to convert and complete the tflite model ended in failure.</p>\n<p>The result of this work was a search and an attempt to create a custom LSTM layer. Somewhere on GitHub, I found a variant of a simple LSTM layer (AI Summer link below). Converting it to keras was the start. The created layer did not lead to the desired result, but laid the foundation for an interesting search process. Everything that was - corresponded in Keras. No less amusing was the combination of everything that was, with new finds.</p>\n<p>My experiments are based on the work of Rohith Ingilela's and MARK WIJKHUIZEN Loading, data processing, model output, minus my little additions are their activities.</p>\n<h2><strong>Details of the submission</strong></h2>\n<p>Once again, I note that I no longer remember which variants were selected by me in the end. I do not exclude that it could even be variants based on the above-mentioned works - pardom. Here are a couple of options for the layers that I used. As I understand it, something similar is often used in work on the classification of poses.</p>\n<p>Simple LSTM layer:</p>\n<pre><code> :\n    def :\n        super(Cell_AI_SUMMER, self).\n        self.input_length = input_length\n\n        self.linear_forget_w1 = tf.keras.layers.\n        self.linear_forget_r1 = tf.keras.layers.\n        self.sigmoid_forget = tf.keras.activations.sigmoid\n\n        self.linear_gate_w2 = tf.keras.layers.\n        self.linear_gate_r2 = tf.keras.layers.\n        self.sigmoid_gate = tf.keras.activations.sigmoid\n\n        self.linear_gate_w3 = tf.keras.layers.\n        self.linear_gate_r3 = tf.keras.layers.\n        self.activation_gate = tf.keras.activations.tanh\n\n        self.linear_gate_w4 = tf.keras.layers.\n        self.linear_gate_r4 = tf.keras.layers.\n        self.sigmoid_hidden_out = tf.keras.activations.sigmoid\n\n        self.drop_forget = tf.keras.layers.\n        self.activation_final = tf.keras.activations.tanh #gelu #tf.keras.activations.tanh\n\n    def forget(self, x, h):\n        x = self.linear\n        x = self.drop\n        h = self.linear\n        h = self.drop\n        return self.sigmoid\n\n    def input:\n        x_temp = self.linear\n        x_temp = self.drop\n        h_temp = self.linear\n        h_temp = self.drop\n        i = self.sigmoid\n        return i\n\n    def cell:\n        x = self.linear\n        h = self.linear\n        x = self.drop\n        h = self.drop\n\n        k = self.activation\n        g = ki\n\n        c = fc_prev\n        c_next = g + c\n        return c_next\n\n    def out:\n        x = self.linear\n        h = self.linear\n        return self.sigmoid\n\n    def call(self, x, tuple_in ):\n        (h, c_prev) = tuple_in\n        i = self.input\n        f = self.forget(x, h)\n        h = tf.cast(h, dtype=tf.float32)\n        c_prev = tf.cast(c_prev, dtype=tf.float32)\n        c_next = self.cell\n        o = self.out\n        h_next = oself.activation\n\n        return h_next, c_next\n</code></pre>\n<p><code>zero_core = 128</code></p>\n<pre><code> (tf.keras.Model):\n     ():\n        (Sequence2, self).__init__(name=)\n        self.input_length = input_length\n        self.rnn1 = Cell_AI_SUMMER(, input_length)\n        self.rnn2 = Cell_AI_SUMMER(, input_length)\n\n        self.add = tf.keras.layers.Add()\n        self.linear = tf.keras.activations.gelu\n\n     ():\n        outputs = []\n        h_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        c_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        h_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n        c_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name=)\n\n        h_t, c_t = self.rnn1(, (h_t, c_t))\n        h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n        output = self.linear(h_t2)\n        outputs += [output]\n\n         i  (future):\n            h_t, c_t = self.rnn1(, (h_t, c_t))\n            h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n            output = self.linear(h_t2)\n            outputs += [output]\n\n        outputs = self.add(outputs)\n        outputs = self.add([outputs, ])\n\n         outputs\n</code></pre>\n<p>Simple Custom Layer:</p>\n<pre><code> :\n    def :\n        super(Convs, self).\n        self.input_length = input_length\n        self.ch1, self.ch2, self.ch3 = , ,  #, , \n        self.conv_0 = tf.keras.layers.\n        self.conv_1 = tf.keras.layers.\n        self.conv_2 = tf.keras.layers.\n        self.conv_3 = tf.keras.layers.\n        self.conv_4 = tf.keras.layers.\n        self.conv_5 = tf.keras.layers.\n        self.conv_6 = tf.keras.layers.\n        self.conv_7 = tf.keras.layers.\n        self.batch1 = tf.keras.layers.\n        self.batch2 = tf.keras.layers.\n        self.batch3 = tf.keras.layers.\n        self.batch4 = tf.keras.layers.\n        self.relu =  tf.keras.layers.\n        self.add = tf.keras.layers.\n        self.drop = tf.keras.layers.\n        self.max2 = tf.keras.layers.\n        self.uns = tf.keras.layers.\n        self.ave = tf.keras.layers.\n\n    def conv0(self, x):\n        x = self.conv\n        x = self.batch1(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv1(self, x):\n        x = self.conv\n        x = self.batch2(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.uns(x)\n\n        return x\n\n    def conv2(self, x):\n        x = self.conv\n        x = self.batch3(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv3(self, x):\n        x = self.conv\n        x = self.batch4(x)\n        x = self.relu(x)\n        x = self.conv\n        x = self.drop(x)\n\n        return x\n\n    def call(self, input):\n        out = self.conv0(input)\n        out = self.conv1(out)\n        out = self.conv2(out)\n        out = self.conv3(out)\n        out = self.add()\n        return out\n</code></pre>\n<p>Model example (this is just one of the smaller options):</p>\n<pre><code>def get:\n    inp = tf.keras.\n\n    x = tf.keras.layers.(inp)\n\n    x = tf.keras.layers.(x)\n    x = tf.keras.layers.(x)\n    pe = tf.cast(positional, dtype=x.dtype)\n    x = x + pe\n    x = (x)\n    x = (x)\n    x1 = (x)\n    x = (x)\n    x = tf.keras.layers.()\n    x = (x)\n    x = (x)\n\n    x = tf.keras.layers.(x)\n\n    model = tf.keras.\n\n    return model\n</code></pre>\n<h2><strong>Which led to an improvement.</strong></h2>\n<p>Augmentations at the initial stage. Subsequently, I abandoned them.<br>\nAlternating UpSampling1D and MaxPooling1D in custom layers.<br>\nReplacing Add with Average.<br>\nLeaving in the model a certain amount of Conv1DBlock, TransformerBlock from the 1st place solution of the last competition, by HOYSO48<br>\nWhat I didn't think to do)</p>\n<h2><strong>Sources</strong></h2>\n<p>My sincere respect to the authors of the following works:<br>\n<a href=\"https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference\" target=\"_blank\">https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference</a><br>\n<a href=\"https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py\" target=\"_blank\">https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py</a><br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406978\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/406978</a><br>\n<a href=\"https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std\" target=\"_blank\">https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std</a></p>\n<p>Something like this. Thanks again for an interesting challenge. Have a good day.</p>",
      "rawMarkdown": "Hello. I had the opportunity to describe my decision and the path by which I came to it) I decided not to deny myself the pleasure of taking this opportunity) Unfortunately, I no longer remember which options I chose for the final submission (I got confused in the list of submitted materials) and it probably doesn't matter considering, that this is a description of the solution that took 429th place. I did a lot of experiments, redid the n-th number of layers, tried to assemble something worthwhile out of them - it helped me become 429) Keras - interesting. Thanks to the organizers for a good time.\n\n##**Context**\n\nBusiness context: https://www.kaggle.com/competitions/asl-fingerspelling/overview\nData context: https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n##**Overview of the Approach**\n\nIn the beginning, I worked with two publicly available, well-known approaches (thanks to their authors). I was able to achieve cv 0.9+ with an encoder-decoder transformer by adding some LSTM or DepthwiseConv1D+ConvLSTM1D layers before the classifier, but all my attempts to convert and complete the tflite model ended in failure.\n\nThe result of this work was a search and an attempt to create a custom LSTM layer. Somewhere on GitHub, I found a variant of a simple LSTM layer (AI Summer link below). Converting it to keras was the start. The created layer did not lead to the desired result, but laid the foundation for an interesting search process. Everything that was - corresponded in Keras. No less amusing was the combination of everything that was, with new finds.\n\nMy experiments are based on the work of Rohith Ingilela's and MARK WIJKHUIZEN Loading, data processing, model output, minus my little additions are their activities.\n\n##**Details of the submission**\n\nOnce again, I note that I no longer remember which variants were selected by me in the end. I do not exclude that it could even be variants based on the above-mentioned works - pardom. Here are a couple of options for the layers that I used. As I understand it, something similar is often used in work on the classification of poses.\n\nSimple LSTM layer:\n\n```\nclass Cell_AI_SUMMER(tf.keras.Model):\n    def __init__(self, numer, input_length=276):\n        super(Cell_AI_SUMMER, self).__init__(name=f'cell_{numer}')\n        self.input_length = input_length\n\n        self.linear_forget_w1 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w1')\n        self.linear_forget_r1 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r1')\n        self.sigmoid_forget = tf.keras.activations.sigmoid\n\n        self.linear_gate_w2 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w2')\n        self.linear_gate_r2 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r2')\n        self.sigmoid_gate = tf.keras.activations.sigmoid\n\n        self.linear_gate_w3 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w3')\n        self.linear_gate_r3 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r3')\n        self.activation_gate = tf.keras.activations.tanh\n\n        self.linear_gate_w4 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w4')\n        self.linear_gate_r4 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r4')\n        self.sigmoid_hidden_out = tf.keras.activations.sigmoid\n\n        self.drop_forget = tf.keras.layers.Dropout(.1)\n        self.activation_final = tf.keras.activations.tanh #gelu #tf.keras.activations.tanh\n\n    def forget(self, x, h):\n        x = self.linear_forget_w1(x)\n        x = self.drop_forget(x)\n        h = self.linear_forget_r1(h)\n        h = self.drop_forget(h)\n        return self.sigmoid_forget(x + h)\n\n    def input_gate(self, x, h):\n        x_temp = self.linear_gate_w2(x)\n        x_temp = self.drop_forget(x_temp)\n        h_temp = self.linear_gate_r2(h)\n        h_temp = self.drop_forget(h_temp)\n        i = self.sigmoid_gate(x_temp + h_temp)\n        return i\n\n    def cell_memory_gate(self, i, f, x, h, c_prev):\n        x = self.linear_gate_w3(x)\n        h = self.linear_gate_r3(h)\n        x = self.drop_forget(x)\n        h = self.drop_forget(h)\n\n        k = self.activation_gate(x + h)\n        g = k * i\n\n        c = f * c_prev\n        c_next = g + c\n        return c_next\n\n    def out_gate(self, x, h):\n        x = self.linear_gate_w4(x)\n        h = self.linear_gate_r4(h)\n        return self.sigmoid_hidden_out(x + h)\n\n    def call(self, x, tuple_in ):\n        (h, c_prev) = tuple_in\n        i = self.input_gate(x, h)\n        f = self.forget(x, h)\n        h = tf.cast(h, dtype=tf.float32)\n        c_prev = tf.cast(c_prev, dtype=tf.float32)\n        c_next = self.cell_memory_gate(i, f, x, h, c_prev)\n        o = self.out_gate(x, h)\n        h_next = o * self.activation_final(c_next)\n\n        return h_next, c_next\n```\n\n`zero_core = 128`\n\n```\nclass Sequence2(tf.keras.Model):\n    def __init__(self, numer, input_length=256):\n        super(Sequence2, self).__init__(name=f'seq_{numer}')\n        self.input_length = input_length\n        self.rnn1 = Cell_AI_SUMMER(f'{numer}_0', input_length)\n        self.rnn2 = Cell_AI_SUMMER(f'{numer}_1', input_length)\n\n        self.add = tf.keras.layers.Add()\n        self.linear = tf.keras.activations.gelu\n\n    def call(self, input, future=5):\n        outputs = []\n        h_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ht')\n        c_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ct')\n        h_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ht2')\n        c_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ct2')\n        \n        h_t, c_t = self.rnn1(input, (h_t, c_t))\n        h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n       \n        output = self.linear(h_t2)\n        outputs += [output]\n\n        for i in range(future):\n            h_t, c_t = self.rnn1(input, (h_t, c_t))\n            h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n            output = self.linear(h_t2)\n            outputs += [output]\n\n        outputs = self.add(outputs)\n        outputs = self.add([outputs, input])\n  \n        return outputs\n```\n\nSimple Custom Layer:\n\n```\nclass Convs(tf.keras.Model):\n    def __init__(self, numer, input_length=256):\n        super(Convs, self).__init__(name=f'cnvs_{numer}')\n        self.input_length = input_length\n        self.ch1, self.ch2, self.ch3 = 128, 384, 768 #64, 128, 512\n        self.conv_0 = tf.keras.layers.Conv1D(self.ch1, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_0', input_shape = [None, 128, 222])\n        self.conv_1 = tf.keras.layers.Conv1D(self.ch1, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_1')\n        self.conv_2 = tf.keras.layers.Conv1D(self.ch2, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_2')\n        self.conv_3 = tf.keras.layers.Conv1D(self.ch2, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_3')\n        self.conv_4 = tf.keras.layers.Conv1D(self.ch3, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_4')\n        self.conv_5 = tf.keras.layers.Conv1D(self.ch3, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_5')\n        self.conv_6 = tf.keras.layers.Conv1D(self.input_length, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_6')\n        self.conv_7 = tf.keras.layers.Conv1D(self.input_length, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_7')\n        self.batch1 = tf.keras.layers.BatchNormalization(momentum=0.1)\n        self.batch2 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[128, 128])\n        self.batch3 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[256, 768])\n        self.batch4 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[128, 256])\n        self.relu =  tf.keras.layers.ReLU()\n        self.add = tf.keras.layers.Add()\n        self.drop = tf.keras.layers.Dropout(0.1)\n        self.max2 = tf.keras.layers.MaxPooling1D(pool_size=2)\n        self.uns = tf.keras.layers.UpSampling1D(size=4)\n        self.ave = tf.keras.layers.AveragePooling1D(pool_size=2)\n\n    def conv0(self, x):\n        x = self.conv_0(x)\n        x = self.batch1(x)\n        x = self.relu(x)\n        x = self.conv_1(x)\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv1(self, x):\n        x = self.conv_2(x)\n        x = self.batch2(x)\n        x = self.relu(x)\n        x = self.conv_3(x)\n        x = self.drop(x)\n        x = self.uns(x)\n\n        return x\n\n    def conv2(self, x):\n        x = self.conv_4(x)\n        x = self.batch3(x)\n        x = self.relu(x)\n        x = self.conv_5(x)\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv3(self, x):\n        x = self.conv_6(x)\n        x = self.batch4(x)\n        x = self.relu(x)\n        x = self.conv_7(x)\n        x = self.drop(x)\n\n        return x\n\n    def call(self, input):\n        out = self.conv0(input)\n        out = self.conv1(out)\n        out = self.conv2(out)\n        out = self.conv3(out)\n        out = self.add([out, input])\n        return out\n```\n\nModel example (this is just one of the smaller options):\n```\ndef get_model(dim = 276, dropout_step=0):\n    inp = tf.keras.Input(INPUT_SHAPE)\n\n    x = tf.keras.layers.Masking(mask_value=0.0)(inp)\n\n    x = tf.keras.layers.Dense(dim, use_bias=False, name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.9, name='bs151')(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[0], dim), dtype=x.dtype)\n    x = x + pe\n    x = Convs3(0, input_length=276)(x)\n    x = Sequence2(0, 276)(x)\n    x1 = SelfAttention1DLayer(similarity=\"linear\",dropout_rate=0.2)(x)\n    x = Conv1DBlock(dim,11,drop_rate=0.2)(x)\n    x = tf.keras.layers.Average()([x,x1])\n    x = Conv1DBlock(dim,11,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(60, activation='linear')(x)\n\n    model = tf.keras.Model(inp, x)\n\n    return model\n```\n\n##**Which led to an improvement.**\n\nAugmentations at the initial stage. Subsequently, I abandoned them.\nAlternating UpSampling1D and MaxPooling1D in custom layers.\nReplacing Add with Average.\nLeaving in the model a certain amount of Conv1DBlock, TransformerBlock from the 1st place solution of the last competition, by HOYSO48\nWhat I didn't think to do)\n\n##**Sources**\nMy sincere respect to the authors of the following works:\n<https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference>\n<https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py>\n<https://www.kaggle.com/competitions/asl-signs/discussion/406978>\n<https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std>\n\nSomething like this. Thanks again for an interesting challenge. Have a good day.",
      "votes": 1
    },
    {
      "id": 2426090,
      "postDate": "2023-09-06T11:55:58.043Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2426090,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T11:55:58.043000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2421980": "Hello. I had the opportunity to describe my decision and the path by which I came to it) I decided not to deny myself the pleasure of taking this opportunity) Unfortunately, I no longer remember which options I chose for the final submission (I got confused in the list of submitted materials) and it probably doesn't matter considering, that this is a description of the solution that took 429th place. I did a lot of experiments, redid the n-th number of layers, tried to assemble something worthwhile out of them - it helped me become 429) Keras - interesting. Thanks to the organizers for a good time.\n\n##**Context**\n\nBusiness context: https://www.kaggle.com/competitions/asl-fingerspelling/overview\nData context: https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n##**Overview of the Approach**\n\nIn the beginning, I worked with two publicly available, well-known approaches (thanks to their authors). I was able to achieve cv 0.9+ with an encoder-decoder transformer by adding some LSTM or DepthwiseConv1D+ConvLSTM1D layers before the classifier, but all my attempts to convert and complete the tflite model ended in failure.\n\nThe result of this work was a search and an attempt to create a custom LSTM layer. Somewhere on GitHub, I found a variant of a simple LSTM layer (AI Summer link below). Converting it to keras was the start. The created layer did not lead to the desired result, but laid the foundation for an interesting search process. Everything that was - corresponded in Keras. No less amusing was the combination of everything that was, with new finds.\n\nMy experiments are based on the work of Rohith Ingilela's and MARK WIJKHUIZEN Loading, data processing, model output, minus my little additions are their activities.\n\n##**Details of the submission**\n\nOnce again, I note that I no longer remember which variants were selected by me in the end. I do not exclude that it could even be variants based on the above-mentioned works - pardom. Here are a couple of options for the layers that I used. As I understand it, something similar is often used in work on the classification of poses.\n\nSimple LSTM layer:\n\n```\nclass Cell_AI_SUMMER(tf.keras.Model):\n    def __init__(self, numer, input_length=276):\n        super(Cell_AI_SUMMER, self).__init__(name=f'cell_{numer}')\n        self.input_length = input_length\n\n        self.linear_forget_w1 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w1')\n        self.linear_forget_r1 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r1')\n        self.sigmoid_forget = tf.keras.activations.sigmoid\n\n        self.linear_gate_w2 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w2')\n        self.linear_gate_r2 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r2')\n        self.sigmoid_gate = tf.keras.activations.sigmoid\n\n        self.linear_gate_w3 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w3')\n        self.linear_gate_r3 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r3')\n        self.activation_gate = tf.keras.activations.tanh\n\n        self.linear_gate_w4 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=True, name=f'{self.name}_w4')\n        self.linear_gate_r4 = tf.keras.layers.Dense(self.input_length, activation=tf.keras.activations.selu, kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_r4')\n        self.sigmoid_hidden_out = tf.keras.activations.sigmoid\n\n        self.drop_forget = tf.keras.layers.Dropout(.1)\n        self.activation_final = tf.keras.activations.tanh #gelu #tf.keras.activations.tanh\n\n    def forget(self, x, h):\n        x = self.linear_forget_w1(x)\n        x = self.drop_forget(x)\n        h = self.linear_forget_r1(h)\n        h = self.drop_forget(h)\n        return self.sigmoid_forget(x + h)\n\n    def input_gate(self, x, h):\n        x_temp = self.linear_gate_w2(x)\n        x_temp = self.drop_forget(x_temp)\n        h_temp = self.linear_gate_r2(h)\n        h_temp = self.drop_forget(h_temp)\n        i = self.sigmoid_gate(x_temp + h_temp)\n        return i\n\n    def cell_memory_gate(self, i, f, x, h, c_prev):\n        x = self.linear_gate_w3(x)\n        h = self.linear_gate_r3(h)\n        x = self.drop_forget(x)\n        h = self.drop_forget(h)\n\n        k = self.activation_gate(x + h)\n        g = k * i\n\n        c = f * c_prev\n        c_next = g + c\n        return c_next\n\n    def out_gate(self, x, h):\n        x = self.linear_gate_w4(x)\n        h = self.linear_gate_r4(h)\n        return self.sigmoid_hidden_out(x + h)\n\n    def call(self, x, tuple_in ):\n        (h, c_prev) = tuple_in\n        i = self.input_gate(x, h)\n        f = self.forget(x, h)\n        h = tf.cast(h, dtype=tf.float32)\n        c_prev = tf.cast(c_prev, dtype=tf.float32)\n        c_next = self.cell_memory_gate(i, f, x, h, c_prev)\n        o = self.out_gate(x, h)\n        h_next = o * self.activation_final(c_next)\n\n        return h_next, c_next\n```\n\n`zero_core = 128`\n\n```\nclass Sequence2(tf.keras.Model):\n    def __init__(self, numer, input_length=256):\n        super(Sequence2, self).__init__(name=f'seq_{numer}')\n        self.input_length = input_length\n        self.rnn1 = Cell_AI_SUMMER(f'{numer}_0', input_length)\n        self.rnn2 = Cell_AI_SUMMER(f'{numer}_1', input_length)\n\n        self.add = tf.keras.layers.Add()\n        self.linear = tf.keras.activations.gelu\n\n    def call(self, input, future=5):\n        outputs = []\n        h_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ht')\n        c_t = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ct')\n        h_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ht2')\n        c_t2 = tf.zeros([zero_core, self.input_length], dtype=tf.double, name='ct2')\n        \n        h_t, c_t = self.rnn1(input, (h_t, c_t))\n        h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n       \n        output = self.linear(h_t2)\n        outputs += [output]\n\n        for i in range(future):\n            h_t, c_t = self.rnn1(input, (h_t, c_t))\n            h_t2, c_t2 = self.rnn2(h_t, (h_t2, c_t2))\n\n            output = self.linear(h_t2)\n            outputs += [output]\n\n        outputs = self.add(outputs)\n        outputs = self.add([outputs, input])\n  \n        return outputs\n```\n\nSimple Custom Layer:\n\n```\nclass Convs(tf.keras.Model):\n    def __init__(self, numer, input_length=256):\n        super(Convs, self).__init__(name=f'cnvs_{numer}')\n        self.input_length = input_length\n        self.ch1, self.ch2, self.ch3 = 128, 384, 768 #64, 128, 512\n        self.conv_0 = tf.keras.layers.Conv1D(self.ch1, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_0', input_shape = [None, 128, 222])\n        self.conv_1 = tf.keras.layers.Conv1D(self.ch1, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_1')\n        self.conv_2 = tf.keras.layers.Conv1D(self.ch2, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_2')\n        self.conv_3 = tf.keras.layers.Conv1D(self.ch2, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_3')\n        self.conv_4 = tf.keras.layers.Conv1D(self.ch3, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_4')\n        self.conv_5 = tf.keras.layers.Conv1D(self.ch3, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_5')\n        self.conv_6 = tf.keras.layers.Conv1D(self.input_length, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_6')\n        self.conv_7 = tf.keras.layers.Conv1D(self.input_length, 1, strides=1, activation='linear', kernel_initializer='random_normal', use_bias=False, name=f'{self.name}_7')\n        self.batch1 = tf.keras.layers.BatchNormalization(momentum=0.1)\n        self.batch2 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[128, 128])\n        self.batch3 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[256, 768])\n        self.batch4 = tf.keras.layers.BatchNormalization(momentum=0.1, input_shape=[128, 256])\n        self.relu =  tf.keras.layers.ReLU()\n        self.add = tf.keras.layers.Add()\n        self.drop = tf.keras.layers.Dropout(0.1)\n        self.max2 = tf.keras.layers.MaxPooling1D(pool_size=2)\n        self.uns = tf.keras.layers.UpSampling1D(size=4)\n        self.ave = tf.keras.layers.AveragePooling1D(pool_size=2)\n\n    def conv0(self, x):\n        x = self.conv_0(x)\n        x = self.batch1(x)\n        x = self.relu(x)\n        x = self.conv_1(x)\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv1(self, x):\n        x = self.conv_2(x)\n        x = self.batch2(x)\n        x = self.relu(x)\n        x = self.conv_3(x)\n        x = self.drop(x)\n        x = self.uns(x)\n\n        return x\n\n    def conv2(self, x):\n        x = self.conv_4(x)\n        x = self.batch3(x)\n        x = self.relu(x)\n        x = self.conv_5(x)\n        x = self.drop(x)\n        x = self.max2(x)\n\n        return x\n\n    def conv3(self, x):\n        x = self.conv_6(x)\n        x = self.batch4(x)\n        x = self.relu(x)\n        x = self.conv_7(x)\n        x = self.drop(x)\n\n        return x\n\n    def call(self, input):\n        out = self.conv0(input)\n        out = self.conv1(out)\n        out = self.conv2(out)\n        out = self.conv3(out)\n        out = self.add([out, input])\n        return out\n```\n\nModel example (this is just one of the smaller options):\n```\ndef get_model(dim = 276, dropout_step=0):\n    inp = tf.keras.Input(INPUT_SHAPE)\n\n    x = tf.keras.layers.Masking(mask_value=0.0)(inp)\n\n    x = tf.keras.layers.Dense(dim, use_bias=False, name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.9, name='bs151')(x)\n    pe = tf.cast(positional_encoding(INPUT_SHAPE[0], dim), dtype=x.dtype)\n    x = x + pe\n    x = Convs3(0, input_length=276)(x)\n    x = Sequence2(0, 276)(x)\n    x1 = SelfAttention1DLayer(similarity=\"linear\",dropout_rate=0.2)(x)\n    x = Conv1DBlock(dim,11,drop_rate=0.2)(x)\n    x = tf.keras.layers.Average()([x,x1])\n    x = Conv1DBlock(dim,11,drop_rate=0.2)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(60, activation='linear')(x)\n\n    model = tf.keras.Model(inp, x)\n\n    return model\n```\n\n##**Which led to an improvement.**\n\nAugmentations at the initial stage. Subsequently, I abandoned them.\nAlternating UpSampling1D and MaxPooling1D in custom layers.\nReplacing Add with Average.\nLeaving in the model a certain amount of Conv1DBlock, TransformerBlock from the 1st place solution of the last competition, by HOYSO48\nWhat I didn't think to do)\n\n##**Sources**\nMy sincere respect to the authors of the following works:\n<https://www.kaggle.com/code/markwijkhuizen/aslfr-transformer-training-inference>\n<https://github.com/The-AI-Summer/RNN_tutorial/blob/master/cutom_LSTM.py>\n<https://www.kaggle.com/competitions/asl-signs/discussion/406978>\n<https://www.kaggle.com/code/irohith/aslfr-preprocess-dataset-tfrecords-mean-std>\n\nSomething like this. Thanks again for an interesting challenge. Have a good day.",
    "2426090": ""
  }
}