{
  "id": 437989,
  "title": "25th Place Solution (time warpping + relative positional encoding + intermediate CTC)",
  "url": "/competitions/asl-fingerspelling/discussion/437989",
  "author_name": "Tong Zou",
  "post_date": "2023-09-09T01:38:13.326000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Google for hosting this competition. This is my first Kaggle competition experence and I have learnt a lot from the previous ASL competition and other participants. I start my model from <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution from the last competition and added some components:</p>\n<p><strong>Data Augmentation</strong><br>\nApplied random time warpping  (LB +0.01)</p>\n<pre><code>def temporal:\n    n = tf.shape(x)\n     n &lt; :\n      return x\n    w = n \n    t1 = tf.clip, , n-)\n    t2 = tf.clip + t1, , n-)\n    new_x1 = interp1d\n    new_x2 = interp1d\n    return tf.concat(, axis=)\n\ndef augment:\n     tf.random.uniform()&lt;  always:\n        x = resample(x, (,))\n     tf.random.uniform()&lt;  always:\n        x = flip\n     max_len is not None:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = spatial\n     tf.random.uniform()&lt;  always:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = spatial\n    return x\n</code></pre>\n<p><strong>Data Preprocessing</strong><br>\nTruncate <code>phrase</code> to have smaller length than the input</p>\n<pre><code> inp_len &lt;= label_len:\n    offset = tf((), , label_len-inp_len+, dtype=tf.int32)\n     = \n    label_len = inp_len-\n</code></pre>\n<p><strong>Multi-headed attention with relative positional encoding</strong> (LB +0.02)</p>\n<pre><code> :\n    def :\n        super.\n        self.dim = dim\n        self.head_dim = dim \n        self.scale = self.head_dim-\n        self.num_heads = num_heads\n        self.num_nbr = num_nbr\n        self.qkv = tf.keras.layers.\n        self.drop1 = tf.keras.layers.\n        self.proj = tf.keras.layers.\n        self.supports_masking = True\n\n        self.wgt_v = self.add,\n                                     ='glorot_uniform',trainable=True,\n                                     name=+str(tf.keras.backend.get))\n        self.wgt_k = self.add,\n                                     ='glorot_uniform',trainable=True,\n                                     name=+str(tf.keras.backend.get))\n\n    def :\n        idx_mat = tf.reshape(tf.tile(tf.range(length), ), )\n        idx_mat = idx_mat - tf.transpose(idx_mat)\n        return tf.clip + self.num_nbr\n\n    def :\n        xy = tf.matmul(x, y, transpose_b=transpose)\n        x = tf.keras.layers.)(x)\n        mul = tf.matmul(x, z, transpose_b=transpose)\n        mul = tf.keras.layers.)(mul)\n        return xy + mul\n\n    def call(self, inputs, mask=None):\n         mask is not None:\n            mask = mask\n        length = inputs.shape\n        idx_mat = self.\n        rpe_k = tf.gather(self.wgt_k, idx_mat)\n        rpe_v = tf.gather(self.wgt_v, idx_mat)\n\n        qkv = self.qkv(inputs)\n        qkv = tf.keras.layers.)(tf.keras.layers.)(qkv))\n        q, k, v = tf.split(qkv, , axis=-)\n        logits = self.self.scale\n        attn = tf.keras.layers.(logits, mask=mask)\n        attn = self.drop1(attn)\n        x = self.\n        x = tf.keras.layers.)(tf.keras.layers.)(x))\n        x = self.proj(x)\n        return x\n</code></pre>\n<p><strong>Intermediate CTC</strong></p>\n<pre><code>def get_base_model(=384, =384):\n    inp = tf.keras.Input((max_len,CHANNELS), =)\n    x = tf.keras.layers.Masking(=PAD,input_shape=(max_len,CHANNELS))(inp)\n    x = tf.keras.layers.Dense(dim, =,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(=0.95,name='stem_bn')(x)\n\n    ksize = 17\n    drop_rate = 0.2\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    inter = tf.keras.layers.Dense(dim, =None)(x)\n    inter = tf.keras.layers.Dense(NUM_CLASSES+1)(inter)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = tf.keras.layers.Dense(dim,=None,name='top_conv')(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES+1, =)(x)\n    return tf.keras.Model(inp, [inter, x])\n\ndef ctc_loss(args):\n    y_pred, inp_len, label, label_len = args\n    return tf.nn.ctc_loss(label, y_pred, label_len, inp_len,\n                          =, =NUM_CLASSES)\n\ndef get_model(base_model):\n    label = tf.keras.Input(shape=(MAX_LEN_OUTPUT,), =, =)\n    label_len = tf.keras.Input(shape=(), =, =)\n    inp_len = tf.keras.Input(shape=(), =, =)\n    loss1 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[0], inp_len, label, label_len])\n    loss2 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[1], inp_len, label, label_len])\n    loss = 0.3 * loss1 + 0.7 * loss2\n    #loss = 0.0 * loss1 + 1.0 * loss2\n    model = tf.keras.Model(\n        inputs=[base_model.input, inp_len, label, label_len],\n        =loss)\n    return model\n</code></pre>\n<p><strong>Fine tuning</strong> (LB +0.005)<br>\nThe model was first trained with both training data and supplementary data with intermediate CTC (0.3 and 0.7) for 300 epochs. Then the pre-trained model was fine tuned only on the training data without intermediate CTC (0.0 and 1.0)</p>\n<p>I will post the training notebook soon.</p>\n<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a></p>",
  "messages": [
    {
      "id": 2430001,
      "postDate": "2023-09-09T01:38:13.327Z",
      "content": "<p>Thanks to Kaggle and Google for hosting this competition. This is my first Kaggle competition experence and I have learnt a lot from the previous ASL competition and other participants. I start my model from <a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution from the last competition and added some components:</p>\n<p><strong>Data Augmentation</strong><br>\nApplied random time warpping  (LB +0.01)</p>\n<pre><code>def temporal:\n    n = tf.shape(x)\n     n &lt; :\n      return x\n    w = n \n    t1 = tf.clip, , n-)\n    t2 = tf.clip + t1, , n-)\n    new_x1 = interp1d\n    new_x2 = interp1d\n    return tf.concat(, axis=)\n\ndef augment:\n     tf.random.uniform()&lt;  always:\n        x = resample(x, (,))\n     tf.random.uniform()&lt;  always:\n        x = flip\n     max_len is not None:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = spatial\n     tf.random.uniform()&lt;  always:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = temporal\n     tf.random.uniform()&lt;  always:\n        x = spatial\n    return x\n</code></pre>\n<p><strong>Data Preprocessing</strong><br>\nTruncate <code>phrase</code> to have smaller length than the input</p>\n<pre><code> inp_len &lt;= label_len:\n    offset = tf((), , label_len-inp_len+, dtype=tf.int32)\n     = \n    label_len = inp_len-\n</code></pre>\n<p><strong>Multi-headed attention with relative positional encoding</strong> (LB +0.02)</p>\n<pre><code> :\n    def :\n        super.\n        self.dim = dim\n        self.head_dim = dim \n        self.scale = self.head_dim-\n        self.num_heads = num_heads\n        self.num_nbr = num_nbr\n        self.qkv = tf.keras.layers.\n        self.drop1 = tf.keras.layers.\n        self.proj = tf.keras.layers.\n        self.supports_masking = True\n\n        self.wgt_v = self.add,\n                                     ='glorot_uniform',trainable=True,\n                                     name=+str(tf.keras.backend.get))\n        self.wgt_k = self.add,\n                                     ='glorot_uniform',trainable=True,\n                                     name=+str(tf.keras.backend.get))\n\n    def :\n        idx_mat = tf.reshape(tf.tile(tf.range(length), ), )\n        idx_mat = idx_mat - tf.transpose(idx_mat)\n        return tf.clip + self.num_nbr\n\n    def :\n        xy = tf.matmul(x, y, transpose_b=transpose)\n        x = tf.keras.layers.)(x)\n        mul = tf.matmul(x, z, transpose_b=transpose)\n        mul = tf.keras.layers.)(mul)\n        return xy + mul\n\n    def call(self, inputs, mask=None):\n         mask is not None:\n            mask = mask\n        length = inputs.shape\n        idx_mat = self.\n        rpe_k = tf.gather(self.wgt_k, idx_mat)\n        rpe_v = tf.gather(self.wgt_v, idx_mat)\n\n        qkv = self.qkv(inputs)\n        qkv = tf.keras.layers.)(tf.keras.layers.)(qkv))\n        q, k, v = tf.split(qkv, , axis=-)\n        logits = self.self.scale\n        attn = tf.keras.layers.(logits, mask=mask)\n        attn = self.drop1(attn)\n        x = self.\n        x = tf.keras.layers.)(tf.keras.layers.)(x))\n        x = self.proj(x)\n        return x\n</code></pre>\n<p><strong>Intermediate CTC</strong></p>\n<pre><code>def get_base_model(=384, =384):\n    inp = tf.keras.Input((max_len,CHANNELS), =)\n    x = tf.keras.layers.Masking(=PAD,input_shape=(max_len,CHANNELS))(inp)\n    x = tf.keras.layers.Dense(dim, =,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(=0.95,name='stem_bn')(x)\n\n    ksize = 17\n    drop_rate = 0.2\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    inter = tf.keras.layers.Dense(dim, =None)(x)\n    inter = tf.keras.layers.Dense(NUM_CLASSES+1)(inter)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,=drop_rate)(x)\n    x = TransformerBlock(dim,=2)(x)\n\n    x = tf.keras.layers.Dense(dim,=None,name='top_conv')(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES+1, =)(x)\n    return tf.keras.Model(inp, [inter, x])\n\ndef ctc_loss(args):\n    y_pred, inp_len, label, label_len = args\n    return tf.nn.ctc_loss(label, y_pred, label_len, inp_len,\n                          =, =NUM_CLASSES)\n\ndef get_model(base_model):\n    label = tf.keras.Input(shape=(MAX_LEN_OUTPUT,), =, =)\n    label_len = tf.keras.Input(shape=(), =, =)\n    inp_len = tf.keras.Input(shape=(), =, =)\n    loss1 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[0], inp_len, label, label_len])\n    loss2 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[1], inp_len, label, label_len])\n    loss = 0.3 * loss1 + 0.7 * loss2\n    #loss = 0.0 * loss1 + 1.0 * loss2\n    model = tf.keras.Model(\n        inputs=[base_model.input, inp_len, label, label_len],\n        =loss)\n    return model\n</code></pre>\n<p><strong>Fine tuning</strong> (LB +0.005)<br>\nThe model was first trained with both training data and supplementary data with intermediate CTC (0.3 and 0.7) for 300 epochs. Then the pre-trained model was fine tuned only on the training data without intermediate CTC (0.0 and 1.0)</p>\n<p>I will post the training notebook soon.</p>\n<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a>'s solution<br>\n<a href=\"https://www.kaggle.com/code/hoyso48/1st-place-solution-training\" target=\"_blank\">https://www.kaggle.com/code/hoyso48/1st-place-solution-training</a></p>",
      "rawMarkdown": "Thanks to Kaggle and Google for hosting this competition. This is my first Kaggle competition experence and I have learnt a lot from the previous ASL competition and other participants. I start my model from @hoyso48's solution from the last competition and added some components:\n\n**Data Augmentation**\nApplied random time warpping  (LB +0.01)\n\n```\ndef temporal_warp(x, factor=8):\n    n = tf.shape(x)[0]\n    if n < 5:\n      return x\n    w = n // factor\n    t1 = tf.clip_by_value(tf.random.uniform([], w, n-w, dtype=tf.int32), 1, n-2)\n    t2 = tf.clip_by_value(tf.random.uniform([], -w, w+1, dtype=tf.int32) + t1, 1, n-2)\n    new_x1 = interp1d_(x[:t1], t2)\n    new_x2 = interp1d_(x[t1:], n-t2)\n    return tf.concat([new_x1, new_x2], axis=0)\n\ndef augment_fn(x, always=False, max_len=None):\n    if tf.random.uniform(())<0.8 or always:\n        x = resample(x, (0.5,1.5))\n    if tf.random.uniform(())<0.5 or always:\n        x = flip_lr(x)\n    if max_len is not None:\n        x = temporal_crop(x, max_len)\n    if tf.random.uniform(())<0.75 or always:\n        x = spatial_random_affine(x)\n    if tf.random.uniform(())<0.75 or always:\n        x = temporal_warp(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = temporal_mask(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = spatial_mask(x)\n    return x\n```\n\n**Data Preprocessing**\nTruncate `phrase` to have smaller length than the input\n```\nif inp_len <= label_len:\n    offset = tf.random.uniform((), 0, label_len-inp_len+2, dtype=tf.int32)\n    label = label[offset:offset+inp_len-1]\n    label_len = inp_len-1\n```\n\n**Multi-headed attention with relative positional encoding** (LB +0.02)\n\n```\nclass MHSAwithRPE(tf.keras.layers.Layer):\n    def __init__(self, dim=256, num_heads=4, num_nbr=8, dropout=0, **kwargs):\n        super().__init__(**kwargs)\n        self.dim = dim\n        self.head_dim = dim // num_heads\n        self.scale = self.head_dim  ** -0.5\n        self.num_heads = num_heads\n        self.num_nbr = num_nbr\n        self.qkv = tf.keras.layers.Dense(3 * dim, use_bias=False)\n        self.drop1 = tf.keras.layers.Dropout(dropout)\n        self.proj = tf.keras.layers.Dense(dim, use_bias=False)\n        self.supports_masking = True\n\n        self.wgt_v = self.add_weight(shape=(num_nbr*2+1, self.head_dim),\n                                     initializer='glorot_uniform',trainable=True,\n                                     name=\"wgt_v\"+str(tf.keras.backend.get_uid(\"wgt_v\")))\n        self.wgt_k = self.add_weight(shape=(num_nbr*2+1, self.head_dim),\n                                     initializer='glorot_uniform',trainable=True,\n                                     name=\"wgt_k\"+str(tf.keras.backend.get_uid(\"wgt_k\")))\n\n    def _get_idx_mat(self, length):\n        idx_mat = tf.reshape(tf.tile(tf.range(length), [length]), [length, length])\n        idx_mat = idx_mat - tf.transpose(idx_mat)\n        return tf.clip_by_value(idx_mat, -self.num_nbr, self.num_nbr) + self.num_nbr\n\n    def _mat_mul(self, x, y, z, transpose):\n        xy = tf.matmul(x, y, transpose_b=transpose)\n        x = tf.keras.layers.Permute((2, 1, 3))(x)\n        mul = tf.matmul(x, z, transpose_b=transpose)\n        mul = tf.keras.layers.Permute((2, 1, 3))(mul)\n        return xy + mul\n\n    def call(self, inputs, mask=None):\n        if mask is not None:\n            mask = mask[:, None, None, :]\n        length = inputs.shape[1]\n        idx_mat = self._get_idx_mat(length)\n        rpe_k = tf.gather(self.wgt_k, idx_mat)\n        rpe_v = tf.gather(self.wgt_v, idx_mat)\n\n        qkv = self.qkv(inputs)\n        qkv = tf.keras.layers.Permute((2, 1, 3))(tf.keras.layers.Reshape((-1, self.num_heads, self.head_dim * 3))(qkv))\n        q, k, v = tf.split(qkv, [self.head_dim] * 3, axis=-1)\n        logits = self._mat_mul(q, k, rpe_k, True) * self.scale\n        attn = tf.keras.layers.Softmax(axis=-1)(logits, mask=mask)\n        attn = self.drop1(attn)\n        x = self._mat_mul(attn, v, rpe_v, False)\n        x = tf.keras.layers.Reshape((-1, self.dim))(tf.keras.layers.Permute((2, 1, 3))(x))\n        x = self.proj(x)\n        return x\n```\n\n**Intermediate CTC**\n\n```\ndef get_base_model(max_len=384, dim=384):\n    inp = tf.keras.Input((max_len,CHANNELS), name=\"inp\")\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n    ksize = 17\n    drop_rate = 0.2\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    inter = tf.keras.layers.Dense(dim, activation=None)(x)\n    inter = tf.keras.layers.Dense(NUM_CLASSES+1)(inter)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation=None,name='top_conv')(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES+1, name='classifier')(x)\n    return tf.keras.Model(inp, [inter, x])\n\ndef ctc_loss(args):\n    y_pred, inp_len, label, label_len = args\n    return tf.nn.ctc_loss(label, y_pred, label_len, inp_len,\n                          logits_time_major=False, blank_index=NUM_CLASSES)\n\ndef get_model(base_model):\n    label = tf.keras.Input(shape=(MAX_LEN_OUTPUT,), dtype='int32', name=\"label\")\n    label_len = tf.keras.Input(shape=(), dtype='int32', name=\"label_len\")\n    inp_len = tf.keras.Input(shape=(), dtype='int32', name=\"inp_len\")\n    loss1 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[0], inp_len, label, label_len])\n    loss2 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[1], inp_len, label, label_len])\n    loss = 0.3 * loss1 + 0.7 * loss2\n    #loss = 0.0 * loss1 + 1.0 * loss2\n    model = tf.keras.Model(\n        inputs=[base_model.input, inp_len, label, label_len],\n        outputs=loss)\n    return model\n```\n\n**Fine tuning** (LB +0.005)\nThe model was first trained with both training data and supplementary data with intermediate CTC (0.3 and 0.7) for 300 epochs. Then the pre-trained model was fine tuned only on the training data without intermediate CTC (0.0 and 1.0)\n\nI will post the training notebook soon.\n\n@hoyso48's solution\nhttps://www.kaggle.com/code/hoyso48/1st-place-solution-training",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2430001": "Thanks to Kaggle and Google for hosting this competition. This is my first Kaggle competition experence and I have learnt a lot from the previous ASL competition and other participants. I start my model from @hoyso48's solution from the last competition and added some components:\n\n**Data Augmentation**\nApplied random time warpping  (LB +0.01)\n\n```\ndef temporal_warp(x, factor=8):\n    n = tf.shape(x)[0]\n    if n < 5:\n      return x\n    w = n // factor\n    t1 = tf.clip_by_value(tf.random.uniform([], w, n-w, dtype=tf.int32), 1, n-2)\n    t2 = tf.clip_by_value(tf.random.uniform([], -w, w+1, dtype=tf.int32) + t1, 1, n-2)\n    new_x1 = interp1d_(x[:t1], t2)\n    new_x2 = interp1d_(x[t1:], n-t2)\n    return tf.concat([new_x1, new_x2], axis=0)\n\ndef augment_fn(x, always=False, max_len=None):\n    if tf.random.uniform(())<0.8 or always:\n        x = resample(x, (0.5,1.5))\n    if tf.random.uniform(())<0.5 or always:\n        x = flip_lr(x)\n    if max_len is not None:\n        x = temporal_crop(x, max_len)\n    if tf.random.uniform(())<0.75 or always:\n        x = spatial_random_affine(x)\n    if tf.random.uniform(())<0.75 or always:\n        x = temporal_warp(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = temporal_mask(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = spatial_mask(x)\n    return x\n```\n\n**Data Preprocessing**\nTruncate `phrase` to have smaller length than the input\n```\nif inp_len <= label_len:\n    offset = tf.random.uniform((), 0, label_len-inp_len+2, dtype=tf.int32)\n    label = label[offset:offset+inp_len-1]\n    label_len = inp_len-1\n```\n\n**Multi-headed attention with relative positional encoding** (LB +0.02)\n\n```\nclass MHSAwithRPE(tf.keras.layers.Layer):\n    def __init__(self, dim=256, num_heads=4, num_nbr=8, dropout=0, **kwargs):\n        super().__init__(**kwargs)\n        self.dim = dim\n        self.head_dim = dim // num_heads\n        self.scale = self.head_dim  ** -0.5\n        self.num_heads = num_heads\n        self.num_nbr = num_nbr\n        self.qkv = tf.keras.layers.Dense(3 * dim, use_bias=False)\n        self.drop1 = tf.keras.layers.Dropout(dropout)\n        self.proj = tf.keras.layers.Dense(dim, use_bias=False)\n        self.supports_masking = True\n\n        self.wgt_v = self.add_weight(shape=(num_nbr*2+1, self.head_dim),\n                                     initializer='glorot_uniform',trainable=True,\n                                     name=\"wgt_v\"+str(tf.keras.backend.get_uid(\"wgt_v\")))\n        self.wgt_k = self.add_weight(shape=(num_nbr*2+1, self.head_dim),\n                                     initializer='glorot_uniform',trainable=True,\n                                     name=\"wgt_k\"+str(tf.keras.backend.get_uid(\"wgt_k\")))\n\n    def _get_idx_mat(self, length):\n        idx_mat = tf.reshape(tf.tile(tf.range(length), [length]), [length, length])\n        idx_mat = idx_mat - tf.transpose(idx_mat)\n        return tf.clip_by_value(idx_mat, -self.num_nbr, self.num_nbr) + self.num_nbr\n\n    def _mat_mul(self, x, y, z, transpose):\n        xy = tf.matmul(x, y, transpose_b=transpose)\n        x = tf.keras.layers.Permute((2, 1, 3))(x)\n        mul = tf.matmul(x, z, transpose_b=transpose)\n        mul = tf.keras.layers.Permute((2, 1, 3))(mul)\n        return xy + mul\n\n    def call(self, inputs, mask=None):\n        if mask is not None:\n            mask = mask[:, None, None, :]\n        length = inputs.shape[1]\n        idx_mat = self._get_idx_mat(length)\n        rpe_k = tf.gather(self.wgt_k, idx_mat)\n        rpe_v = tf.gather(self.wgt_v, idx_mat)\n\n        qkv = self.qkv(inputs)\n        qkv = tf.keras.layers.Permute((2, 1, 3))(tf.keras.layers.Reshape((-1, self.num_heads, self.head_dim * 3))(qkv))\n        q, k, v = tf.split(qkv, [self.head_dim] * 3, axis=-1)\n        logits = self._mat_mul(q, k, rpe_k, True) * self.scale\n        attn = tf.keras.layers.Softmax(axis=-1)(logits, mask=mask)\n        attn = self.drop1(attn)\n        x = self._mat_mul(attn, v, rpe_v, False)\n        x = tf.keras.layers.Reshape((-1, self.dim))(tf.keras.layers.Permute((2, 1, 3))(x))\n        x = self.proj(x)\n        return x\n```\n\n**Intermediate CTC**\n\n```\ndef get_base_model(max_len=384, dim=384):\n    inp = tf.keras.Input((max_len,CHANNELS), name=\"inp\")\n    x = tf.keras.layers.Masking(mask_value=PAD,input_shape=(max_len,CHANNELS))(inp)\n    x = tf.keras.layers.Dense(dim, use_bias=False,name='stem_conv')(x)\n    x = tf.keras.layers.BatchNormalization(momentum=0.95,name='stem_bn')(x)\n\n    ksize = 17\n    drop_rate = 0.2\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    inter = tf.keras.layers.Dense(dim, activation=None)(x)\n    inter = tf.keras.layers.Dense(NUM_CLASSES+1)(inter)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = Conv1DBlock(dim,ksize,drop_rate=drop_rate)(x)\n    x = TransformerBlock(dim,expand=2)(x)\n\n    x = tf.keras.layers.Dense(dim*2,activation=None,name='top_conv')(x)\n    x = tf.keras.layers.Dense(NUM_CLASSES+1, name='classifier')(x)\n    return tf.keras.Model(inp, [inter, x])\n\ndef ctc_loss(args):\n    y_pred, inp_len, label, label_len = args\n    return tf.nn.ctc_loss(label, y_pred, label_len, inp_len,\n                          logits_time_major=False, blank_index=NUM_CLASSES)\n\ndef get_model(base_model):\n    label = tf.keras.Input(shape=(MAX_LEN_OUTPUT,), dtype='int32', name=\"label\")\n    label_len = tf.keras.Input(shape=(), dtype='int32', name=\"label_len\")\n    inp_len = tf.keras.Input(shape=(), dtype='int32', name=\"inp_len\")\n    loss1 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[0], inp_len, label, label_len])\n    loss2 = tf.keras.layers.Lambda(ctc_loss, output_shape=())(\n        [base_model.output[1], inp_len, label, label_len])\n    loss = 0.3 * loss1 + 0.7 * loss2\n    #loss = 0.0 * loss1 + 1.0 * loss2\n    model = tf.keras.Model(\n        inputs=[base_model.input, inp_len, label, label_len],\n        outputs=loss)\n    return model\n```\n\n**Fine tuning** (LB +0.005)\nThe model was first trained with both training data and supplementary data with intermediate CTC (0.3 and 0.7) for 300 epochs. Then the pre-trained model was fine tuned only on the training data without intermediate CTC (0.0 and 1.0)\n\nI will post the training notebook soon.\n\n@hoyso48's solution\nhttps://www.kaggle.com/code/hoyso48/1st-place-solution-training"
  }
}