{
  "id": 436331,
  "title": "122nd Solution for Google - American Sign Language Fingerspelling Recognition Competition",
  "url": "/competitions/asl-fingerspelling/discussion/436331",
  "author_name": "Tony_Zhang_2004",
  "post_date": "2023-09-01T23:17:28.548000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>We did this in our freshman summer, defintely learnt a lot.</p>\n<p><strong>Context</strong></p>\n<ul>\n<li>Business context:<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></li>\n<li>Data context:<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></li>\n</ul>\n<p><strong>Overview of the Approach</strong><br>\nOur work is basically some improvements based on the <a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">public notebook baseline</a> w/ LB score 0.699, and we improved it w/ +0.007.<br>\nIt conbines Transformer and 1D-CNN.<br>\nWe basically improved it by doing some feature engineering to the frames.</p>\n<p><strong>Details on the Submission</strong><br>\nOriginally our work was based on another public notebook baseline which only uses the Transformer, and we did some improvement on that by rescheduling the learning rate, specifically, to increase the learning rate at the end, but it didn't work for our final notebook.<br>\nWe also processed the supplemental data, but it didn't have some improvements(maybe we did it the wrong way).<br>\nOur data augmentation includes some techniques like random cropping, rotation, affine transformations, etc.</p>\n<p><strong>Code Samples for Data Augmentation</strong></p>\n<pre><code>@tf.function()\ndef :\n     tf.[0] &lt; FRAME_LEN:\n        x = tf.pad(x, ([[, FRAME_LEN-tf.shape(x)[]], [, ], [, ]]), constant_values=())\n    :\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[]))\n     x\n\n@tf.function()\ndef resize_pad_channel(x):\n     tf.shape(x)[] &lt; FRAME_LEN:\n        x = tf.pad(x, ([[, FRAME_LEN-tf.shape(x)[]], [, ], [, ]]), constant_values=())\n    :\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[]))\n     x\n\ndef interp1d_(x, target_len, method=):\n    target_len = tf.maximum(,target_len)\n     method == :\n        \n         tf.random.uniform(()) &lt; :\n            x = tf.image.resize(x, (target_len,),)\n        :\n             tf.random.uniform(()) &lt; :\n                x = tf.image.resize(x, (target_len,),)\n            :\n                x = tf.image.resize(x, (target_len,),)\n    :\n        x = tf.image.resize(x, (target_len,),method)\n     resize_pad(x)\n\ndef flip_lr(x):\n    \n    dx, dy, dz = tf.unstack(x, axis=) \n    dx = -dx\n    new_x = tf.([dx,dy,dz], )\n     new_x\n\ndef resample(x, rate=(,)):\n    rate = tf.random.uniform((), rate[], rate[])\n    length = tf.shape(x)[]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_x = tf.image.resize(x, (new_size, ))\n     resize_pad_channel(new_x)\n\ndef spatial_random_affine(xyz,\n    scale  = (,),\n    shear = (,),\n    shift  = (,),\n    degree = (,),\n):\n    center = tf.constant([,])\n     scale is not None:\n        scale = tf.random.uniform((),*scale)\n        xyz = scale*xyz\n\n     shear is not None:\n        xy = xyz[...,:]\n        z = xyz[...,:]\n        shear_x = shear_y = tf.random.uniform((),*shear)\n         tf.random.uniform(()) &lt; :\n            shear_x = \n        :\n            shear_y = \n        shear_mat = tf.identity([\n            [,shear_x],\n            [shear_y,]\n        ])\n        xy = xy @ shear_mat\n        center = center + [shear_y, shear_x]\n        xyz = tf.concat([xy,z], axis=)\n\n     degree is not None:\n        xy = xyz[...,:]\n        z = xyz[...,:]\n        xy -= center\n        degree = tf.random.uniform((),*degree)\n        radian = degree/*np.pi\n        c = tf.math.(radian)\n        s = tf.math.(radian)\n        rotate_mat = tf.identity([\n            [c,s],\n            [-s, c],\n        ])\n        xy = xy @ rotate_mat\n        xy = xy + center\n        xyz = tf.concat([xy,z], axis=)\n\n     shift is not None:\n        shift = tf.random.uniform((),*shift)\n        xyz = xyz + shift\n\n     xyz\n\ndef temporal_crop(x, length=FRAME_LEN):\n    l = tf.shape(x)[]\n    offset = tf.random.uniform((), , tf.clip_by_value(l-length,,length), dtype=tf.int32)\n    x = x[offset:offset+length]\n     x\n\ndef temporal_mask0(x, size=(,), mask_value=()):\n    l = tf.shape(x)[]\n    mask_size = tf.random.uniform((), *size)\n    mask_size = tf.cast(tf.cast(l, tf.float32) * mask_size, tf.int32)\n    mask_offset = tf.random.uniform((), , tf.clip_by_value(l-mask_size,,l), dtype=tf.int32)\n    x = tf.tensor_scatter_nd_update(x,tf.range(mask_offset, mask_offset+mask_size)[...,None],tf.fill([mask_size,,],mask_value))\n     x\n\ndef temporal_mask(x, rate=, mask_value=()):\n    \n    mask_size=(FRAME_LEN*rate)\n    mask = tf.squeeze(tf.random.categorical(np.mat([/FRAME_LEN  i in range(FRAME_LEN)]),mask_size))\n    \n    x = tf.tensor_scatter_nd_update(x,mask[...,None],tf.fill([mask_size,,],mask_value))\n     x\n\ndef spatial_mask(x, size=(,), mask_value=()):\n    \n    mask_offset_y = tf.random.uniform(())\n    mask_offset_x = tf.random.uniform(())\n    mask_size = tf.random.uniform((), *size)\n    mask_x = (mask_offset_x&lt;x[...,]) &amp; (x[...,] &lt; mask_offset_x + mask_size)\n    mask_y = (mask_offset_y&lt;x[...,]) &amp; (x[...,] &lt; mask_offset_y + mask_size)\n    mask = mask_x &amp; mask_y\n    x = tf.where(mask[...,None], mask_value, x)\n     x\n\ndef augment_fn(x, always=False):\n     tf.random.uniform(())&lt; or always:\n        x = resample(x, (,))\n     tf.random.uniform(())&lt; or always:\n        x = flip_lr(x)\n     tf.random.uniform(())&lt; or always:\n        x = spatial_random_affine(x)\n     tf.random.uniform(())&lt; or always:\n        x = temporal_mask(x)\n     tf.random.uniform(())&lt; or always:\n        x = spatial_mask(x)\n     x\n\n@tf.function(jit_compile=True)\ndef pre_process0(x):\n    lip_x = tf.gather(x, LIP_IDX_X, axis=)\n    lip_y = tf.gather(x, LIP_IDX_Y, axis=)\n    lip_z = tf.gather(x, LIP_IDX_Z, axis=)\n\n    rhand_x = tf.gather(x, RHAND_IDX_X, axis=)\n    rhand_y = tf.gather(x, RHAND_IDX_Y, axis=)\n    rhand_z = tf.gather(x, RHAND_IDX_Z, axis=)\n\n    lhand_x = tf.gather(x, LHAND_IDX_X, axis=)\n    lhand_y = tf.gather(x, LHAND_IDX_Y, axis=)\n    lhand_z = tf.gather(x, LHAND_IDX_Z, axis=)\n\n    rpose_x = tf.gather(x, RPOSE_IDX_X, axis=)\n    rpose_y = tf.gather(x, RPOSE_IDX_Y, axis=)\n    rpose_z = tf.gather(x, RPOSE_IDX_Z, axis=)\n\n    lpose_x = tf.gather(x, LPOSE_IDX_X, axis=)\n    lpose_y = tf.gather(x, LPOSE_IDX_Y, axis=)\n    lpose_z = tf.gather(x, LPOSE_IDX_Z, axis=)\n\n    lip   = tf.concat([lip_x[..., tf.newaxis], lip_y[..., tf.newaxis], lip_z[..., tf.newaxis]], axis=)\n    rhand = tf.concat([rhand_x[..., tf.newaxis], rhand_y[..., tf.newaxis], rhand_z[..., tf.newaxis]], axis=)\n    lhand = tf.concat([lhand_x[..., tf.newaxis], lhand_y[..., tf.newaxis], lhand_z[..., tf.newaxis]], axis=)\n    rpose = tf.concat([rpose_x[..., tf.newaxis], rpose_y[..., tf.newaxis], rpose_z[..., tf.newaxis]], axis=)\n    lpose = tf.concat([lpose_x[..., tf.newaxis], lpose_y[..., tf.newaxis], lpose_z[..., tf.newaxis]], axis=)\n\n    hand =  tf.concat([rhand, lhand], axis=)\n    hand = tf.where(tf.math.is_nan(hand), , hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n\n    lip = lip[mask]\n    rhand = rhand[mask]\n    lhand = lhand[mask]\n    rpose = rpose[mask]\n    lpose = lpose[mask]\n\n     lip, rhand,lhand,  rpose, lpose \n\n@tf.function()\ndef pre_process1(lip, rhand,lhand,  rpose, lpose, augment = False): \n    lip   = (resize_pad(lip) - LIPM) / LIPS\n    rhand = (resize_pad(rhand) - RHM) / RHS\n    lhand = (resize_pad(lhand) - LHM) / LHS\n    rpose = (resize_pad(rpose) - RPM) / RPS\n    lpose = (resize_pad(lpose) - LPM) / LPS\n\n    x = tf.concat([lip, rhand, lhand, rpose, lpose], axis=) \n    x = tf.where(tf.math.is_nan(x), , x)\n     augment:\n        x = augment_fn(x)\n    s = tf.shape(x)\n    x = tf.reshape(x, (s[], s[]*s[]))\n     x\n\npre0 = pre_process0(frames)\npre1 = pre_process1(*pre0)\nINPUT_SHAPE = (pre1.shape)\nprint(INPUT_SHAPE)\n</code></pre>\n<p><strong>Some Useful References</strong><br>\n<a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">https://www.kaggle.com/code/royalacecat/the-deeper-the-better</a><br>\n<a href=\"https://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing\" target=\"_blank\">https://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing</a><br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438</a></p>",
  "messages": [
    {
      "id": 2419335,
      "postDate": "2023-09-01T23:17:28.547Z",
      "content": "<p>We did this in our freshman summer, defintely learnt a lot.</p>\n<p><strong>Context</strong></p>\n<ul>\n<li>Business context:<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/overview\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/overview</a></li>\n<li>Data context:<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/data\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/data</a></li>\n</ul>\n<p><strong>Overview of the Approach</strong><br>\nOur work is basically some improvements based on the <a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">public notebook baseline</a> w/ LB score 0.699, and we improved it w/ +0.007.<br>\nIt conbines Transformer and 1D-CNN.<br>\nWe basically improved it by doing some feature engineering to the frames.</p>\n<p><strong>Details on the Submission</strong><br>\nOriginally our work was based on another public notebook baseline which only uses the Transformer, and we did some improvement on that by rescheduling the learning rate, specifically, to increase the learning rate at the end, but it didn't work for our final notebook.<br>\nWe also processed the supplemental data, but it didn't have some improvements(maybe we did it the wrong way).<br>\nOur data augmentation includes some techniques like random cropping, rotation, affine transformations, etc.</p>\n<p><strong>Code Samples for Data Augmentation</strong></p>\n<pre><code>@tf.function()\ndef :\n     tf.[0] &lt; FRAME_LEN:\n        x = tf.pad(x, ([[, FRAME_LEN-tf.shape(x)[]], [, ], [, ]]), constant_values=())\n    :\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[]))\n     x\n\n@tf.function()\ndef resize_pad_channel(x):\n     tf.shape(x)[] &lt; FRAME_LEN:\n        x = tf.pad(x, ([[, FRAME_LEN-tf.shape(x)[]], [, ], [, ]]), constant_values=())\n    :\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[]))\n     x\n\ndef interp1d_(x, target_len, method=):\n    target_len = tf.maximum(,target_len)\n     method == :\n        \n         tf.random.uniform(()) &lt; :\n            x = tf.image.resize(x, (target_len,),)\n        :\n             tf.random.uniform(()) &lt; :\n                x = tf.image.resize(x, (target_len,),)\n            :\n                x = tf.image.resize(x, (target_len,),)\n    :\n        x = tf.image.resize(x, (target_len,),method)\n     resize_pad(x)\n\ndef flip_lr(x):\n    \n    dx, dy, dz = tf.unstack(x, axis=) \n    dx = -dx\n    new_x = tf.([dx,dy,dz], )\n     new_x\n\ndef resample(x, rate=(,)):\n    rate = tf.random.uniform((), rate[], rate[])\n    length = tf.shape(x)[]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_x = tf.image.resize(x, (new_size, ))\n     resize_pad_channel(new_x)\n\ndef spatial_random_affine(xyz,\n    scale  = (,),\n    shear = (,),\n    shift  = (,),\n    degree = (,),\n):\n    center = tf.constant([,])\n     scale is not None:\n        scale = tf.random.uniform((),*scale)\n        xyz = scale*xyz\n\n     shear is not None:\n        xy = xyz[...,:]\n        z = xyz[...,:]\n        shear_x = shear_y = tf.random.uniform((),*shear)\n         tf.random.uniform(()) &lt; :\n            shear_x = \n        :\n            shear_y = \n        shear_mat = tf.identity([\n            [,shear_x],\n            [shear_y,]\n        ])\n        xy = xy @ shear_mat\n        center = center + [shear_y, shear_x]\n        xyz = tf.concat([xy,z], axis=)\n\n     degree is not None:\n        xy = xyz[...,:]\n        z = xyz[...,:]\n        xy -= center\n        degree = tf.random.uniform((),*degree)\n        radian = degree/*np.pi\n        c = tf.math.(radian)\n        s = tf.math.(radian)\n        rotate_mat = tf.identity([\n            [c,s],\n            [-s, c],\n        ])\n        xy = xy @ rotate_mat\n        xy = xy + center\n        xyz = tf.concat([xy,z], axis=)\n\n     shift is not None:\n        shift = tf.random.uniform((),*shift)\n        xyz = xyz + shift\n\n     xyz\n\ndef temporal_crop(x, length=FRAME_LEN):\n    l = tf.shape(x)[]\n    offset = tf.random.uniform((), , tf.clip_by_value(l-length,,length), dtype=tf.int32)\n    x = x[offset:offset+length]\n     x\n\ndef temporal_mask0(x, size=(,), mask_value=()):\n    l = tf.shape(x)[]\n    mask_size = tf.random.uniform((), *size)\n    mask_size = tf.cast(tf.cast(l, tf.float32) * mask_size, tf.int32)\n    mask_offset = tf.random.uniform((), , tf.clip_by_value(l-mask_size,,l), dtype=tf.int32)\n    x = tf.tensor_scatter_nd_update(x,tf.range(mask_offset, mask_offset+mask_size)[...,None],tf.fill([mask_size,,],mask_value))\n     x\n\ndef temporal_mask(x, rate=, mask_value=()):\n    \n    mask_size=(FRAME_LEN*rate)\n    mask = tf.squeeze(tf.random.categorical(np.mat([/FRAME_LEN  i in range(FRAME_LEN)]),mask_size))\n    \n    x = tf.tensor_scatter_nd_update(x,mask[...,None],tf.fill([mask_size,,],mask_value))\n     x\n\ndef spatial_mask(x, size=(,), mask_value=()):\n    \n    mask_offset_y = tf.random.uniform(())\n    mask_offset_x = tf.random.uniform(())\n    mask_size = tf.random.uniform((), *size)\n    mask_x = (mask_offset_x&lt;x[...,]) &amp; (x[...,] &lt; mask_offset_x + mask_size)\n    mask_y = (mask_offset_y&lt;x[...,]) &amp; (x[...,] &lt; mask_offset_y + mask_size)\n    mask = mask_x &amp; mask_y\n    x = tf.where(mask[...,None], mask_value, x)\n     x\n\ndef augment_fn(x, always=False):\n     tf.random.uniform(())&lt; or always:\n        x = resample(x, (,))\n     tf.random.uniform(())&lt; or always:\n        x = flip_lr(x)\n     tf.random.uniform(())&lt; or always:\n        x = spatial_random_affine(x)\n     tf.random.uniform(())&lt; or always:\n        x = temporal_mask(x)\n     tf.random.uniform(())&lt; or always:\n        x = spatial_mask(x)\n     x\n\n@tf.function(jit_compile=True)\ndef pre_process0(x):\n    lip_x = tf.gather(x, LIP_IDX_X, axis=)\n    lip_y = tf.gather(x, LIP_IDX_Y, axis=)\n    lip_z = tf.gather(x, LIP_IDX_Z, axis=)\n\n    rhand_x = tf.gather(x, RHAND_IDX_X, axis=)\n    rhand_y = tf.gather(x, RHAND_IDX_Y, axis=)\n    rhand_z = tf.gather(x, RHAND_IDX_Z, axis=)\n\n    lhand_x = tf.gather(x, LHAND_IDX_X, axis=)\n    lhand_y = tf.gather(x, LHAND_IDX_Y, axis=)\n    lhand_z = tf.gather(x, LHAND_IDX_Z, axis=)\n\n    rpose_x = tf.gather(x, RPOSE_IDX_X, axis=)\n    rpose_y = tf.gather(x, RPOSE_IDX_Y, axis=)\n    rpose_z = tf.gather(x, RPOSE_IDX_Z, axis=)\n\n    lpose_x = tf.gather(x, LPOSE_IDX_X, axis=)\n    lpose_y = tf.gather(x, LPOSE_IDX_Y, axis=)\n    lpose_z = tf.gather(x, LPOSE_IDX_Z, axis=)\n\n    lip   = tf.concat([lip_x[..., tf.newaxis], lip_y[..., tf.newaxis], lip_z[..., tf.newaxis]], axis=)\n    rhand = tf.concat([rhand_x[..., tf.newaxis], rhand_y[..., tf.newaxis], rhand_z[..., tf.newaxis]], axis=)\n    lhand = tf.concat([lhand_x[..., tf.newaxis], lhand_y[..., tf.newaxis], lhand_z[..., tf.newaxis]], axis=)\n    rpose = tf.concat([rpose_x[..., tf.newaxis], rpose_y[..., tf.newaxis], rpose_z[..., tf.newaxis]], axis=)\n    lpose = tf.concat([lpose_x[..., tf.newaxis], lpose_y[..., tf.newaxis], lpose_z[..., tf.newaxis]], axis=)\n\n    hand =  tf.concat([rhand, lhand], axis=)\n    hand = tf.where(tf.math.is_nan(hand), , hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[, ]), )\n\n    lip = lip[mask]\n    rhand = rhand[mask]\n    lhand = lhand[mask]\n    rpose = rpose[mask]\n    lpose = lpose[mask]\n\n     lip, rhand,lhand,  rpose, lpose \n\n@tf.function()\ndef pre_process1(lip, rhand,lhand,  rpose, lpose, augment = False): \n    lip   = (resize_pad(lip) - LIPM) / LIPS\n    rhand = (resize_pad(rhand) - RHM) / RHS\n    lhand = (resize_pad(lhand) - LHM) / LHS\n    rpose = (resize_pad(rpose) - RPM) / RPS\n    lpose = (resize_pad(lpose) - LPM) / LPS\n\n    x = tf.concat([lip, rhand, lhand, rpose, lpose], axis=) \n    x = tf.where(tf.math.is_nan(x), , x)\n     augment:\n        x = augment_fn(x)\n    s = tf.shape(x)\n    x = tf.reshape(x, (s[], s[]*s[]))\n     x\n\npre0 = pre_process0(frames)\npre1 = pre_process1(*pre0)\nINPUT_SHAPE = (pre1.shape)\nprint(INPUT_SHAPE)\n</code></pre>\n<p><strong>Some Useful References</strong><br>\n<a href=\"https://www.kaggle.com/code/royalacecat/the-deeper-the-better\" target=\"_blank\">https://www.kaggle.com/code/royalacecat/the-deeper-the-better</a><br>\n<a href=\"https://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing\" target=\"_blank\">https://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing</a><br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438</a></p>",
      "rawMarkdown": "We did this in our freshman summer, defintely learnt a lot.\n\n**Context**\n- Business context:https://www.kaggle.com/competitions/asl-fingerspelling/overview\n- Data context:https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n**Overview of the Approach**\nOur work is basically some improvements based on the [public notebook baseline](https://www.kaggle.com/code/royalacecat/the-deeper-the-better) w/ LB score 0.699, and we improved it w/ +0.007.\nIt conbines Transformer and 1D-CNN.\nWe basically improved it by doing some feature engineering to the frames.\n\n**Details on the Submission**\nOriginally our work was based on another public notebook baseline which only uses the Transformer, and we did some improvement on that by rescheduling the learning rate, specifically, to increase the learning rate at the end, but it didn't work for our final notebook.\nWe also processed the supplemental data, but it didn't have some improvements(maybe we did it the wrong way).\nOur data augmentation includes some techniques like random cropping, rotation, affine transformations, etc.\n\n**Code Samples for Data Augmentation**\n```c\n@tf.function()\ndef resize_pad(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]), constant_values=float(\"NaN\"))\n    else:\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\n@tf.function()\ndef resize_pad_channel(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]), constant_values=float(0))\n    else:\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\ndef interp1d_(x, target_len, method='random'):\n    target_len = tf.maximum(1,target_len)\n    if method == 'random':\n        # random interpolation\n        if tf.random.uniform(()) < 0.33:\n            x = tf.image.resize(x, (target_len,92),'bilinear')\n        else:\n            if tf.random.uniform(()) < 0.5:\n                x = tf.image.resize(x, (target_len,92),'bicubic')\n            else:\n                x = tf.image.resize(x, (target_len,92),'nearest')\n    else:\n        x = tf.image.resize(x, (target_len,92),method)\n    return resize_pad(x)\n\ndef flip_lr(x):\n    # flip left and right\n    dx, dy, dz = tf.unstack(x, axis=-1) # unstack in a dimension\n    dx = -dx\n    new_x = tf.stack([dx,dy,dz], -1)\n    return new_x\n\ndef resample(x, rate=(0.8,1.2)):\n    rate = tf.random.uniform((), rate[0], rate[1])\n    length = tf.shape(x)[0]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_x = tf.image.resize(x, (new_size, 92))\n    return resize_pad_channel(new_x)\n\ndef spatial_random_affine(xyz,\n    scale  = (0.8,1.2),\n    shear = (-0.15,0.15),\n    shift  = (-0.2,0.2),\n    degree = (-30,30),\n):\n    center = tf.constant([0.0,0.0])\n    if scale is not None:\n        scale = tf.random.uniform((),*scale)\n        xyz = scale*xyz\n\n    if shear is not None:\n        xy = xyz[...,:2]\n        z = xyz[...,2:]\n        shear_x = shear_y = tf.random.uniform((),*shear)\n        if tf.random.uniform(()) < 0.5:\n            shear_x = 0.\n        else:\n            shear_y = 0.\n        shear_mat = tf.identity([\n            [1.,shear_x],\n            [shear_y,1.]\n        ])\n        xy = xy @ shear_mat\n        center = center + [shear_y, shear_x]\n        xyz = tf.concat([xy,z], axis=-1)\n\n    if degree is not None:\n        xy = xyz[...,:2]\n        z = xyz[...,2:]\n        xy -= center\n        degree = tf.random.uniform((),*degree)\n        radian = degree/180*np.pi\n        c = tf.math.cos(radian)\n        s = tf.math.sin(radian)\n        rotate_mat = tf.identity([\n            [c,s],\n            [-s, c],\n        ])\n        xy = xy @ rotate_mat\n        xy = xy + center\n        xyz = tf.concat([xy,z], axis=-1)\n\n    if shift is not None:\n        shift = tf.random.uniform((),*shift)\n        xyz = xyz + shift\n\n    return xyz\n\ndef temporal_crop(x, length=FRAME_LEN):\n    l = tf.shape(x)[0]\n    offset = tf.random.uniform((), 0, tf.clip_by_value(l-length,1,length), dtype=tf.int32)\n    x = x[offset:offset+length]\n    return x\n\ndef temporal_mask0(x, size=(0.2,0.4), mask_value=float(0.0)):\n    l = tf.shape(x)[0]\n    mask_size = tf.random.uniform((), *size)\n    mask_size = tf.cast(tf.cast(l, tf.float32) * mask_size, tf.int32)\n    mask_offset = tf.random.uniform((), 0, tf.clip_by_value(l-mask_size,1,l), dtype=tf.int32)\n    x = tf.tensor_scatter_nd_update(x,tf.range(mask_offset, mask_offset+mask_size)[...,None],tf.fill([mask_size,92,3],mask_value))\n    return x\n\ndef temporal_mask(x, rate=0.2, mask_value=float(0.0)):\n    # mask 1/10 frames randomly\n    mask_size=int(FRAME_LEN*rate)\n    mask = tf.squeeze(tf.random.categorical(np.mat([1/FRAME_LEN for i in range(FRAME_LEN)]),mask_size))\n    # print(mask)\n    x = tf.tensor_scatter_nd_update(x,mask[...,None],tf.fill([mask_size,92,3],mask_value))\n    return x\n\ndef spatial_mask(x, size=(0.2,0.4), mask_value=float(0.0)):\n    # mask by a randome rectangle\n    mask_offset_y = tf.random.uniform(())\n    mask_offset_x = tf.random.uniform(())\n    mask_size = tf.random.uniform((), *size)\n    mask_x = (mask_offset_x<x[...,0]) & (x[...,0] < mask_offset_x + mask_size)\n    mask_y = (mask_offset_y<x[...,1]) & (x[...,1] < mask_offset_y + mask_size)\n    mask = mask_x & mask_y\n    x = tf.where(mask[...,None], mask_value, x)\n    return x\n\ndef augment_fn(x, always=False):\n    if tf.random.uniform(())<0.5 or always:\n        x = resample(x, (0.6,1.4))\n    if tf.random.uniform(())<0.5 or always:\n        x = flip_lr(x)\n    if tf.random.uniform(())<0.7 or always:\n        x = spatial_random_affine(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = temporal_mask(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = spatial_mask(x)\n    return x\n\n@tf.function(jit_compile=True)\ndef pre_process0(x):\n    lip_x = tf.gather(x, LIP_IDX_X, axis=1)\n    lip_y = tf.gather(x, LIP_IDX_Y, axis=1)\n    lip_z = tf.gather(x, LIP_IDX_Z, axis=1)\n\n    rhand_x = tf.gather(x, RHAND_IDX_X, axis=1)\n    rhand_y = tf.gather(x, RHAND_IDX_Y, axis=1)\n    rhand_z = tf.gather(x, RHAND_IDX_Z, axis=1)\n    \n    lhand_x = tf.gather(x, LHAND_IDX_X, axis=1)\n    lhand_y = tf.gather(x, LHAND_IDX_Y, axis=1)\n    lhand_z = tf.gather(x, LHAND_IDX_Z, axis=1)\n\n    rpose_x = tf.gather(x, RPOSE_IDX_X, axis=1)\n    rpose_y = tf.gather(x, RPOSE_IDX_Y, axis=1)\n    rpose_z = tf.gather(x, RPOSE_IDX_Z, axis=1)\n    \n    lpose_x = tf.gather(x, LPOSE_IDX_X, axis=1)\n    lpose_y = tf.gather(x, LPOSE_IDX_Y, axis=1)\n    lpose_z = tf.gather(x, LPOSE_IDX_Z, axis=1)\n    \n    lip   = tf.concat([lip_x[..., tf.newaxis], lip_y[..., tf.newaxis], lip_z[..., tf.newaxis]], axis=-1)\n    rhand = tf.concat([rhand_x[..., tf.newaxis], rhand_y[..., tf.newaxis], rhand_z[..., tf.newaxis]], axis=-1)\n    lhand = tf.concat([lhand_x[..., tf.newaxis], lhand_y[..., tf.newaxis], lhand_z[..., tf.newaxis]], axis=-1)\n    rpose = tf.concat([rpose_x[..., tf.newaxis], rpose_y[..., tf.newaxis], rpose_z[..., tf.newaxis]], axis=-1)\n    lpose = tf.concat([lpose_x[..., tf.newaxis], lpose_y[..., tf.newaxis], lpose_z[..., tf.newaxis]], axis=-1)\n    \n    hand =  tf.concat([rhand, lhand], axis=1)\n    hand = tf.where(tf.math.is_nan(hand), 0.0, hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n\n    lip = lip[mask]\n    rhand = rhand[mask]\n    lhand = lhand[mask]\n    rpose = rpose[mask]\n    lpose = lpose[mask]\n\n    return lip, rhand,lhand,  rpose, lpose #lhand\n\n@tf.function()\ndef pre_process1(lip, rhand,lhand,  rpose, lpose, augment = False): #lhand,\n    lip   = (resize_pad(lip) - LIPM) / LIPS\n    rhand = (resize_pad(rhand) - RHM) / RHS\n    lhand = (resize_pad(lhand) - LHM) / LHS\n    rpose = (resize_pad(rpose) - RPM) / RPS\n    lpose = (resize_pad(lpose) - LPM) / LPS\n\n    x = tf.concat([lip, rhand, lhand, rpose, lpose], axis=1) #lhand,\n    x = tf.where(tf.math.is_nan(x), 0.0, x)\n    if augment:\n        x = augment_fn(x)\n    s = tf.shape(x)\n    x = tf.reshape(x, (s[0], s[1]*s[2]))\n    return x\n\npre0 = pre_process0(frames)\npre1 = pre_process1(*pre0)\nINPUT_SHAPE = list(pre1.shape)\nprint(INPUT_SHAPE)\n```\n\n**Some Useful References**\nhttps://www.kaggle.com/code/royalacecat/the-deeper-the-better\nhttps://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438"
    },
    {
      "id": 2426094,
      "postDate": "2023-09-06T12:00:06.533Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2426094,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-06T12:00:06.533000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2419335": "We did this in our freshman summer, defintely learnt a lot.\n\n**Context**\n- Business context:https://www.kaggle.com/competitions/asl-fingerspelling/overview\n- Data context:https://www.kaggle.com/competitions/asl-fingerspelling/data\n\n**Overview of the Approach**\nOur work is basically some improvements based on the [public notebook baseline](https://www.kaggle.com/code/royalacecat/the-deeper-the-better) w/ LB score 0.699, and we improved it w/ +0.007.\nIt conbines Transformer and 1D-CNN.\nWe basically improved it by doing some feature engineering to the frames.\n\n**Details on the Submission**\nOriginally our work was based on another public notebook baseline which only uses the Transformer, and we did some improvement on that by rescheduling the learning rate, specifically, to increase the learning rate at the end, but it didn't work for our final notebook.\nWe also processed the supplemental data, but it didn't have some improvements(maybe we did it the wrong way).\nOur data augmentation includes some techniques like random cropping, rotation, affine transformations, etc.\n\n**Code Samples for Data Augmentation**\n```c\n@tf.function()\ndef resize_pad(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]), constant_values=float(\"NaN\"))\n    else:\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\n@tf.function()\ndef resize_pad_channel(x):\n    if tf.shape(x)[0] < FRAME_LEN:\n        x = tf.pad(x, ([[0, FRAME_LEN-tf.shape(x)[0]], [0, 0], [0, 0]]), constant_values=float(0))\n    else:\n        x = tf.image.resize(x, (FRAME_LEN, tf.shape(x)[1]))\n    return x\n\ndef interp1d_(x, target_len, method='random'):\n    target_len = tf.maximum(1,target_len)\n    if method == 'random':\n        # random interpolation\n        if tf.random.uniform(()) < 0.33:\n            x = tf.image.resize(x, (target_len,92),'bilinear')\n        else:\n            if tf.random.uniform(()) < 0.5:\n                x = tf.image.resize(x, (target_len,92),'bicubic')\n            else:\n                x = tf.image.resize(x, (target_len,92),'nearest')\n    else:\n        x = tf.image.resize(x, (target_len,92),method)\n    return resize_pad(x)\n\ndef flip_lr(x):\n    # flip left and right\n    dx, dy, dz = tf.unstack(x, axis=-1) # unstack in a dimension\n    dx = -dx\n    new_x = tf.stack([dx,dy,dz], -1)\n    return new_x\n\ndef resample(x, rate=(0.8,1.2)):\n    rate = tf.random.uniform((), rate[0], rate[1])\n    length = tf.shape(x)[0]\n    new_size = tf.cast(rate*tf.cast(length,tf.float32), tf.int32)\n    new_x = tf.image.resize(x, (new_size, 92))\n    return resize_pad_channel(new_x)\n\ndef spatial_random_affine(xyz,\n    scale  = (0.8,1.2),\n    shear = (-0.15,0.15),\n    shift  = (-0.2,0.2),\n    degree = (-30,30),\n):\n    center = tf.constant([0.0,0.0])\n    if scale is not None:\n        scale = tf.random.uniform((),*scale)\n        xyz = scale*xyz\n\n    if shear is not None:\n        xy = xyz[...,:2]\n        z = xyz[...,2:]\n        shear_x = shear_y = tf.random.uniform((),*shear)\n        if tf.random.uniform(()) < 0.5:\n            shear_x = 0.\n        else:\n            shear_y = 0.\n        shear_mat = tf.identity([\n            [1.,shear_x],\n            [shear_y,1.]\n        ])\n        xy = xy @ shear_mat\n        center = center + [shear_y, shear_x]\n        xyz = tf.concat([xy,z], axis=-1)\n\n    if degree is not None:\n        xy = xyz[...,:2]\n        z = xyz[...,2:]\n        xy -= center\n        degree = tf.random.uniform((),*degree)\n        radian = degree/180*np.pi\n        c = tf.math.cos(radian)\n        s = tf.math.sin(radian)\n        rotate_mat = tf.identity([\n            [c,s],\n            [-s, c],\n        ])\n        xy = xy @ rotate_mat\n        xy = xy + center\n        xyz = tf.concat([xy,z], axis=-1)\n\n    if shift is not None:\n        shift = tf.random.uniform((),*shift)\n        xyz = xyz + shift\n\n    return xyz\n\ndef temporal_crop(x, length=FRAME_LEN):\n    l = tf.shape(x)[0]\n    offset = tf.random.uniform((), 0, tf.clip_by_value(l-length,1,length), dtype=tf.int32)\n    x = x[offset:offset+length]\n    return x\n\ndef temporal_mask0(x, size=(0.2,0.4), mask_value=float(0.0)):\n    l = tf.shape(x)[0]\n    mask_size = tf.random.uniform((), *size)\n    mask_size = tf.cast(tf.cast(l, tf.float32) * mask_size, tf.int32)\n    mask_offset = tf.random.uniform((), 0, tf.clip_by_value(l-mask_size,1,l), dtype=tf.int32)\n    x = tf.tensor_scatter_nd_update(x,tf.range(mask_offset, mask_offset+mask_size)[...,None],tf.fill([mask_size,92,3],mask_value))\n    return x\n\ndef temporal_mask(x, rate=0.2, mask_value=float(0.0)):\n    # mask 1/10 frames randomly\n    mask_size=int(FRAME_LEN*rate)\n    mask = tf.squeeze(tf.random.categorical(np.mat([1/FRAME_LEN for i in range(FRAME_LEN)]),mask_size))\n    # print(mask)\n    x = tf.tensor_scatter_nd_update(x,mask[...,None],tf.fill([mask_size,92,3],mask_value))\n    return x\n\ndef spatial_mask(x, size=(0.2,0.4), mask_value=float(0.0)):\n    # mask by a randome rectangle\n    mask_offset_y = tf.random.uniform(())\n    mask_offset_x = tf.random.uniform(())\n    mask_size = tf.random.uniform((), *size)\n    mask_x = (mask_offset_x<x[...,0]) & (x[...,0] < mask_offset_x + mask_size)\n    mask_y = (mask_offset_y<x[...,1]) & (x[...,1] < mask_offset_y + mask_size)\n    mask = mask_x & mask_y\n    x = tf.where(mask[...,None], mask_value, x)\n    return x\n\ndef augment_fn(x, always=False):\n    if tf.random.uniform(())<0.5 or always:\n        x = resample(x, (0.6,1.4))\n    if tf.random.uniform(())<0.5 or always:\n        x = flip_lr(x)\n    if tf.random.uniform(())<0.7 or always:\n        x = spatial_random_affine(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = temporal_mask(x)\n    if tf.random.uniform(())<0.5 or always:\n        x = spatial_mask(x)\n    return x\n\n@tf.function(jit_compile=True)\ndef pre_process0(x):\n    lip_x = tf.gather(x, LIP_IDX_X, axis=1)\n    lip_y = tf.gather(x, LIP_IDX_Y, axis=1)\n    lip_z = tf.gather(x, LIP_IDX_Z, axis=1)\n\n    rhand_x = tf.gather(x, RHAND_IDX_X, axis=1)\n    rhand_y = tf.gather(x, RHAND_IDX_Y, axis=1)\n    rhand_z = tf.gather(x, RHAND_IDX_Z, axis=1)\n    \n    lhand_x = tf.gather(x, LHAND_IDX_X, axis=1)\n    lhand_y = tf.gather(x, LHAND_IDX_Y, axis=1)\n    lhand_z = tf.gather(x, LHAND_IDX_Z, axis=1)\n\n    rpose_x = tf.gather(x, RPOSE_IDX_X, axis=1)\n    rpose_y = tf.gather(x, RPOSE_IDX_Y, axis=1)\n    rpose_z = tf.gather(x, RPOSE_IDX_Z, axis=1)\n    \n    lpose_x = tf.gather(x, LPOSE_IDX_X, axis=1)\n    lpose_y = tf.gather(x, LPOSE_IDX_Y, axis=1)\n    lpose_z = tf.gather(x, LPOSE_IDX_Z, axis=1)\n    \n    lip   = tf.concat([lip_x[..., tf.newaxis], lip_y[..., tf.newaxis], lip_z[..., tf.newaxis]], axis=-1)\n    rhand = tf.concat([rhand_x[..., tf.newaxis], rhand_y[..., tf.newaxis], rhand_z[..., tf.newaxis]], axis=-1)\n    lhand = tf.concat([lhand_x[..., tf.newaxis], lhand_y[..., tf.newaxis], lhand_z[..., tf.newaxis]], axis=-1)\n    rpose = tf.concat([rpose_x[..., tf.newaxis], rpose_y[..., tf.newaxis], rpose_z[..., tf.newaxis]], axis=-1)\n    lpose = tf.concat([lpose_x[..., tf.newaxis], lpose_y[..., tf.newaxis], lpose_z[..., tf.newaxis]], axis=-1)\n    \n    hand =  tf.concat([rhand, lhand], axis=1)\n    hand = tf.where(tf.math.is_nan(hand), 0.0, hand)\n    mask = tf.math.not_equal(tf.reduce_sum(hand, axis=[1, 2]), 0.0)\n\n    lip = lip[mask]\n    rhand = rhand[mask]\n    lhand = lhand[mask]\n    rpose = rpose[mask]\n    lpose = lpose[mask]\n\n    return lip, rhand,lhand,  rpose, lpose #lhand\n\n@tf.function()\ndef pre_process1(lip, rhand,lhand,  rpose, lpose, augment = False): #lhand,\n    lip   = (resize_pad(lip) - LIPM) / LIPS\n    rhand = (resize_pad(rhand) - RHM) / RHS\n    lhand = (resize_pad(lhand) - LHM) / LHS\n    rpose = (resize_pad(rpose) - RPM) / RPS\n    lpose = (resize_pad(lpose) - LPM) / LPS\n\n    x = tf.concat([lip, rhand, lhand, rpose, lpose], axis=1) #lhand,\n    x = tf.where(tf.math.is_nan(x), 0.0, x)\n    if augment:\n        x = augment_fn(x)\n    s = tf.shape(x)\n    x = tf.reshape(x, (s[0], s[1]*s[2]))\n    return x\n\npre0 = pre_process0(frames)\npre1 = pre_process1(*pre0)\nINPUT_SHAPE = list(pre1.shape)\nprint(INPUT_SHAPE)\n```\n\n**Some Useful References**\nhttps://www.kaggle.com/code/royalacecat/the-deeper-the-better\nhttps://www.kaggle.com/code/hebasaleh00/aslfr-eda-preprocessing\nhttps://www.kaggle.com/competitions/asl-fingerspelling/discussion/409438",
    "2426094": ""
  }
}