{
  "id": 435231,
  "title": "Some Notes",
  "url": "/competitions/asl-fingerspelling/discussion/435231",
  "author_name": "Scenery SunFireInk",
  "post_date": "2023-08-28T12:43:55.165000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many more skilled participants have shared very rich and impressive solutions. I'm only documenting a few experimental insights to avoid adding something redundant to the discussion.</p>\n<h1>LB History</h1>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>max_len=188</td>\n<td>0.770</td>\n<td>0.741</td>\n</tr>\n<tr>\n<td>max_len=320</td>\n<td>0.776</td>\n<td>0.744</td>\n</tr>\n<tr>\n<td>Post Process by Chris Deotte</td>\n<td>0.778</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>expand ratio 2-&gt;4</td>\n<td>0.779</td>\n<td>0.752</td>\n</tr>\n<tr>\n<td>epochs 126-&gt;200</td>\n<td>0.780</td>\n<td>0.752</td>\n</tr>\n</tbody>\n</table>\n<h1>What Didn't Work</h1>\n<ol>\n<li><p>I attempted various normalization techniques, including global normalization, local normalization, and normalization centered around the nose. However, I'm unsure if there were any experimental errors, as I found that these approaches merely expedited the training convergence without improving the leaderboard performance. I'm uncertain if my conclusion is accurate.</p></li>\n<li><p>Lowering the dropout of the final layer from 0.4 to 0.1 and removing dropout from all other layers would result in a decrease of 0.001 points in both the public test set and the private test set.</p></li>\n<li><p>To place the batch normalization (BN) before the depthwise convolution (DW Conv) instead of after would result in a decrease of 0.001 points for both the public LB (Leaderboard) and private LB (Leaderboard).</p></li>\n<li><p>Spatial Mask. It may lead to training instability when used. If other aspects are handled well (such as normalization), it might be usable.</p></li>\n</ol>",
  "messages": [
    {
      "id": 2412691,
      "postDate": "2023-08-28T12:43:55.167Z",
      "content": "<p>Many more skilled participants have shared very rich and impressive solutions. I'm only documenting a few experimental insights to avoid adding something redundant to the discussion.</p>\n<h1>LB History</h1>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>max_len=188</td>\n<td>0.770</td>\n<td>0.741</td>\n</tr>\n<tr>\n<td>max_len=320</td>\n<td>0.776</td>\n<td>0.744</td>\n</tr>\n<tr>\n<td>Post Process by Chris Deotte</td>\n<td>0.778</td>\n<td>0.747</td>\n</tr>\n<tr>\n<td>expand ratio 2-&gt;4</td>\n<td>0.779</td>\n<td>0.752</td>\n</tr>\n<tr>\n<td>epochs 126-&gt;200</td>\n<td>0.780</td>\n<td>0.752</td>\n</tr>\n</tbody>\n</table>\n<h1>What Didn't Work</h1>\n<ol>\n<li><p>I attempted various normalization techniques, including global normalization, local normalization, and normalization centered around the nose. However, I'm unsure if there were any experimental errors, as I found that these approaches merely expedited the training convergence without improving the leaderboard performance. I'm uncertain if my conclusion is accurate.</p></li>\n<li><p>Lowering the dropout of the final layer from 0.4 to 0.1 and removing dropout from all other layers would result in a decrease of 0.001 points in both the public test set and the private test set.</p></li>\n<li><p>To place the batch normalization (BN) before the depthwise convolution (DW Conv) instead of after would result in a decrease of 0.001 points for both the public LB (Leaderboard) and private LB (Leaderboard).</p></li>\n<li><p>Spatial Mask. It may lead to training instability when used. If other aspects are handled well (such as normalization), it might be usable.</p></li>\n</ol>",
      "rawMarkdown": "Many more skilled participants have shared very rich and impressive solutions. I'm only documenting a few experimental insights to avoid adding something redundant to the discussion.\n\n# LB History\n| | public LB | private LB|\n| ---  | --- | --- |\n| max_len=188 | 0.770| 0.741|\n| max_len=320 | 0.776| 0.744|\n| Post Process by Chris Deotte | 0.778| 0.747|\n| expand ratio 2->4 | 0.779| 0.752|\n| epochs 126->200 | 0.780| 0.752|\n\n# What Didn't Work\n1. I attempted various normalization techniques, including global normalization, local normalization, and normalization centered around the nose. However, I'm unsure if there were any experimental errors, as I found that these approaches merely expedited the training convergence without improving the leaderboard performance. I'm uncertain if my conclusion is accurate.\n\n2. Lowering the dropout of the final layer from 0.4 to 0.1 and removing dropout from all other layers would result in a decrease of 0.001 points in both the public test set and the private test set.\n\n3. To place the batch normalization (BN) before the depthwise convolution (DW Conv) instead of after would result in a decrease of 0.001 points for both the public LB (Leaderboard) and private LB (Leaderboard).\n\n4. Spatial Mask. It may lead to training instability when used. If other aspects are handled well (such as normalization), it might be usable.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2412691": "Many more skilled participants have shared very rich and impressive solutions. I'm only documenting a few experimental insights to avoid adding something redundant to the discussion.\n\n# LB History\n| | public LB | private LB|\n| ---  | --- | --- |\n| max_len=188 | 0.770| 0.741|\n| max_len=320 | 0.776| 0.744|\n| Post Process by Chris Deotte | 0.778| 0.747|\n| expand ratio 2->4 | 0.779| 0.752|\n| epochs 126->200 | 0.780| 0.752|\n\n# What Didn't Work\n1. I attempted various normalization techniques, including global normalization, local normalization, and normalization centered around the nose. However, I'm unsure if there were any experimental errors, as I found that these approaches merely expedited the training convergence without improving the leaderboard performance. I'm uncertain if my conclusion is accurate.\n\n2. Lowering the dropout of the final layer from 0.4 to 0.1 and removing dropout from all other layers would result in a decrease of 0.001 points in both the public test set and the private test set.\n\n3. To place the batch normalization (BN) before the depthwise convolution (DW Conv) instead of after would result in a decrease of 0.001 points for both the public LB (Leaderboard) and private LB (Leaderboard).\n\n4. Spatial Mask. It may lead to training instability when used. If other aspects are handled well (such as normalization), it might be usable."
  }
}