{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# How well can we do with just a fixed prediction?\n\nThe goal of this notebook is two-fold:\n1. Understanding the _baseline_ performance on the task: the normalised Levenshtein distance is unusual and I wanted to get a feel for what the minimum reasonable LB score is. **Understanding the metric is vital to understand how to build the best models.**\n2. Understanding the minimum possible submission. This is my first time using TFLite, and so making this baseline helped understand what steps are needed.\n\nThank you to @wonderingalice for their minimal submission notebook! https://www.kaggle.com/code/wonderingalice/working-sample-submission-and-inference","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras import layers, optimizers, constraints, regularizers\nimport numpy as np\nimport pandas as pd\nimport json\n\nprint(\"Tensorflow\", tf.__version__)\n!python --version\n\nbasedir = \"/kaggle/working/\"\nNUM_CHARACTERS = 59","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-27T19:33:08.289278Z","iopub.execute_input":"2023-05-27T19:33:08.289736Z","iopub.status.idle":"2023-05-27T19:33:09.414224Z","shell.execute_reply.started":"2023-05-27T19:33:08.289701Z","shell.execute_reply":"2023-05-27T19:33:09.412883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv('/kaggle/input/asl-fingerspelling/train.csv')\n# Dummy features: we don't actually use any features\nSEL_FEATURES = ['x_right_hand_0','y_right_hand_0']\n\nc2p = json.load(open('/kaggle/input/asl-fingerspelling/character_to_prediction_index.json', 'r'))\np2c = {p: c for c, p in c2p.items()}","metadata":{"execution":{"iopub.status.busy":"2023-05-27T19:33:09.417179Z","iopub.execute_input":"2023-05-27T19:33:09.418136Z","iopub.status.idle":"2023-05-27T19:33:09.551033Z","shell.execute_reply.started":"2023-05-27T19:33:09.418084Z","shell.execute_reply":"2023-05-27T19:33:09.549802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Finding the best constant prediction\n\n[Levenshtein distance](https://en.wikipedia.org/wiki/Levenshtein_distance) gives us the smallest number of insertions, deletions and **substitutions** possible between our predicted string as the actual string.\n\n**There's a subtle point here:** a substitution is the same as a deletion, so if we know the average string has 12 characters, it makes sense to predict 12 characters. Any character you predict which ends up in the actual string gets you a point, while any character you mis-predict gets you the same as if you made no prediction. \n\nThis is different to other metrics that incorporate recall, where over-predicting can harm you.\n\nTo find the \"average\" string, which has the shortest distance to all the strings in the dataset, we use a greedy algorithm. We start with the empty string, and repeatedly find the best single character to insert _anywhere_ in the string, until we can no longer improve the training score.","metadata":{}},{"cell_type":"code","source":"from Levenshtein import distance\nally = df_train['phrase'].values\ntotaly = sum([len(y) for y in ally])\n\n# Evaluate a constant prediction on the training set\ndef eval_string(s):\n    d = 0\n    for y in ally:\n        d += distance(s, y)\n    return (totaly - d) / totaly\n\n# Greedy algorithm\nbest_str = ''\nbest_score = 0\nchars = list(c2p.keys())\n\nfor i in range(20): # max length\n    inner_best = best_str\n    inner_best_score = best_score\n    \n    for position in range(len(best_str)+1): # at all insertion points\n        for newchar in chars:               # try all characters\n            new_str = best_str[:position] + str(newchar) + best_str[position:]\n            score = eval_string(new_str)\n\n            if score > inner_best_score:\n                inner_best = new_str\n                inner_best_score = score\n                print(f'New best @ {len(inner_best)}=\"{inner_best}\", score {inner_best_score:.4f}')\n\n    if best_score >= inner_best_score:\n        print('No improvement, best is', best_str)\n        break\n        \n    best_str = inner_best\n    best_score = inner_best_score\n    print(f'Best str @ {len(best_str)}=\"{best_str}\", score {best_score:.4f}')","metadata":{"execution":{"iopub.status.busy":"2023-05-27T20:06:55.167112Z","iopub.execute_input":"2023-05-27T20:06:55.167755Z","iopub.status.idle":"2023-05-27T20:06:58.884239Z","shell.execute_reply.started":"2023-05-27T20:06:55.167712Z","shell.execute_reply":"2023-05-27T20:06:58.881784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Turning this into a TFLite model\n\nI found that the path of least resistance here was to turn this constant prediction into a Keras model, and then convert that into TFLie. It's unclear to me what operations are allowed in TFLite models, but this provides a framework for embedding any code you might want into an arbitrary Keras layer.\n\nIt appears that Kaggle's evalution work like this:\n- Your model is run **once per video** in the test set (each video is one batch).\n- You receive an input of shape `(N_FRAMES, N_FEATURES)`. The normal \"time-series\" way of doing this would be `(1, N_FRAMES, N_FEATURES)`, so note this is different.\n- You return an output of shape `(N_CHARS, 59)` where 59 is the number of possible characters. `N_CHARS` is up to your model, in our case it's constant.\n- Evaluation is done on **the argmax of your predictions**, so it doesn't matter what the actual probability is","metadata":{}},{"cell_type":"code","source":"const_pred = np.zeros((len(best_str), 59))\nfor i, c in enumerate(best_str):\n    const_pred[i, c2p[c]] = 1","metadata":{"execution":{"iopub.status.busy":"2023-05-27T19:14:59.300965Z","iopub.execute_input":"2023-05-27T19:14:59.301872Z","iopub.status.idle":"2023-05-27T19:14:59.309479Z","shell.execute_reply.started":"2023-05-27T19:14:59.301824Z","shell.execute_reply":"2023-05-27T19:14:59.307687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the custom layer\nfrom tensorflow.keras.layers import Layer, Input\nclass ConstantLayer(Layer):\n    def __init__(self, constant_vector, name=None):\n        super(ConstantLayer, self).__init__(name=name)\n        self.constant_vector = tf.Variable(initial_value=constant_vector, trainable=False, dtype=tf.float32)\n\n    def call(self, inputs):\n        return self.constant_vector","metadata":{"execution":{"iopub.status.busy":"2023-05-27T19:15:14.368237Z","iopub.execute_input":"2023-05-27T19:15:14.368615Z","iopub.status.idle":"2023-05-27T19:15:14.377292Z","shell.execute_reply.started":"2023-05-27T19:15:14.368585Z","shell.execute_reply":"2023-05-27T19:15:14.375473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_layer = Input(shape=(len(SEL_FEATURES),), name='inputs')  # Let's assume we are inputting vectors of size 10\noutput_layer = ConstantLayer(const_pred, name='outputs')(input_layer)\n\nmodel = tf.keras.models.Model(inputs=input_layer, outputs=output_layer)","metadata":{"execution":{"iopub.status.busy":"2023-05-27T19:15:30.403492Z","iopub.execute_input":"2023-05-27T19:15:30.404011Z","iopub.status.idle":"2023-05-27T19:15:30.431153Z","shell.execute_reply.started":"2023-05-27T19:15:30.403971Z","shell.execute_reply":"2023-05-27T19:15:30.430131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tflite_model = tf.lite.TFLiteConverter.from_keras_model(model).convert()\nmodel_path = 'model.tflite'\n\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)\n\n!zip submission.zip  './model.tflite' './inference_args.json'","metadata":{"execution":{"iopub.status.busy":"2023-05-27T19:15:32.601137Z","iopub.execute_input":"2023-05-27T19:15:32.601576Z","iopub.status.idle":"2023-05-27T19:15:34.82182Z","shell.execute_reply.started":"2023-05-27T19:15:32.601541Z","shell.execute_reply":"2023-05-27T19:15:34.820086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Overall, we see 0.160 on CV, and 0.157 on LB - pretty consistent!","metadata":{}}]}