{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# quickdraw-doodle-recognition\n\nHello, everyone. This competition is a level3 image classification. \n\n**Dataset's name:description(size, count of train data, count of test data)**  \nLevel 1 : **MNIST** : pictures of 0~9, grey(28\\*28, 60000, 10000)  \nLevel 2 : **CIFAR-10** : pictures of 10 objects, rgb(32\\*32, 50000, 10000)  \nLevel 3 : **CIFAR-100** : pictures of 100 objects, rgb(32\\*32, 50000, 10000)  \nLevel 4 : **ImageNet** : pictures of 1000 objects, rgb(224\\*224)\n(I set this level by my criteria.)\n\nIn image classification, generally, people go through the following steps:  \n1. Check data  \n    1-1. Check data's size  \n    1-2. Draw image(when data is not an image)  \n    1-3. Check image's info  \n2. Construct Model  \n    2-1. VGG, ResNet, GoogleNet, your own model etc...  \n    2-2. Set parameters  \n    2-3. optimizers, annealing  \n3. Training  \n4. Predict test data  \n    4-1. If you don't satisfy, go to step2.  \n5. Make submissions  \n\nThen let's start."},{"metadata":{"_cell_guid":"","_uuid":"","trusted":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport tensorflow as tf\nimport json\nimport cv2\nimport matplotlib.pyplot as plt\nimport datetime as dt\nfrom tqdm import tqdm\nfrom tensorflow import keras\nfrom tensorflow.keras.layers import Conv2D, MaxPooling2D\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, Activation, BatchNormalization\nfrom tensorflow.keras.metrics import categorical_accuracy, top_k_categorical_accuracy, categorical_crossentropy\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.applications import MobileNet\nfrom tensorflow.keras.applications.mobilenet import preprocess_input","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Hyper Parameter"},{"metadata":{"trusted":true},"cell_type":"code","source":"DP_DIR = '../input/shuffle-csvs/'\nINPUT_DIR = '../input/quickdraw-doodle-recognition/'\nNCSVS = 100\nNCATS = 340\nBASE_SIZE = 256\nsize = 64\nepochs = 30\nbatch_size = 100\nstart = dt.datetime.now()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 1. Check data\n\n## 1-1. Check data's size"},{"metadata":{},"cell_type":"markdown","source":"It's too bigger than memory given to us(140GB >> 13GB). So, we will use other data.  \n<https://www.kaggle.com/gaborfodor/shuffle-csvs>  \nIn this kernel, gaborfodor makes shuffle-csvs. Each of csv files includes all kinds of pictures. A file is fully available within the memory given to us.  "},{"metadata":{},"cell_type":"markdown","source":"## 1-2. Draw Image"},{"metadata":{},"cell_type":"markdown","source":"Draw image with line function in cv2 module."},{"metadata":{"trusted":true},"cell_type":"code","source":"def draw_img(lines):\n    img = np.zeros((BASE_SIZE, BASE_SIZE))\n    for line in lines:\n        for i in range(len(line[0]) - 1):\n            _ = cv2.line(img, (line[0][i], line[1][i]), (line[0][i + 1], line[1][i + 1]), 255, 6)\n    return cv2.resize(img, (size, size))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Make image randomly."},{"metadata":{"trusted":true},"cell_type":"code","source":"def image_gen(batchsize, cnt):\n    while True:\n        for k in np.random.permutation(cnt):\n            filename = os.path.join(DP_DIR, 'train_k{}.csv.gz'.format(k))\n            for df in pd.read_csv(filename, chunksize=batchsize):\n                df['drawing'] = df['drawing'].apply(json.loads)\n                x = np.zeros((len(df), size, size, 1))\n                for i, lines in enumerate(df.drawing.values):\n                    x[i, :, :, 0] = draw_img(lines)\n                    \n                x = preprocess_input(x).astype(np.float32)\n                y = keras.utils.to_categorical(df.y, num_classes=NCATS)\n                \n                yield x, y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def df_to_image(df):\n    df['drawing'] = df['drawing'].apply(json.loads)\n    x = np.zeros((len(df), size, size, 1))\n    for i, lines in enumerate(df.drawing.values):\n        x[i, :, :, 0] = draw_img(lines)\n    x = preprocess_input(x).astype(np.float32)\n    return x","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This code is based on [this kernel](https://www.kaggle.com/gaborfodor/greyscale-mobilenet-lb-0-892)."},{"metadata":{"trusted":true},"cell_type":"code","source":"train_datagen = image_gen(batch_size, range(NCSVS - 1))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Check image's info"},{"metadata":{},"cell_type":"markdown","source":"Make map with number and category's name."},{"metadata":{"trusted":true},"cell_type":"code","source":"files = sorted(os.listdir('../input/quickdraw-doodle-recognition/train_simplified/'), reverse=False, key=str.lower)\nclass_dict = {file[:-4].replace(\" \", \"_\"): i for i, file in enumerate(files)}\nclassreverse_dict = {v: k for k, v in class_dict.items()}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 2. Construct Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"def CNN_model():\n    model = MobileNet(input_shape=(size, size, 1), alpha=1., weights=None, classes=NCATS)\n    \n#     My Own Model\n#     model = Sequential()\n\n#     model.add(Conv2D(32,kernel_size=3,activation='relu',padding='same',input_shape=(size,size,1)))\n#     model.add(BatchNormalization())\n#     model.add(Conv2D(32,kernel_size=3,activation='relu', padding='same'))\n#     model.add(BatchNormalization())\n#     model.add(Conv2D(32,kernel_size=5,strides=2,padding='same',activation='relu'))\n#     model.add(BatchNormalization())\n#     model.add(Dropout(0.4))\n\n#     model.add(Conv2D(64,kernel_size=3,activation='relu', padding='same'))\n#     model.add(BatchNormalization())\n#     model.add(Conv2D(64,kernel_size=3,activation='relu', padding='same'))\n#     model.add(BatchNormalization())\n#     model.add(Conv2D(64,kernel_size=5,strides=2,padding='same',activation='relu'))\n#     model.add(BatchNormalization())\n#     model.add(Dropout(0.4))\n\n#     model.add(Flatten())\n#     model.add(Dense(2 * NCATS, activation='relu'))\n#     model.add(BatchNormalization())\n#     model.add(Dropout(0.4))\n#     model.add(Dense(NCATS, activation='softmax'))\n\n    model.summary()\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":""},{"metadata":{"trusted":true},"cell_type":"code","source":"def top_3_accuracy(y_true, y_pred):\n    return top_k_categorical_accuracy(y_true, y_pred, k=3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"Use Adam optimizers."},{"metadata":{"trusted":true},"cell_type":"code","source":"model = CNN_model()\n\nmodel.compile(optimizer=Adam(lr=0.0024), loss='categorical_crossentropy', metrics=[categorical_crossentropy, categorical_accuracy, top_3_accuracy])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## annealing"},{"metadata":{"trusted":true},"cell_type":"code","source":"callbacks = [ReduceLROnPlateau(monitor='val_acc', patience=3, verbose=1, factor=0.5, min_lr=0.00001)]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 3. Training"},{"metadata":{},"cell_type":"markdown","source":"Make validation data."},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_df = pd.read_csv(os.path.join(DP_DIR, 'train_k{}.csv.gz'.format(NCSVS - 1)), nrows=34000)\nx_valid = df_to_image(valid_df)\ny_valid = keras.utils.to_categorical(valid_df.y, num_classes=NCATS)\nprint(x_valid.shape, y_valid.shape)\nprint('Validation array memory {:.2f} GB'.format(x_valid.nbytes / 1024.**3 ))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit_generator(train_datagen, epochs = epochs, verbose = 1, \n                              validation_data=(x_valid, y_valid),\n                              steps_per_epoch=x_valid.shape[0] // batch_size, callbacks=callbacks)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 4. Predict test data"},{"metadata":{"trusted":true},"cell_type":"code","source":"test = pd.read_csv(os.path.join(INPUT_DIR, 'test_simplified.csv'))\ntest.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x_test = df_to_image(test)\nprint(test.shape, x_test.shape)\nprint('Test array memory {:.2f} GB'.format(x_test.nbytes / 1024.**3 ))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_predictions = model.predict(x_test, batch_size=batch_size)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 5. Make submissions"},{"metadata":{},"cell_type":"markdown","source":"Select top3 category."},{"metadata":{"trusted":true},"cell_type":"code","source":"top3 = pd.DataFrame(np.argsort(-test_predictions, axis=1)[:, :3])\ntop3.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Change number to category's name, and submit submissions."},{"metadata":{"trusted":true},"cell_type":"code","source":"word = top3.replace(classreverse_dict)\ntest['word'] = word[0] + ' ' + word[1] + ' ' + word[2]\nsubmission = test[['key_id', 'word']]\nsubmission.to_csv('submission.csv', index=False)\nsubmission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"end = dt.datetime.now()\nprint('Latest run {}.\\nTotal time {}s'.format(end, (end - start).seconds))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}