{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 日本語も英語もできる方へ\n- 筆者の英語力は壊滅的で、この文章も機械翻訳の力を借りて用意しています。明らかな間違いや「こっちの表現の方が良いと思う」といった部分が沢山あると思うので、その様な部分があればコメントで教えてください。\n\n# 概要とか | summary\n- このnotebookは、[mayo clinic strip aiコンペ](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/overview)で最終29位を取った解法に関するnotebookです。(ただ、実際にこれをSubmitしたわけではないです)\n\n- 最終提出物には\n```\n1: Public LBで0.4を出したやつ(Private LBで0.7)\n2: 精度を落として(Public LBで0.6)安定感を高めたやつ(Private LBで0.6)\n``` \nの2種類あるのですが、このnotebookは2に関するものです。\n\n- This notebook is about a solution method that took 29th place in [mayo clinic strip ai](https://www.kaggle.com/competitions/mayo-clinic-strip-ai/overview)\n- There are two types of methods, \n```\n1: scored 0.4 in Public LB(this method scored 0.7 in Private LB)\n2: scored 0.6 in Public LB(this method scored 0.6 in Private LB)\n```\nThis document deals with 2.\n\n# モデル | models\n- EfficientNetベースのモデルを16個使いました。\n- モデルはここに置いてあります: https://www.kaggle.com/datasets/yuahyodo/mayo-sub12-models\n\n- I used 16 EfficientNet-based models.\n- models are here: https://www.kaggle.com/datasets/yuahyodo/mayo-sub12-models\n\n# 訓練データ | train data\n- 患者1人につき画像1枚になるように、同じ患者の2枚目以降は削除しています。\n- 訓練データは回転して4倍に水増ししています。\n- 「CEとLAA、どっちも分類できないとスコアが上がらない」という事が分かっていたので、割合が低いLAAの画像のみミラー反転を行ってさらに増やしています。\n- To ensure that there is only one image per patient, the second and subsequent images of the same patient will be deleted.\n- The image is rotated and the number of images is padded by a factor of 4.\n- LAA's images are padded by mirror inversion due to the small number of images.\n\n# テストデータ | valudation data\n- テストデータの割合は全体の10%で、水増し処理の前に分割されます。\n- Test data accounts for 10% of the total and is split before the padding process.\n\n# 損失関数・最適化関数 | loss and optimizer\n- 以下のような設定です(学習率については後述)\n- set up as follows(see below for learning rate)\n```\nloss='categorical_crossentropy'\noptimizer=Adam(learning_rate=LearningRate)\nmetrics='acc'\n```\n\n# 訓練の流れ | training flow\n- 訓練は2段階に分かれています。\n- 1段階目は学習済みEfficientNetの重みを固定して、出力付近に追加した層のみを学習させます。\n- 2段階目は全体を学習させます。\n- 1段階目は0.000001、2段階目はというかなり低い学習率で学習させました。\n- 学習の停止はEarlyStoppingで決めています。具体的には、\n```\n1段階目: EarlyStopping(monitor='val_loss', min_delta=0.0001, restore_best_weights=True, patience=10)\n2段階目: EarlyStopping(monitor='val_loss', min_delta=0.00001, restore_best_weights=True, patience=10)\n```\nという設定にしています。\n\n- Training is divided into two phases.\n- The first stage fixes the weights of the learned model and trains only the layers added near the output layer.\n- In the second step, all weights are learned.\n- The first stage had a learning rate of 0.000001 and the second stage had a learning rate of 0.0000001.\n- Learning stops are determined by EarlyStopping.The settings are as follows\n```\nfirst stage: EarlyStopping(monitor='val_loss', min_delta=0.0001, restore_best_weights=True, patience=10)\nsecond stage: EarlyStopping(monitor='val_loss', min_delta=0.00001, restore_best_weights=True, patience=10)\n```","metadata":{}},{"cell_type":"markdown","source":"# 提出用コード | code for submission","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\n\nfrom tensorflow.keras.applications.efficientnet import preprocess_input\nfrom tensorflow.keras.models import load_model\n\nfrom tifffile import TiffFile as TIF \nfrom PIL import Image, ImageOps \nimport sys \nimport gc\n\nimg_size = (224, 224, 3)\ndire = '../input/mayo-sub12-models/'\nmodel_list = [dire + i for i in os.listdir(dire) if '.h5' in i]\ncsv_files_list = ['model' + str(i) + '.csv' for i in range(len(model_list))]\n\ndef make_jpg_imgs(df, out_dire, imgs_dire='../input/mayo-clinic-strip-ai/test/'):\n    global img_size\n    gc.collect()\n    os.makedirs(out_dire, exist_ok=True)\n    files = [imgs_dire + i + '.tif' for i in df['image_id']]\n    output_files = [out_dire + i + '.jpg' for i in df['image_id']]\n    del df\n    gc.collect()\n    for i in range(len(files)):\n        with TIF(files[i]) as t:\n            data = t.asarray(out='memmap')\n        del t\n        gc.collect()\n        \n        if data.shape[0] >= 50000:\n            a = data.shape[0] - 50000\n            a = a // 2\n            data = data[a:a+50000, :, :]\n            gc.collect()\n            if data.shape[1] >= 35000:\n                a = data.shape[1] - 35000\n                a = a // 2\n                data = data[:, a:a+35000, :]\n                gc.collect()\n        if data.shape[1] >= 50000:\n            a = data.shape[1] - 50000\n            a = a // 2\n            data = data[:, a:a+50000, :]\n            gc.collect()\n            if data.shape[0] >= 35000:\n                a = data.shape[0] - 35000\n                a = a // 2\n                data = data[a:a+35000, :, :]\n                gc.collect()\n        data = np.asarray(data.astype(np.uint8))\n        gc.collect()\n        img = Image.fromarray(data)\n        del data\n        gc.collect()\n        img = img.resize((img_size[1], img_size[0]))\n        gc.collect()\n        img.save(output_files[i])\n        del img\n        gc.collect()\n    del files, output_files\n    gc.collect()\n    return\n\ndef load_JPG(df, dire):\n    gc.collect()\n    file = dire + df['image_id'] + '.jpg'\n    img = Image.open(file).resize((img_size[1], img_size[0]))\n    output = np.empty((4, img_size[0], img_size[1], img_size[2]), dtype=np.float32)\n    R = (0, 90, 180, 270)\n    for i in range(len(R)):\n        output[i] = np.array(img.rotate(R[i]))\n    output = preprocess_input(output)\n    gc.collect()\n    return output    \n\ndire = '../input/mayo-clinic-strip-ai/'\nout_dire = '../working/mayo-clinic-test-JPG/'\ntest_df = pd.read_csv(dire + 'test.csv')\nmake_jpg_imgs(test_df, out_dire)\n\ntest_df = test_df.drop_duplicates(subset=['patient_id'])\n\n\nfor model_num in range(len(model_list)):\n    model_file = model_list[model_num]\n    NN = load_model(model_file)\n    X_CE = []\n    X_LAA = []\n    for i in range(len(test_df)):\n        X = load_JPG(test_df.iloc[i], out_dire)\n        pred = NN.predict(X, batch_size=1)\n        del X\n        Y = [0, 0]\n        for B in range(len(pred)):\n            Y[0] += pred[B][0]\n            Y[1] += pred[B][1]\n        Y[0] /= len(pred)\n        Y[1] /= len(pred)\n        gc.collect()\n        X_CE.append(Y[0])\n        X_LAA.append(Y[1])\n        del pred\n        gc.collect()\n    model_submit = pd.read_csv(dire + 'sample_submission.csv')\n    model_submit['CE'] = list(X_CE)\n    model_submit['LAA'] = list(X_LAA)\n    model_submit.to_csv(csv_files_list[model_num])\n\nsubmit = pd.read_csv(dire + 'sample_submission.csv')\nCE = np.zeros(len(submit))\nLAA = np.zeros(len(submit))\nfor i in range(len(csv_files_list)):\n    df = pd.read_csv(csv_files_list[i])\n    CE += (df['CE'] / len(csv_files_list))\n    LAA += (df['LAA'] / len(csv_files_list))\nsubmit['CE'] = list(CE)\nsubmit['LAA'] = list(LAA)\noutput = 'submission.csv'\nsubmit.to_csv(output, index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 色々 | various\n- 以降は作者が書きたいことを書くだけの部分なので、あなたの貴重な時間を削ってまで読む部分ではないです。\n- After that, it's just the part where the author writes what he wants to write, so it's not the part where you have to spend your precious time to read it!\n\n## 使った計算資源\n- ローカルにもGPUがある(ただし、GPUメモリが4GBしかないのですぐ溢れる)ので、序盤はそれを使って軽い実験をしていました。\n- その後、「Kaggle notebookはGPUとTPUが無料でたくさん使えて便利」という事に気づいて、中盤以降はそれを利用するようになりました。\n- 「某G社のタダでGPUとかTPUが使えるサービスが使いにくくなったので、戦えない」との意見も見られますが、私は少なくともそれは感じませんでした。\n\n## 私が大規模なShakeを乗り越えられた理由\n- 私はPublic LBで銀メダル圏内に入る前には既に、1: Public LBの計算に使われるデータが非常に少ないこと、2: 医療系のデータなのでバラツキが激しい可能性が高い事、3: 当時の金メダル圏内にはExpert以下の称号の方しか居ないこと、などから「大規模なShakeが起きそう」と予想はしていました。\n- その後、Public LBで0.4を出した時「とても嬉しいが、スコアの上がり方が明らかにおかしい。これはマグレか過学習なのではないか？」と考え、「順位が暴落する可能性が非常に高いので、安定感のある選択肢も用意するべき」との結論に至り最終日の分の権利を使って「安全な選択肢」を用意しました。\n- 最終提出物を選ぶ際、「Public LBでのスコアが良い2つのモデルを選びたい」という自分に「ここで惑わされると痛い目を見る」と言い聞かせて「安全な選択肢」を2つの内の1つに選びました。\n- 予想は的中し、「安全な選択肢」を選んでいたおかげで何とか助かった。\n\n## 今後の抱負\n- とりあえずExpertを目指したいです。\n- ただ、しばらくは機械学習の理論の勉強と休養と将棋AIやオセロAIの開発のためにKaggleはお休みしようかと考えています。","metadata":{}}]}