{
  "id": 391055,
  "title": "Multi-GPU Inference with TensorRT in Kaggle Notebooks",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391055",
  "author_name": "tmyok",
  "post_date": "2023-02-28T08:53:40.389000",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I would like to thank the competition organizers, the Kaggle team, and all the participants. I have learned a lot from this competition. </p>\n<p>In this topic, I will share how to run multiple GPUs simultaneously using Python's \"threading\" and \"queue\" modules.</p>\n<p>In Kaggle Notebook, we can use two Tesla T4 GPUs. Moreover, since T4 GPUs are equipped with units of Tensor Cores, we can accelerate inference using TensorRT (discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375881\" target=\"_blank\">here</a>). However, TensorRT does not seem to be able to perform multi-GPU processing like PyTorch's DataParallel (<a href=\"https://docs.nvidia.com/deeplearning/tensorrt/developer-guide/index.html#faq\" target=\"_blank\">reference</a>). I wanted to efficiently use TensorRT models on two GPUs simultaneously.</p>\n<p>Here is the notebook that we used for submission.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm</a></p>\n<p>Below is a summary of the process</p>\n<p>Implement pre-processing, GPU0 processing, and GPU1 processing, respectively.</p>\n<pre><code>def preprocess(args, params):\n    return result\n\ndef GPU0_process(args, params):\n    return result\n\ndef GPU1_process(args, params):\n    return result\n</code></pre>\n<p>Implement queues and threads.</p>\n<pre><code>import threading\nimport queue\n\ndef wrap_func_for_mt(func, params):\n    def wrap_func(queue_input, queue_output):\n        while True:\n            input = queue_input.get()\n            if input is None:\n                queue_output.put(None)\n                continue\n\n            result = func(input, params)\n\n            queue_output.put(result)\n\n    return wrap_func\n\ndef prepare_multithreading(params):\n\n    # funcs to proc in pipeline\n    func_params = [\n        (preprocess, (params)),\n        (GPU0_process, (params)),\n        (GPU1_process, (params)),\n    ]\n    wrap_funcs = list(map(lambda func_param: wrap_func_for_mt(func_param[0], func_param[1]), func_params))\n\n    # prepare queues\n    queues_input = [queue.Queue() for _ in range(len(wrap_funcs))]\n    queues_output = [queue.Queue() for _ in range(len(wrap_funcs))]\n\n    # create Threads\n    threads = []\n    for wrap_func, queue_input, queue_output in zip(wrap_funcs, queues_input, queues_output):\n        t = threading.Thread(target=wrap_func, args=(queue_input, queue_output), daemon=True)\n        threads.append(t)\n\n    for t in threads:\n        t.start()\n\n    return queues_input, queues_output, len(wrap_funcs)\n\ndef loop_proc(queues_input, queues_output, inputs):\n    for queue_input, input in zip(queues_input, inputs):\n        queue_input.put(input)\n\n    outputs = []\n    for queue_output in queues_output:\n        output = queue_output.get()\n        outputs.append(output)\n\n    return outputs\n</code></pre>\n<p>Execute the process.</p>\n<pre><code>queues_input, queues_output, len_wrap_funcs = prepare_multithreading(params)\nwhile len(prediction_id_list) &lt; len(prediction_ids):\n\n    if idx &gt;= len(prediction_ids):\n        args = None\n    else:\n        args = {\n            \"prediction_id\": prediction_ids[idx],\n        }\n\n    if idx == 0:\n        init_inputs = [args] + [None]*(len_wrap_funcs - 1)  # [[], None, None, ...]\n        inputs = init_inputs\n    else:\n        inputs = [args] + outputs[:-1]\n\n    outputs = loop_proc(queues_input, queues_output, inputs)\n    result = outputs[-1]\n\n    if result is not None:\n        prediction_id_list.append(result[\"prediction_id\"])\n        cancer_list.append(result[\"cancer\"])\n\n    idx = idx + 1\n</code></pre>\n<p>Since three processes run simultaneously, ideally, the process that takes the longest time will be hidden by other processes. By using this technique, it is expected that inference on T4 GPU can be executed more efficiently. If you have any other good ideas, please let me know.</p>",
  "messages": [
    {
      "id": 2162512,
      "postDate": "2023-02-28T08:53:40.390Z",
      "content": "<p>First of all, I would like to thank the competition organizers, the Kaggle team, and all the participants. I have learned a lot from this competition. </p>\n<p>In this topic, I will share how to run multiple GPUs simultaneously using Python's \"threading\" and \"queue\" modules.</p>\n<p>In Kaggle Notebook, we can use two Tesla T4 GPUs. Moreover, since T4 GPUs are equipped with units of Tensor Cores, we can accelerate inference using TensorRT (discussed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375881\" target=\"_blank\">here</a>). However, TensorRT does not seem to be able to perform multi-GPU processing like PyTorch's DataParallel (<a href=\"https://docs.nvidia.com/deeplearning/tensorrt/developer-guide/index.html#faq\" target=\"_blank\">reference</a>). I wanted to efficiently use TensorRT models on two GPUs simultaneously.</p>\n<p>Here is the notebook that we used for submission.</p>\n<p><a href=\"https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm\" target=\"_blank\">https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm</a></p>\n<p>Below is a summary of the process</p>\n<p>Implement pre-processing, GPU0 processing, and GPU1 processing, respectively.</p>\n<pre><code>def preprocess(args, params):\n    return result\n\ndef GPU0_process(args, params):\n    return result\n\ndef GPU1_process(args, params):\n    return result\n</code></pre>\n<p>Implement queues and threads.</p>\n<pre><code>import threading\nimport queue\n\ndef wrap_func_for_mt(func, params):\n    def wrap_func(queue_input, queue_output):\n        while True:\n            input = queue_input.get()\n            if input is None:\n                queue_output.put(None)\n                continue\n\n            result = func(input, params)\n\n            queue_output.put(result)\n\n    return wrap_func\n\ndef prepare_multithreading(params):\n\n    # funcs to proc in pipeline\n    func_params = [\n        (preprocess, (params)),\n        (GPU0_process, (params)),\n        (GPU1_process, (params)),\n    ]\n    wrap_funcs = list(map(lambda func_param: wrap_func_for_mt(func_param[0], func_param[1]), func_params))\n\n    # prepare queues\n    queues_input = [queue.Queue() for _ in range(len(wrap_funcs))]\n    queues_output = [queue.Queue() for _ in range(len(wrap_funcs))]\n\n    # create Threads\n    threads = []\n    for wrap_func, queue_input, queue_output in zip(wrap_funcs, queues_input, queues_output):\n        t = threading.Thread(target=wrap_func, args=(queue_input, queue_output), daemon=True)\n        threads.append(t)\n\n    for t in threads:\n        t.start()\n\n    return queues_input, queues_output, len(wrap_funcs)\n\ndef loop_proc(queues_input, queues_output, inputs):\n    for queue_input, input in zip(queues_input, inputs):\n        queue_input.put(input)\n\n    outputs = []\n    for queue_output in queues_output:\n        output = queue_output.get()\n        outputs.append(output)\n\n    return outputs\n</code></pre>\n<p>Execute the process.</p>\n<pre><code>queues_input, queues_output, len_wrap_funcs = prepare_multithreading(params)\nwhile len(prediction_id_list) &lt; len(prediction_ids):\n\n    if idx &gt;= len(prediction_ids):\n        args = None\n    else:\n        args = {\n            \"prediction_id\": prediction_ids[idx],\n        }\n\n    if idx == 0:\n        init_inputs = [args] + [None]*(len_wrap_funcs - 1)  # [[], None, None, ...]\n        inputs = init_inputs\n    else:\n        inputs = [args] + outputs[:-1]\n\n    outputs = loop_proc(queues_input, queues_output, inputs)\n    result = outputs[-1]\n\n    if result is not None:\n        prediction_id_list.append(result[\"prediction_id\"])\n        cancer_list.append(result[\"cancer\"])\n\n    idx = idx + 1\n</code></pre>\n<p>Since three processes run simultaneously, ideally, the process that takes the longest time will be hidden by other processes. By using this technique, it is expected that inference on T4 GPU can be executed more efficiently. If you have any other good ideas, please let me know.</p>",
      "rawMarkdown": "First of all, I would like to thank the competition organizers, the Kaggle team, and all the participants. I have learned a lot from this competition. \n\nIn this topic, I will share how to run multiple GPUs simultaneously using Python's \"threading\" and \"queue\" modules.\n\nIn Kaggle Notebook, we can use two Tesla T4 GPUs. Moreover, since T4 GPUs are equipped with units of Tensor Cores, we can accelerate inference using TensorRT (discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375881)). However, TensorRT does not seem to be able to perform multi-GPU processing like PyTorch's DataParallel ([reference](https://docs.nvidia.com/deeplearning/tensorrt/developer-guide/index.html#faq)). I wanted to efficiently use TensorRT models on two GPUs simultaneously.\n\nHere is the notebook that we used for submission.\n\n[https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm](https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm)\n\nBelow is a summary of the process\n\nImplement pre-processing, GPU0 processing, and GPU1 processing, respectively.\n```\ndef preprocess(args, params):\n    return result\n\ndef GPU0_process(args, params):\n    return result\n\ndef GPU1_process(args, params):\n    return result\n```\n\nImplement queues and threads.\n```\nimport threading\nimport queue\n\ndef wrap_func_for_mt(func, params):\n    def wrap_func(queue_input, queue_output):\n        while True:\n            input = queue_input.get()\n            if input is None:\n                queue_output.put(None)\n                continue\n\n            result = func(input, params)\n\n            queue_output.put(result)\n\n    return wrap_func\n\ndef prepare_multithreading(params):\n\n    # funcs to proc in pipeline\n    func_params = [\n        (preprocess, (params)),\n        (GPU0_process, (params)),\n        (GPU1_process, (params)),\n    ]\n    wrap_funcs = list(map(lambda func_param: wrap_func_for_mt(func_param[0], func_param[1]), func_params))\n\n    # prepare queues\n    queues_input = [queue.Queue() for _ in range(len(wrap_funcs))]\n    queues_output = [queue.Queue() for _ in range(len(wrap_funcs))]\n\n    # create Threads\n    threads = []\n    for wrap_func, queue_input, queue_output in zip(wrap_funcs, queues_input, queues_output):\n        t = threading.Thread(target=wrap_func, args=(queue_input, queue_output), daemon=True)\n        threads.append(t)\n\n    for t in threads:\n        t.start()\n\n    return queues_input, queues_output, len(wrap_funcs)\n\ndef loop_proc(queues_input, queues_output, inputs):\n    for queue_input, input in zip(queues_input, inputs):\n        queue_input.put(input)\n\n    outputs = []\n    for queue_output in queues_output:\n        output = queue_output.get()\n        outputs.append(output)\n\n    return outputs\n```\n\nExecute the process.\n```\nqueues_input, queues_output, len_wrap_funcs = prepare_multithreading(params)\nwhile len(prediction_id_list) < len(prediction_ids):\n\n    if idx >= len(prediction_ids):\n        args = None\n    else:\n        args = {\n            \"prediction_id\": prediction_ids[idx],\n        }\n\n    if idx == 0:\n        init_inputs = [args] + [None]*(len_wrap_funcs - 1)  # [[], None, None, ...]\n        inputs = init_inputs\n    else:\n        inputs = [args] + outputs[:-1]\n\n    outputs = loop_proc(queues_input, queues_output, inputs)\n    result = outputs[-1]\n\n    if result is not None:\n        prediction_id_list.append(result[\"prediction_id\"])\n        cancer_list.append(result[\"cancer\"])\n\n    idx = idx + 1\n\n```\n\nSince three processes run simultaneously, ideally, the process that takes the longest time will be hidden by other processes. By using this technique, it is expected that inference on T4 GPU can be executed more efficiently. If you have any other good ideas, please let me know.",
      "votes": 10
    },
    {
      "id": 2162599,
      "postDate": "2023-02-28T10:01:47.537Z",
      "content": "<p>Nice grounding work and great notebook!</p>",
      "rawMarkdown": "Nice grounding work and great notebook!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2162599,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T10:01:47.537000",
      "content": "<p>Nice grounding work and great notebook!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2162512": "First of all, I would like to thank the competition organizers, the Kaggle team, and all the participants. I have learned a lot from this competition. \n\nIn this topic, I will share how to run multiple GPUs simultaneously using Python's \"threading\" and \"queue\" modules.\n\nIn Kaggle Notebook, we can use two Tesla T4 GPUs. Moreover, since T4 GPUs are equipped with units of Tensor Cores, we can accelerate inference using TensorRT (discussed [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/375881)). However, TensorRT does not seem to be able to perform multi-GPU processing like PyTorch's DataParallel ([reference](https://docs.nvidia.com/deeplearning/tensorrt/developer-guide/index.html#faq)). I wanted to efficiently use TensorRT models on two GPUs simultaneously.\n\nHere is the notebook that we used for submission.\n\n[https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm](https://www.kaggle.com/code/tmyok1984/rsna-inference-lightgbm)\n\nBelow is a summary of the process\n\nImplement pre-processing, GPU0 processing, and GPU1 processing, respectively.\n```\ndef preprocess(args, params):\n    return result\n\ndef GPU0_process(args, params):\n    return result\n\ndef GPU1_process(args, params):\n    return result\n```\n\nImplement queues and threads.\n```\nimport threading\nimport queue\n\ndef wrap_func_for_mt(func, params):\n    def wrap_func(queue_input, queue_output):\n        while True:\n            input = queue_input.get()\n            if input is None:\n                queue_output.put(None)\n                continue\n\n            result = func(input, params)\n\n            queue_output.put(result)\n\n    return wrap_func\n\ndef prepare_multithreading(params):\n\n    # funcs to proc in pipeline\n    func_params = [\n        (preprocess, (params)),\n        (GPU0_process, (params)),\n        (GPU1_process, (params)),\n    ]\n    wrap_funcs = list(map(lambda func_param: wrap_func_for_mt(func_param[0], func_param[1]), func_params))\n\n    # prepare queues\n    queues_input = [queue.Queue() for _ in range(len(wrap_funcs))]\n    queues_output = [queue.Queue() for _ in range(len(wrap_funcs))]\n\n    # create Threads\n    threads = []\n    for wrap_func, queue_input, queue_output in zip(wrap_funcs, queues_input, queues_output):\n        t = threading.Thread(target=wrap_func, args=(queue_input, queue_output), daemon=True)\n        threads.append(t)\n\n    for t in threads:\n        t.start()\n\n    return queues_input, queues_output, len(wrap_funcs)\n\ndef loop_proc(queues_input, queues_output, inputs):\n    for queue_input, input in zip(queues_input, inputs):\n        queue_input.put(input)\n\n    outputs = []\n    for queue_output in queues_output:\n        output = queue_output.get()\n        outputs.append(output)\n\n    return outputs\n```\n\nExecute the process.\n```\nqueues_input, queues_output, len_wrap_funcs = prepare_multithreading(params)\nwhile len(prediction_id_list) < len(prediction_ids):\n\n    if idx >= len(prediction_ids):\n        args = None\n    else:\n        args = {\n            \"prediction_id\": prediction_ids[idx],\n        }\n\n    if idx == 0:\n        init_inputs = [args] + [None]*(len_wrap_funcs - 1)  # [[], None, None, ...]\n        inputs = init_inputs\n    else:\n        inputs = [args] + outputs[:-1]\n\n    outputs = loop_proc(queues_input, queues_output, inputs)\n    result = outputs[-1]\n\n    if result is not None:\n        prediction_id_list.append(result[\"prediction_id\"])\n        cancer_list.append(result[\"cancer\"])\n\n    idx = idx + 1\n\n```\n\nSince three processes run simultaneously, ideally, the process that takes the longest time will be hidden by other processes. By using this technique, it is expected that inference on T4 GPU can be executed more efficiently. If you have any other good ideas, please let me know.",
    "2162599": "Nice grounding work and great notebook!"
  }
}