{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px; \n            color:#150d0a;\n            margin:10px;\n            font-size:220%;\n            text-align:center;\n            display:fill;\n            border-radius:20px;\n            border-width: 5px;\n            border-style: solid;\n            border-color: #150d0a;\n            background-color:#eca912;\n            overflow:hidden;\n            font-weight:500\">A Typical Torch Pipeline for training Neural Networks</div>\n<p><center style=\"color:purple; font-family:arial;\">From Nobody to Somebody</center></p>\n\n***","metadata":{}},{"cell_type":"markdown","source":"![looking.png](https://geekflare.com/wp-content/uploads/2022/11/pytorch-installation.png)\n\n<cite>Image from https://geekflare.com/pytorch-installation-windows-and-linux/</cite>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌 <b>The Purpose of This Notebook:</b><br><br>\n    1. Deliver a simple, easy-to-follow, yet on-point tutorial on how PyTorch works for tarining neural networks.<br><br>\n    2. Help people who are quite new to ML projects, by giving them a concise but insightful explanation regarding Torch basics and fundamentals.<br><br>\n    3. Make it easier for beginners to efficiently build their own neural networks for competitions.\n</div>","metadata":{}},{"cell_type":"markdown","source":"# 0. Import Packages","metadata":{}},{"cell_type":"code","source":"import torch\nimport numpy as np\nfrom tqdm.autonotebook import tqdm\nprint(f'\\nVersion and device of Torch used: {torch.__version__}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:03.063731Z","iopub.execute_input":"2023-08-11T02:09:03.064288Z","iopub.status.idle":"2023-08-11T02:09:07.523896Z","shell.execute_reply.started":"2023-08-11T02:09:03.064234Z","shell.execute_reply":"2023-08-11T02:09:07.522671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. It is all about Tensors!\n> ML is genuinely about representing things into tensors, processing tensors, and then output tensors","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n       display:fill;\n       border-radius:5px;\n       background-color:#5642C5;\n       font-size:110%;\n       font-family:Verdana;\n       letter-spacing:0.5px\">\n    <p style=\"padding: 10px;\n          color:white;\">\n        It may sound funny that the most important thing about training neural networks using PyTorch is the concept Tensor, for beginners. People who are in the early stage of leanring ML and its implementations with Torch or Tensorflow tend to underestimate the importance of Tensor, which could lead to different types of issues as they are learning more about things in ML. Therefore, it is essential to know Tensors more.\n    </p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<blockquote><p style=\"font-size:16px; color:#159364; font-family:verdana;\">💬 As per ChatGPT: Tensors are a fundamental concept in PyTorch, serving as the core data structure that enables efficient and flexible computation in deep learning. Similar to multidimensional arrays, tensors store and manipulate numerical data, such as images, audio, and text, in a way that aligns with the underlying hardware, such as GPUs. Tensors facilitate mathematical operations, like matrix multiplication and differentiation, crucial for training and optimizing neural networks. PyTorch's tensor operations are highly optimized, allowing researchers and developers to build and experiment with complex machine learning models more efficiently, making it an essential component of modern deep learning frameworks.</p></blockquote>","metadata":{}},{"cell_type":"code","source":"# we can use torch.randn() to generate multi-dimensional tensors. torch.rand(), torch.empty() will do as well.\nx_1 = torch.randn(2) \nx_3 = torch.randn(2, 3, 4)\nprint(f'Create tensors with 1 dimension and 3 dimensions:\\n\\n {x_1}\\n and\\n {x_3}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.526219Z","iopub.execute_input":"2023-08-11T02:09:07.527633Z","iopub.status.idle":"2023-08-11T02:09:07.651823Z","shell.execute_reply.started":"2023-08-11T02:09:07.527583Z","shell.execute_reply":"2023-08-11T02:09:07.650162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the Size property\nprint(f'Size of x_1 and x_3 are {x_1.size()} and {x_3.size()}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.653903Z","iopub.execute_input":"2023-08-11T02:09:07.654691Z","iopub.status.idle":"2023-08-11T02:09:07.660498Z","shell.execute_reply.started":"2023-08-11T02:09:07.654637Z","shell.execute_reply":"2023-08-11T02:09:07.659466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the Shape property\nprint(f'Shape of x_1 and x_3 are {x_1.shape} and {x_3.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.663674Z","iopub.execute_input":"2023-08-11T02:09:07.664087Z","iopub.status.idle":"2023-08-11T02:09:07.673353Z","shell.execute_reply.started":"2023-08-11T02:09:07.664049Z","shell.execute_reply":"2023-08-11T02:09:07.671841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check data type of, e.g., the Tensor x_3\nprint(f'Data type of x_3: {x_3.dtype}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.675177Z","iopub.execute_input":"2023-08-11T02:09:07.675862Z","iopub.status.idle":"2023-08-11T02:09:07.685001Z","shell.execute_reply.started":"2023-08-11T02:09:07.675821Z","shell.execute_reply":"2023-08-11T02:09:07.683907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can specify dtype when we create Tensor\nx_float16 = torch.empty(2, 3, 4, dtype=torch.float64)\nprint(f'x_float16 = {x_float16}\\n\\nData type of x_float16: {x_float16.dtype}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.686672Z","iopub.execute_input":"2023-08-11T02:09:07.687598Z","iopub.status.idle":"2023-08-11T02:09:07.700456Z","shell.execute_reply.started":"2023-08-11T02:09:07.687555Z","shell.execute_reply":"2023-08-11T02:09:07.699266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can also create Tensors using torch.tensor() method\nx_torch_tensor = torch.tensor([1, 2, 3.0])\nprint(x_torch_tensor, x_torch_tensor.shape, x_torch_tensor.dtype)","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.702152Z","iopub.execute_input":"2023-08-11T02:09:07.702847Z","iopub.status.idle":"2023-08-11T02:09:07.711453Z","shell.execute_reply.started":"2023-08-11T02:09:07.702805Z","shell.execute_reply":"2023-08-11T02:09:07.710331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Operations with tensors\nx, y = torch.rand(1, 3), torch.ones(1, 3)\n\nx_plus_y = x + y # torch.add(x, y)\nx_times_y = x * y # torch.mul(x, y)\nx_substract_y = x - y # torch.sub(x, y)\nx_divide_y = x / y # torch.div(x, y)\n\nprint(f'x = {x}\\ny = {y}')\nprint(f'x+y = {x_plus_y}\\nx*y = {x_times_y}\\nx-y = {x_substract_y}\\nx/y = {x_divide_y}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.713708Z","iopub.execute_input":"2023-08-11T02:09:07.714622Z","iopub.status.idle":"2023-08-11T02:09:07.73909Z","shell.execute_reply.started":"2023-08-11T02:09:07.714571Z","shell.execute_reply":"2023-08-11T02:09:07.738077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**One of the most important things that need to be beared in mind is the shpae of tensors, as one of the most common errors we will encounter during training. we can use torch.view() to manipulate shape.**","metadata":{}},{"cell_type":"code","source":"x = torch.empty(4, 5)\nx_flattened = x.view(-1) # -1 is one of the most things we will find during the construction of neural networks, it means that the size will be infereed automatically by torch.\nx_reshaped_left = x.view(-1, 2) # the position marked by -1 will be automatically inferred, by formula 4*5/2.\nx_reshaped_right = x.view(10, -1) # the position marked by -1 will be automatically inferred, by formula 4*5/10.\n\nprint(f'Shape of x, x_flattened, x_reshaped_left, x_reshaped_right:\\n\\n{x.shape}\\n{x_flattened.shape}\\n{x_reshaped_left.shape}\\n{x_reshaped_right.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.740877Z","iopub.execute_input":"2023-08-11T02:09:07.74122Z","iopub.status.idle":"2023-08-11T02:09:07.748484Z","shell.execute_reply.started":"2023-08-11T02:09:07.74119Z","shell.execute_reply":"2023-08-11T02:09:07.747651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. The device where Torch deals with training should be specified!","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n       display:fill;\n       border-radius:5px;\n       background-color:#5642C5;\n       font-size:110%;\n       font-family:Verdana;\n       letter-spacing:0.5px\">\n    <p style=\"padding: 10px;\n          color:white;\">\n            Always use the following chunk of agnostic code to specify Device\n    </p>\n</div>","metadata":{}},{"cell_type":"code","source":"device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') # when possible, GPU is preferred over CPU for parallel computing.","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.751586Z","iopub.execute_input":"2023-08-11T02:09:07.752181Z","iopub.status.idle":"2023-08-11T02:09:07.759555Z","shell.execute_reply.started":"2023-08-11T02:09:07.752142Z","shell.execute_reply":"2023-08-11T02:09:07.758687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(x.device)\nx.to(device) # we use .to() method to deliver variable to the device that we specified. \nprint(x.device)","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.761161Z","iopub.execute_input":"2023-08-11T02:09:07.761612Z","iopub.status.idle":"2023-08-11T02:09:07.777213Z","shell.execute_reply.started":"2023-08-11T02:09:07.761569Z","shell.execute_reply":"2023-08-11T02:09:07.776056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. An example of Linear Regression\n> Linear regression is a basic machine learning method for predicting a continuous output by finding the best linear relationship between input features and the target.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n       display:fill;\n       border-radius:5px;\n       background-color:#5642C5;\n       font-size:110%;\n       font-family:Verdana;\n       letter-spacing:0.5px\">\n    <p style=\"padding: 10px;\n          color:white;\">\n            We will try to find the Weight of a function, via training Torch-based nueral networks with data we coined\n    </p>\n</div>","metadata":{}},{"cell_type":"code","source":"# Create the dataset for the Linear Regression, i.e., those on the x-axis and y-axis\nX = torch.tensor(np.arange(1, 11), dtype=torch.float32)\nY = X * 6\n\nprint(X, Y)\nprint(X.requires_grad, Y.grad_fn)","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.778899Z","iopub.execute_input":"2023-08-11T02:09:07.779312Z","iopub.status.idle":"2023-08-11T02:09:07.791913Z","shell.execute_reply.started":"2023-08-11T02:09:07.779271Z","shell.execute_reply":"2023-08-11T02:09:07.790805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# If we tell Torch to track the gradient, we can check the difference as below.\nX.requires_grad_(True)\nY = X * 6\n\nprint(X, Y)\nprint(X.requires_grad, Y.grad_fn)\n\n# we can use .detach() method to tell Torch not to track gradient\nX.detach_() # the underscore mark in the tail means that this method will work in_place, do Google what does in_place mean if you are unclear about.\nY = X * 6\nprint(X.requires_grad, Y.grad_fn)","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.793428Z","iopub.execute_input":"2023-08-11T02:09:07.793771Z","iopub.status.idle":"2023-08-11T02:09:07.80741Z","shell.execute_reply.started":"2023-08-11T02:09:07.793731Z","shell.execute_reply":"2023-08-11T02:09:07.806147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"weight = torch.tensor(0, dtype=torch.float32, requires_grad=True) # we aim to use Torch to learn this variable, therefore Torch has to track its gradient.","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.808754Z","iopub.execute_input":"2023-08-11T02:09:07.809661Z","iopub.status.idle":"2023-08-11T02:09:07.814915Z","shell.execute_reply.started":"2023-08-11T02:09:07.809626Z","shell.execute_reply":"2023-08-11T02:09:07.813886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the forward function\ndef compute_forward(input_data, weight):\n    return weight * input_data\n\n# Define the loss function\ndef compute_loss(predicted_values, target_values):\n    return torch.mean((predicted_values - target_values)**2)\n\n# Hyperparameters\nlearning_rate = 0.001\nnum_epochs = 100\n\n# Initial weight value\nweight = torch.tensor(1.0, requires_grad=True)\n\n# Training loop\nfor epoch in tqdm(range(num_epochs)):\n    # Forward pass\n    predictions = compute_forward(input_data=X, weight=weight)\n    \n    # Compute loss\n    current_loss = compute_loss(predicted_values=predictions, target_values=Y)\n    \n    # Backpropagation\n    current_loss.backward()\n    \n    # Update weight using gradient descent\n    with torch.no_grad():\n        weight -= learning_rate * weight.grad\n        \n    # Reset the gradient\n    weight.grad.zero_()\n    \n    # Print progress every 10 epochs\n    if (epoch + 1) % 10 == 0:\n        print(f'Epoch {epoch + 1}: Weight = {weight.item()}, Loss = {current_loss.item()}')\n","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:09:07.816547Z","iopub.execute_input":"2023-08-11T02:09:07.817752Z","iopub.status.idle":"2023-08-11T02:09:07.879827Z","shell.execute_reply.started":"2023-08-11T02:09:07.817711Z","shell.execute_reply":"2023-08-11T02:09:07.879045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"padding:20px; \n            color:#36FF00;\n            margin:10px;\n            font-size:110%;\n            display:fill;\n            border-radius:10px;\n            border-style: solid;\n            border-color: #36FF00;\n            background-color:#000000;\n            overflow:hidden;\n            font-weight:500\"><b>The typical training loop:</b> Forward the input to generate output --> Calculate the Loss function --> Backpropergate the gradient --> Update weights --> Zero the grident after weight updating. Then, come to the next loop, i.e., epoch herein.</div>","metadata":{}},{"cell_type":"markdown","source":"# 4. Training the Example with Model, Loss and Optimizer provided by Torch","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\n\n# Define the training data\nX = torch.tensor([[1], [2], [3], [4], [5], [6], [7], [8]], dtype=torch.float32)\nY = torch.tensor([[6], [8], [18], [24], [30], [36], [42], [48]], dtype=torch.float32)\n\n# Define a single test sample\nX_test = torch.tensor([5], dtype=torch.float32)\n\n# Determine the number of samples and features\nn_samples, n_features = X.shape\nprint(f'n_samples = {n_samples}, n_features = {n_features}')\n\n# Define the linear regression model\nclass LinearRegression(nn.Module):\n    def __init__(self, input_dim, output_dim):\n        super(LinearRegression, self).__init__()\n        # Define a linear layer with specified input and output dimensions\n        self.lin = nn.Linear(input_dim, output_dim)\n\n    def forward(self, x):\n        # Perform a forward pass through the linear layer\n        return self.lin(x)\n\n# Create the linear regression model\ninput_size, output_size = n_features, n_features\nmodel = LinearRegression(input_size, output_size)\n\n# Print the prediction before training\nprint(f'Prediction before training: f({X_test.item()}) = {model(X_test).item():.3f}')\n\n# Define the loss function (Mean Squared Error) and the optimizer (Stochastic Gradient Descent)\nlearning_rate = 0.001\nn_epochs = 100\n\nloss_fn = nn.MSELoss()\noptimizer = torch.optim.SGD(model.parameters(), lr=learning_rate)\n\n# Training loop\nfor epoch in range(n_epochs):\n    # Forward pass: predict the output using the current model\n    y_predicted = model(X)\n\n    # Compute the loss between the predicted and actual values\n    loss_value = loss_fn(Y, y_predicted)\n\n    # Backward pass: compute gradients and perform optimization step\n    loss_value.backward()\n    optimizer.step()\n\n    # Zero the gradients to prevent accumulation in the next iteration\n    optimizer.zero_grad()\n\n    # Print progress every 10 epochs\n    if (epoch+1) % 10 == 0:\n        # Retrieve the learned weight parameter from the model\n        w, _ = model.parameters()\n        print('Epoch ', epoch+1, ': w = ', w[0][0].item(), ' loss = ', loss_value.item())\n\n# Print the prediction after training\nprint(f'Prediction after training: f({X_test.item()}) = {model(X_test).item():.3f}')\n","metadata":{"execution":{"iopub.status.busy":"2023-08-11T02:24:04.895142Z","iopub.execute_input":"2023-08-11T02:24:04.895704Z","iopub.status.idle":"2023-08-11T02:24:04.945456Z","shell.execute_reply.started":"2023-08-11T02:24:04.895666Z","shell.execute_reply":"2023-08-11T02:24:04.943781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Description of the above snippet of Torch code:\n\nThis code snippet performs linear regression using PyTorch:\n\n1. **Data**: It defines input (`X`) and target (`Y`) data, and a test sample (`X_test`).\n\n2. **Model**: A linear regression model is created as a subclass of `nn.Module`.\n\n3. **Training**: It trains the model to fit the input-output relationship in the data using MSE loss and SGD optimizer.\n\n4. **Progress**: It prints progress (epoch, weight, loss) during training.\n\n5. **Prediction**: After training, it prints the model's prediction on the test sample.\n\nThe goal is to teach the model to predict output (`Y`) from input (`X`) and then validate its prediction on the test sample.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:25px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp; If you found this simple notebook is benificial for you to understand how Torch works, please upvote. Cheers, mate!\n</div>","metadata":{}},{"cell_type":"markdown","source":"# Acknowledgement: This notebook is inspired by [AssemblyAI](https://www.assemblyai.com), do check their terrific tutorials of ML\n","metadata":{}}]}