{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## **Problem Statement**\n\n**Objective**\n\nYou are given a data of US Airline tweets and their sentiment. The task is to do sentiment analysis about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as \"late flight\" or \"rude service\").","metadata":{"id":"KppRrzQ7s-Gb"}},{"cell_type":"code","source":"#importing req. Lib.\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport re\nimport nltk\nfrom nltk.corpus import stopwords\nfrom sklearn.model_selection import train_test_split\nfrom mlxtend.plotting import plot_confusion_matrix\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score,confusion_matrix,classification_report","metadata":{"id":"hBSeD3pqrOYL","execution":{"iopub.status.busy":"2022-08-27T08:25:12.58007Z","iopub.execute_input":"2022-08-27T08:25:12.580628Z","iopub.status.idle":"2022-08-27T08:25:14.185621Z","shell.execute_reply.started":"2022-08-27T08:25:12.58058Z","shell.execute_reply":"2022-08-27T08:25:14.184266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#load our data set\ndata = pd.read_csv('../input/twitter-airline-sentiment/Tweets.csv')","metadata":{"id":"wbFyZ_0ktf2_","execution":{"iopub.status.busy":"2022-08-27T08:25:15.755715Z","iopub.execute_input":"2022-08-27T08:25:15.756651Z","iopub.status.idle":"2022-08-27T08:25:15.900503Z","shell.execute_reply.started":"2022-08-27T08:25:15.756614Z","shell.execute_reply":"2022-08-27T08:25:15.899483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.shape","metadata":{"id":"41FYBCaS_YzV","outputId":"27cccad5-a6dc-4b27-ee78-f41a86b89ba4","execution":{"iopub.status.busy":"2022-08-27T08:25:17.39658Z","iopub.execute_input":"2022-08-27T08:25:17.397685Z","iopub.status.idle":"2022-08-27T08:25:17.406043Z","shell.execute_reply.started":"2022-08-27T08:25:17.39764Z","shell.execute_reply":"2022-08-27T08:25:17.405162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#looking into our data\ndata.head()","metadata":{"id":"bzcg27aitx5y","outputId":"8416e314-0b3e-489c-ee82-ee2ff827a1eb","execution":{"iopub.status.busy":"2022-08-27T08:25:18.10978Z","iopub.execute_input":"2022-08-27T08:25:18.11075Z","iopub.status.idle":"2022-08-27T08:25:18.139196Z","shell.execute_reply.started":"2022-08-27T08:25:18.110708Z","shell.execute_reply":"2022-08-27T08:25:18.138386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking last 5 entries\ndata.tail()","metadata":{"id":"JJqtEe-0tzy3","outputId":"dd1161c9-a3c8-442b-e953-2a9315b3c804","execution":{"iopub.status.busy":"2022-08-27T08:25:18.537564Z","iopub.execute_input":"2022-08-27T08:25:18.537972Z","iopub.status.idle":"2022-08-27T08:25:18.558208Z","shell.execute_reply.started":"2022-08-27T08:25:18.537938Z","shell.execute_reply":"2022-08-27T08:25:18.557095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking columns in our data\ndata.columns","metadata":{"id":"t7YWhWR_uDlk","outputId":"1ea9587a-e79a-4926-921e-9010789e2529","execution":{"iopub.status.busy":"2022-08-27T08:25:19.07178Z","iopub.execute_input":"2022-08-27T08:25:19.072944Z","iopub.status.idle":"2022-08-27T08:25:19.08014Z","shell.execute_reply.started":"2022-08-27T08:25:19.0729Z","shell.execute_reply":"2022-08-27T08:25:19.079047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking info our data\ndata.info()","metadata":{"id":"K3IhqOUhup7A","outputId":"10f4658b-0934-4706-b021-9355b3194464","execution":{"iopub.status.busy":"2022-08-27T08:25:19.485574Z","iopub.execute_input":"2022-08-27T08:25:19.486721Z","iopub.status.idle":"2022-08-27T08:25:19.523255Z","shell.execute_reply.started":"2022-08-27T08:25:19.486675Z","shell.execute_reply":"2022-08-27T08:25:19.521932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking unique values \ndata.nunique()","metadata":{"id":"pgcq-jhquKPT","outputId":"24100084-3f2b-4afd-b3a5-309322bddfc4","execution":{"iopub.status.busy":"2022-08-27T08:25:19.781559Z","iopub.execute_input":"2022-08-27T08:25:19.781961Z","iopub.status.idle":"2022-08-27T08:25:19.817358Z","shell.execute_reply.started":"2022-08-27T08:25:19.781929Z","shell.execute_reply":"2022-08-27T08:25:19.816158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#checking null values in our data\ndata.isnull().sum()","metadata":{"id":"VuKJ_u2cufiX","outputId":"7c0c3333-d81c-4ca9-982a-1d8ec31d4c4d","execution":{"iopub.status.busy":"2022-08-27T08:25:20.122693Z","iopub.execute_input":"2022-08-27T08:25:20.123106Z","iopub.status.idle":"2022-08-27T08:25:20.140708Z","shell.execute_reply.started":"2022-08-27T08:25:20.123067Z","shell.execute_reply":"2022-08-27T08:25:20.139665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Preprocessing on data**","metadata":{"id":"iTTDQvM3v7R8"}},{"cell_type":"markdown","source":"tweet_created column got the date recorts and showing type is object we have to change it of date time format","metadata":{"id":"wkQmnNagwGQR"}},{"cell_type":"code","source":"data['tweet_created'] = pd.to_datetime(data['tweet_created']).dt.date","metadata":{"id":"-BcWc6xUvFq5","execution":{"iopub.status.busy":"2022-08-27T08:25:21.033419Z","iopub.execute_input":"2022-08-27T08:25:21.033852Z","iopub.status.idle":"2022-08-27T08:25:21.135833Z","shell.execute_reply.started":"2022-08-27T08:25:21.033807Z","shell.execute_reply":"2022-08-27T08:25:21.134779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['tweet_created'] = pd.to_datetime(data['tweet_created'])","metadata":{"id":"_bXJqLGGy4mT","execution":{"iopub.status.busy":"2022-08-27T08:25:21.358032Z","iopub.execute_input":"2022-08-27T08:25:21.358413Z","iopub.status.idle":"2022-08-27T08:25:21.369038Z","shell.execute_reply.started":"2022-08-27T08:25:21.35838Z","shell.execute_reply":"2022-08-27T08:25:21.367948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.info()","metadata":{"id":"C4WmcHVixDhk","outputId":"130983ea-0c7d-411b-b4c3-1eb93bbd4170","execution":{"iopub.status.busy":"2022-08-27T08:25:21.646718Z","iopub.execute_input":"2022-08-27T08:25:21.64735Z","iopub.status.idle":"2022-08-27T08:25:21.667942Z","shell.execute_reply.started":"2022-08-27T08:25:21.647314Z","shell.execute_reply":"2022-08-27T08:25:21.666679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()","metadata":{"id":"Iu3o-YqqxFdc","outputId":"ad9987d5-1996-43be-fa09-7e53dd00d809","execution":{"iopub.status.busy":"2022-08-27T08:25:21.928946Z","iopub.execute_input":"2022-08-27T08:25:21.929578Z","iopub.status.idle":"2022-08-27T08:25:21.950357Z","shell.execute_reply.started":"2022-08-27T08:25:21.929541Z","shell.execute_reply":"2022-08-27T08:25:21.949047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['tweet_created'].min()","metadata":{"id":"yf9UXXcJxL4Z","outputId":"baf409f7-c372-483e-d6e6-c6e795f5f619","execution":{"iopub.status.busy":"2022-08-27T08:25:22.255518Z","iopub.execute_input":"2022-08-27T08:25:22.256226Z","iopub.status.idle":"2022-08-27T08:25:22.264517Z","shell.execute_reply.started":"2022-08-27T08:25:22.256189Z","shell.execute_reply":"2022-08-27T08:25:22.263318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['tweet_created'].max()","metadata":{"id":"0d17zq3XxbVr","outputId":"45f472ae-ecd5-4307-8782-f803b3b4c561","execution":{"iopub.status.busy":"2022-08-27T08:25:22.498192Z","iopub.execute_input":"2022-08-27T08:25:22.499359Z","iopub.status.idle":"2022-08-27T08:25:22.506098Z","shell.execute_reply.started":"2022-08-27T08:25:22.499294Z","shell.execute_reply":"2022-08-27T08:25:22.505052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we have data from 16th feb 2015 to 25 feb 2015 mins we have data of 9 days.","metadata":{"id":"LTR0nBypxsQz"}},{"cell_type":"code","source":"#checking uniques values in tweet_created columns\ndata['tweet_created'].nunique()","metadata":{"id":"rKCkw51RxgjW","outputId":"013f3a45-8ed2-4e97-d28d-b8a18bc6cac2","execution":{"iopub.status.busy":"2022-08-27T08:25:23.015446Z","iopub.execute_input":"2022-08-27T08:25:23.016397Z","iopub.status.idle":"2022-08-27T08:25:23.024483Z","shell.execute_reply.started":"2022-08-27T08:25:23.016353Z","shell.execute_reply":"2022-08-27T08:25:23.023525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numberoftweets = data.groupby('tweet_created').size()","metadata":{"id":"JiLPH_IoyPFB","execution":{"iopub.status.busy":"2022-08-27T08:25:23.253173Z","iopub.execute_input":"2022-08-27T08:25:23.254224Z","iopub.status.idle":"2022-08-27T08:25:23.261238Z","shell.execute_reply.started":"2022-08-27T08:25:23.254173Z","shell.execute_reply":"2022-08-27T08:25:23.260156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numberoftweets.dtype","metadata":{"id":"O4Cn0Mhg0Ei3","outputId":"296938b8-d436-49b3-aeb6-4f14de1f7e36","execution":{"iopub.status.busy":"2022-08-27T08:25:23.510211Z","iopub.execute_input":"2022-08-27T08:25:23.510862Z","iopub.status.idle":"2022-08-27T08:25:23.517157Z","shell.execute_reply.started":"2022-08-27T08:25:23.510822Z","shell.execute_reply":"2022-08-27T08:25:23.515963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numberoftweets","metadata":{"id":"K7a6eyk_2E-f","outputId":"ab2c3b61-b667-4146-c085-57d16244688c","execution":{"iopub.status.busy":"2022-08-27T08:25:23.762272Z","iopub.execute_input":"2022-08-27T08:25:23.763182Z","iopub.status.idle":"2022-08-27T08:25:23.77115Z","shell.execute_reply.started":"2022-08-27T08:25:23.76314Z","shell.execute_reply":"2022-08-27T08:25:23.770175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"here we can see tweets created every day","metadata":{"id":"jv86CeVU2pPd"}},{"cell_type":"markdown","source":"# **treating with null values**","metadata":{"id":"-M-u6zr59tgU"}},{"cell_type":"code","source":"data.isna().sum()","metadata":{"id":"joUDquqs2Rsn","outputId":"84e54f21-197b-4cd4-933f-1a71708f55bf","execution":{"iopub.status.busy":"2022-08-27T08:25:24.529422Z","iopub.execute_input":"2022-08-27T08:25:24.529855Z","iopub.status.idle":"2022-08-27T08:25:24.549248Z","shell.execute_reply.started":"2022-08-27T08:25:24.529817Z","shell.execute_reply":"2022-08-27T08:25:24.547764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"in ","metadata":{"id":"xQDgU1Jp-cmj"}},{"cell_type":"code","source":"print(\"Percentage null or na values in df\")\n((data.isnull() | data.isna()).sum() * 100 / data.index.size).round(2)","metadata":{"id":"7a58BVJI2Zll","outputId":"995695e0-6e87-4f69-aa69-958b8b483b22","execution":{"iopub.status.busy":"2022-08-25T20:22:43.623055Z","iopub.execute_input":"2022-08-25T20:22:43.623497Z","iopub.status.idle":"2022-08-25T20:22:43.648997Z","shell.execute_reply.started":"2022-08-25T20:22:43.623458Z","shell.execute_reply":"2022-08-25T20:22:43.648153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**airline_sentiment_gold, negativereason_gold** have more than 99% missing data And **tweet_coord** have nearly 93% missing data. It will be better to delete these columns as they will not provide any constructive information","metadata":{"id":"stVZW6wr_6qi"}},{"cell_type":"code","source":"del data['tweet_coord']\ndel data['airline_sentiment_gold']\ndel data['negativereason_gold']\ndata.head()","metadata":{"id":"AcYVR9S0_y8x","outputId":"4baf337a-cee0-4342-fde9-c7db844542be","execution":{"iopub.status.busy":"2022-08-25T20:22:44.634142Z","iopub.execute_input":"2022-08-25T20:22:44.63487Z","iopub.status.idle":"2022-08-25T20:22:44.65607Z","shell.execute_reply.started":"2022-08-25T20:22:44.634821Z","shell.execute_reply":"2022-08-25T20:22:44.654823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"freq = data.groupby('negativereason').size()","metadata":{"id":"Wj8C_4gpE4Kz","execution":{"iopub.status.busy":"2022-08-25T20:22:45.099787Z","iopub.execute_input":"2022-08-25T20:22:45.100238Z","iopub.status.idle":"2022-08-25T20:22:45.107427Z","shell.execute_reply.started":"2022-08-25T20:22:45.100203Z","shell.execute_reply":"2022-08-25T20:22:45.106526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"freq","metadata":{"id":"4-7xlYZUJC2G","outputId":"784f62cf-ad2d-4b93-df29-9986edaa30fa","execution":{"iopub.status.busy":"2022-08-25T20:22:45.525016Z","iopub.execute_input":"2022-08-25T20:22:45.525454Z","iopub.status.idle":"2022-08-25T20:22:45.533718Z","shell.execute_reply.started":"2022-08-25T20:22:45.525416Z","shell.execute_reply":"2022-08-25T20:22:45.532484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we cant fill it will affect in bad way for example we have positive reviwe and we fill the values with mode that means with Customer Service Issue  it is missmatch and can be affect on train model so we keep the data as it is.","metadata":{"id":"bpzwuePiJRqh"}},{"cell_type":"markdown","source":"# **EDA**","metadata":{"id":"5n5ziM4GJlhX"}},{"cell_type":"markdown","source":"### **Count of Type of Sentiment**","metadata":{"id":"4TZAoB0FKiAn"}},{"cell_type":"code","source":"counter = data.airline_sentiment.value_counts()\nindex = [1,2,3]\nplt.figure(1,figsize=(12,6))\nplt.bar(index,counter,color=['green','red','blue'])\nplt.xticks(index,['negative','neutral','positive'],rotation=0)\nplt.xlabel('Sentiment Type')\nplt.ylabel('Sentiment Count')\nplt.title('Count of Type of Sentiment')","metadata":{"id":"x7A2OUXsKiif","outputId":"4194c7bb-33f0-4d49-d980-4e490bf32782","execution":{"iopub.status.busy":"2022-08-25T20:22:47.307406Z","iopub.execute_input":"2022-08-25T20:22:47.308228Z","iopub.status.idle":"2022-08-25T20:22:47.531152Z","shell.execute_reply.started":"2022-08-25T20:22:47.308173Z","shell.execute_reply":"2022-08-25T20:22:47.52999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Airline sentiments for each airline**","metadata":{"id":"zWDn_qAdJtPR"}},{"cell_type":"code","source":"#checking differtent airlines we have\ndata['airline'].unique()","metadata":{"id":"V6OuAP9RJ4Ev","outputId":"05a14703-b58f-4add-def4-ff02b7bf126f","execution":{"iopub.status.busy":"2022-08-25T20:22:48.113073Z","iopub.execute_input":"2022-08-25T20:22:48.113487Z","iopub.status.idle":"2022-08-25T20:22:48.122959Z","shell.execute_reply.started":"2022-08-25T20:22:48.113453Z","shell.execute_reply":"2022-08-25T20:22:48.121636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Total number of tweets for each airline \\n \",data.groupby('airline')['airline_sentiment'].count().sort_values(ascending=False))\nairlines= ['US Airways','United','American','Southwest','Delta','Virgin America']\nplt.figure(1,figsize=(12, 12))\nfor i in airlines:\n    indices= airlines.index(i)\n    plt.subplot(2,3,indices+1)\n    new_df=data[data['airline']==i]\n    count=new_df['airline_sentiment'].value_counts()\n    Index = [1,2,3]\n    plt.bar(Index,count, color=['red', 'green', 'blue'])\n    plt.xticks(Index,['negative','neutral','positive'])\n    plt.ylabel('Mood Count')\n    plt.xlabel('Mood')\n    plt.title('Count of Moods of '+i)","metadata":{"id":"lP898gTFJJUQ","outputId":"8243e1e8-b5a2-4737-a91d-db9715758e9c","execution":{"iopub.status.busy":"2022-08-25T20:22:48.522767Z","iopub.execute_input":"2022-08-25T20:22:48.5239Z","iopub.status.idle":"2022-08-25T20:22:49.285012Z","shell.execute_reply.started":"2022-08-25T20:22:48.523858Z","shell.execute_reply":"2022-08-25T20:22:49.283783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks like people are not having pleasant flights these days. It is important to know which airline pleases their costumers the most and vice versa, so we sill be looking at the percentage of the negative reviews for each airline.","metadata":{"id":"6xGPy0L3K0v-"}},{"cell_type":"code","source":"neg_tweets = data.groupby(['airline','airline_sentiment']).count().iloc[:,0]\ntotal_tweets = data.groupby(['airline'])['airline_sentiment'].count()\n\nmy_dict = {'American':neg_tweets[0] / total_tweets[0],'Delta':neg_tweets[3] / total_tweets[1],'Southwest': neg_tweets[6] / total_tweets[2],\n'US Airways': neg_tweets[9] / total_tweets[3],'United': neg_tweets[12] / total_tweets[4],'Virgin': neg_tweets[15] / total_tweets[5]}\nperc = pd.DataFrame.from_dict(my_dict, orient = 'index')\nperc.columns = ['Percent Negative']\nprint(perc)\nax = perc.plot(kind = 'bar', rot=0, colormap = 'Greens_r', figsize = (15,6))\nax.set_xlabel('Airlines')\nax.set_ylabel('Percentage of negative tweets')\nplt.show()","metadata":{"id":"W8Qr1Fu1KTNE","outputId":"b096fc6c-cd64-40b9-a847-33de63771811","execution":{"iopub.status.busy":"2022-08-25T20:22:49.543461Z","iopub.execute_input":"2022-08-25T20:22:49.5439Z","iopub.status.idle":"2022-08-25T20:22:49.795003Z","shell.execute_reply.started":"2022-08-25T20:22:49.543864Z","shell.execute_reply":"2022-08-25T20:22:49.793756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* United, US Airways, American substantially get negative reactions.\n* Tweets for Virgin America are the most balanced.","metadata":{"id":"6fe6Jz1CNA0n"}},{"cell_type":"code","source":"figure_2 = data.groupby(['airline', 'airline_sentiment']).size()\nfigure_2.unstack().plot(kind='bar', stacked=True, figsize=(15,10))","metadata":{"id":"qH5WQkh_K6UO","outputId":"81b2439a-518a-4def-897d-3141b37afdfd","execution":{"iopub.status.busy":"2022-08-25T20:22:54.871006Z","iopub.execute_input":"2022-08-25T20:22:54.871494Z","iopub.status.idle":"2022-08-25T20:22:55.196837Z","shell.execute_reply.started":"2022-08-25T20:22:54.871445Z","shell.execute_reply":"2022-08-25T20:22:55.195268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(figure_2)","metadata":{"id":"VUa7VE2qNY50","outputId":"80aa5cef-0265-436b-ed45-a2f078648152","execution":{"iopub.status.busy":"2022-08-25T20:22:55.493154Z","iopub.execute_input":"2022-08-25T20:22:55.493551Z","iopub.status.idle":"2022-08-25T20:22:55.501181Z","shell.execute_reply.started":"2022-08-25T20:22:55.49352Z","shell.execute_reply":"2022-08-25T20:22:55.499751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Last but not least, people complain for many reasons about their flights; 10 reasons to be specific**","metadata":{"id":"uD07weQkNmOl"}},{"cell_type":"code","source":"negative_reasons = data.groupby('airline')['negativereason'].value_counts(ascending=True)\nnegative_reasons.groupby(['airline','negativereason']).sum().unstack().plot(kind='bar',figsize=(22,12))\nplt.xlabel('Airline Company')\nplt.ylabel('Number of Negative reasons')\nplt.title(\"The number of the count of negative reasons for airlines\")\nplt.show()","metadata":{"id":"XRrHtPB5Nc01","outputId":"e3aaf8a7-1257-4438-9dff-aa256ed50f0f","execution":{"iopub.status.busy":"2022-08-25T20:22:56.470138Z","iopub.execute_input":"2022-08-25T20:22:56.470868Z","iopub.status.idle":"2022-08-25T20:22:57.092918Z","shell.execute_reply.started":"2022-08-25T20:22:56.470826Z","shell.execute_reply":"2022-08-25T20:22:57.091706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### What are the reasons for negative sentimental tweets for each airline ?","metadata":{"id":"EgoOpCRwOQjD"}},{"cell_type":"markdown","source":"We will explore the negative reason column of our dataframe to extract conclusions about negative sentiments in the tweets by the customers","metadata":{"id":"gItQ7h-COYpY"}},{"cell_type":"code","source":"#get the number of negative reasons\ndata['negativereason'].nunique()\n\nNR_Count=dict(data['negativereason'].value_counts(sort=False))\ndef NR_Count(Airline):\n    if Airline=='All':\n        a=data\n    else:\n        a=data[data['airline']==Airline]\n    count=dict(a['negativereason'].value_counts())\n    Unique_reason=list(data['negativereason'].unique())\n    Unique_reason=[x for x in Unique_reason if str(x) != 'nan']\n    Reason_frame=pd.DataFrame({'Reasons':Unique_reason})\n    Reason_frame['count']=Reason_frame['Reasons'].apply(lambda x: count[x])\n    return Reason_frame\ndef plot_reason(Airline):\n    \n    a=NR_Count(Airline)\n    count=a['count']\n    Index = range(1,(len(a)+1))\n    plt.bar(Index,count, color=['red','yellow','blue','green','black','brown','gray','cyan','purple','orange'])\n    plt.xticks(Index,a['Reasons'],rotation=90)\n    plt.ylabel('Count')\n    plt.xlabel('Reason')\n    plt.title('Count of Reasons for '+Airline)\n    \nplot_reason('All')\nplt.figure(2,figsize=(13, 13))\nfor i in airlines:\n    indices= airlines.index(i)\n    plt.subplot(2,3,indices+1)\n    plt.subplots_adjust(hspace=0.9)\n    plot_reason(i)","metadata":{"id":"8cnfjR3SNrnt","outputId":"992a5b5d-4049-4a7b-ccde-4c685894e288","execution":{"iopub.status.busy":"2022-08-25T20:22:58.185468Z","iopub.execute_input":"2022-08-25T20:22:58.185943Z","iopub.status.idle":"2022-08-25T20:22:59.564449Z","shell.execute_reply.started":"2022-08-25T20:22:58.185902Z","shell.execute_reply":"2022-08-25T20:22:59.563264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Customer Service Issue is the main neagtive reason for US Airways,United,American,Southwest,Virgin America\n* Late Flight is the main negative reason for Delta\n* Interestingly, Virgin America has the least count of negative reasons (all less than 60)\n* Contrastingly to Virgin America, airlines like US Airways,United,American have more than 500 negative reasons (Late flight, Customer Service Issue)","metadata":{"id":"X6KgRaTEPZLB"}},{"cell_type":"markdown","source":"## **Is there a relationship between negative sentiments and date?**","metadata":{"id":"wDDipNSKPo02"}},{"cell_type":"markdown","source":"It will be interesting to see if the date has any effect on the sentiments of the tweets(especially negative !). We can draw various coclusions by visualizing this.","metadata":{"id":"ELrn4eLJPyEi"}},{"cell_type":"code","source":"date = data.reset_index()\n#convert the Date column to pandas datetime\ndate.tweet_created = pd.to_datetime(date.tweet_created)\n#Reduce the dates in the date column to only the date and no time stamp using the 'dt.date' method\ndate.tweet_created = date.tweet_created.dt.date\ndate.tweet_created.head()\ndf = date\nday_df = df.groupby(['tweet_created','airline','airline_sentiment']).size()\n# day_df = day_df.reset_index()\nday_df","metadata":{"id":"9KhrrUrsOh5l","outputId":"e4b97859-e0ac-4d75-da0b-71af9af4b5d5","execution":{"iopub.status.busy":"2022-08-25T20:23:00.32595Z","iopub.execute_input":"2022-08-25T20:23:00.3266Z","iopub.status.idle":"2022-08-25T20:23:00.366294Z","shell.execute_reply.started":"2022-08-25T20:23:00.326553Z","shell.execute_reply":"2022-08-25T20:23:00.365294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"plot this and get better visualization for negative tweets.","metadata":{"id":"tUjzfy0HQJuX"}},{"cell_type":"code","source":"day_df = day_df.loc(axis=0)[:,:,'negative']\n\n#groupby and plot data\nax2 = day_df.groupby(['tweet_created','airline']).sum().unstack().plot(kind = 'bar', color=['red', 'green', 'blue','yellow','purple','orange'], figsize = (15,6), rot = 70)\nlabels = ['American','Delta','Southwest','US Airways','United','Virgin America']\nax2.legend(labels = labels)\nax2.set_xlabel('Date')\nax2.set_ylabel('Negative Tweets')\nplt.show()","metadata":{"id":"otUN8jUSP8LT","outputId":"4e65d6b3-2f19-4fd2-e515-3f9223bf522d","execution":{"iopub.status.busy":"2022-08-25T20:23:01.270498Z","iopub.execute_input":"2022-08-25T20:23:01.271212Z","iopub.status.idle":"2022-08-25T20:23:01.660599Z","shell.execute_reply.started":"2022-08-25T20:23:01.271169Z","shell.execute_reply":"2022-08-25T20:23:01.65936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Interestingly, American has a sudden upsurge in negative sentimental tweets on 2015-02-23, which reduced to half the very next day 2015-02-24. (I hope American is doing better these days and resolved their Customer Service Issue as we saw before)\n* Virgin America has the least number of negative tweets throughout the weekly data that we have. It should be noted that the total number of tweets for Virgin America was also significantly less as compared to the rest airlines, and hence the least negative tweets.\n* The negative tweets for all the rest airlines is slightly skewed towards the end of the week !","metadata":{"id":"74OB9XFhQT0z"}},{"cell_type":"markdown","source":"## **Wordcloud for positive reasons**","metadata":{"id":"Gi25bGX5Rcv7"}},{"cell_type":"code","source":"from wordcloud import WordCloud,STOPWORDS","metadata":{"id":"FHfE5Ru_RznH","execution":{"iopub.status.busy":"2022-08-25T20:23:02.688806Z","iopub.execute_input":"2022-08-25T20:23:02.689546Z","iopub.status.idle":"2022-08-25T20:23:02.725566Z","shell.execute_reply.started":"2022-08-25T20:23:02.68951Z","shell.execute_reply":"2022-08-25T20:23:02.724084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"new_df=data[data['airline_sentiment']=='positive']\nwords = ' '.join(new_df['text'])\ncleaned_word = \" \".join([word for word in words.split()\n                            if 'http' not in word\n                                and not word.startswith('@')\n                                and word != 'RT'\n                            ])\nwordcloud = WordCloud(stopwords=STOPWORDS,\n                      background_color='black',\n                      width=3000,\n                      height=2500\n                     ).generate(cleaned_word)\nplt.figure(1,figsize=(12, 12))\nplt.imshow(wordcloud)\nplt.axis('off')\nplt.show()","metadata":{"id":"oqmI8nxURYSW","outputId":"db424286-6098-4fb0-de05-8f530b021fa0","execution":{"iopub.status.busy":"2022-08-25T20:23:03.130383Z","iopub.execute_input":"2022-08-25T20:23:03.130801Z","iopub.status.idle":"2022-08-25T20:23:19.752514Z","shell.execute_reply.started":"2022-08-25T20:23:03.130761Z","shell.execute_reply":"2022-08-25T20:23:19.750857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wordcloud for Negative sentiments of tweets","metadata":{"id":"3lahWy6lRq_L"}},{"cell_type":"code","source":"new_df=data[data['airline_sentiment']=='negative']\nwords = ' '.join(new_df['text'])\ncleaned_word = \" \".join([word for word in words.split()\n                            if 'http' not in word\n                                and not word.startswith('@')\n                                and word != 'RT'\n                            ])\nwordcloud = WordCloud(stopwords=STOPWORDS,\n                      background_color='black',\n                      width=3000,\n                      height=2500\n                     ).generate(cleaned_word)\nplt.figure(1,figsize=(12, 12))\nplt.imshow(wordcloud)\nplt.axis('off')\nplt.show()","metadata":{"id":"cojhLDJMRrwM","outputId":"3d0dca95-bf59-4ab7-aab2-dd1fdf15c90c","execution":{"iopub.status.busy":"2022-08-25T20:23:19.754869Z","iopub.execute_input":"2022-08-25T20:23:19.75527Z","iopub.status.idle":"2022-08-25T20:23:36.268956Z","shell.execute_reply.started":"2022-08-25T20:23:19.755237Z","shell.execute_reply":"2022-08-25T20:23:36.26773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Dropng the rows with neutral sentiments**","metadata":{"id":"8iXmSTRqZy5q"}},{"cell_type":"code","source":"data.drop(data.loc[data['airline_sentiment']=='neutral'].index, inplace=True)","metadata":{"id":"toJnCjjwa85N","execution":{"iopub.status.busy":"2022-08-25T20:23:36.270763Z","iopub.execute_input":"2022-08-25T20:23:36.271837Z","iopub.status.idle":"2022-08-25T20:23:36.284643Z","shell.execute_reply.started":"2022-08-25T20:23:36.271795Z","shell.execute_reply":"2022-08-25T20:23:36.282958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **label encoding on airline_sentiment**","metadata":{"id":"E815CYR1cAMy"}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\nle = LabelEncoder()\nle.fit(data['airline_sentiment'])\n\ndata['airline_sentiment_encoded'] = le.transform(data['airline_sentiment'])\ndata.head()","metadata":{"id":"QmefSE6ub2Xo","outputId":"025836c9-33c6-45cf-977a-57091cb04be1","execution":{"iopub.status.busy":"2022-08-25T20:23:36.287329Z","iopub.execute_input":"2022-08-25T20:23:36.288006Z","iopub.status.idle":"2022-08-25T20:23:36.314504Z","shell.execute_reply.started":"2022-08-25T20:23:36.287968Z","shell.execute_reply":"2022-08-25T20:23:36.313408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Preprocessing the tweet text data**","metadata":{"id":"bpzU8MOfQuOd"}},{"cell_type":"markdown","source":"Now, we will clean the tweet text data and apply classification algorithms on it","metadata":{"id":"SV8b0HkuQy7N"}},{"cell_type":"code","source":"def tweet_to_words(tweet):\n    letters_only = re.sub(\"[^a-zA-Z]\", \" \",tweet) \n    words = letters_only.lower().split()                             \n    stops = set(stopwords.words(\"english\"))                  \n    meaningful_words = [w for w in words if not w in stops] \n    return( \" \".join( meaningful_words ))","metadata":{"id":"9_Hw4yp_QNs2","execution":{"iopub.status.busy":"2022-08-25T20:23:36.315806Z","iopub.execute_input":"2022-08-25T20:23:36.317105Z","iopub.status.idle":"2022-08-25T20:23:36.325299Z","shell.execute_reply.started":"2022-08-25T20:23:36.317069Z","shell.execute_reply":"2022-08-25T20:23:36.323904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nltk.download('stopwords')\ndata['clean_tweet']=data['text'].apply(lambda x: tweet_to_words(x))","metadata":{"id":"Be5lq5exQ2Eo","outputId":"3123a313-9b64-4603-899d-9511a474d646","execution":{"iopub.status.busy":"2022-08-25T20:23:36.326665Z","iopub.execute_input":"2022-08-25T20:23:36.327205Z","iopub.status.idle":"2022-08-25T20:23:38.531257Z","shell.execute_reply.started":"2022-08-25T20:23:36.327168Z","shell.execute_reply":"2022-08-25T20:23:38.530074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Vectorization**","metadata":{"id":"xW8yjSQ1WwNS"}},{"cell_type":"code","source":"x = data.clean_tweet\ny = data.airline_sentiment\n\nprint(len(x), len(y))","metadata":{"id":"BAbhoS5HWnJb","outputId":"15518f77-2806-4391-d601-c1de175a77b0","execution":{"iopub.status.busy":"2022-08-25T20:23:38.532873Z","iopub.execute_input":"2022-08-25T20:23:38.533264Z","iopub.status.idle":"2022-08-25T20:23:38.540398Z","shell.execute_reply.started":"2022-08-25T20:23:38.533229Z","shell.execute_reply":"2022-08-25T20:23:38.539167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The data is split in the standard 80,20 ratio","metadata":{"id":"HnxSANA1Ssof"}},{"cell_type":"code","source":"x_train, x_test, y_train, y_test = train_test_split(x, y, random_state=42)\nprint(len(x_train), len(y_train))\nprint(len(x_test), len(y_test))","metadata":{"id":"uzQrwwy7UEK_","outputId":"410dbf4a-fe1d-40f0-f470-21e0058cde33","execution":{"iopub.status.busy":"2022-08-25T20:23:38.542292Z","iopub.execute_input":"2022-08-25T20:23:38.542642Z","iopub.status.idle":"2022-08-25T20:23:38.556351Z","shell.execute_reply.started":"2022-08-25T20:23:38.542611Z","shell.execute_reply":"2022-08-25T20:23:38.555169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.feature_extraction.text import CountVectorizer\n\n# instantiate the vectorizer\nvect = CountVectorizer()\nvect.fit(x_train)","metadata":{"id":"-qzI5IbgQ7_q","outputId":"a8cf7911-b4b5-44f1-ee63-ea2ab6bb8278","execution":{"iopub.status.busy":"2022-08-25T20:23:38.557982Z","iopub.execute_input":"2022-08-25T20:23:38.558854Z","iopub.status.idle":"2022-08-25T20:23:38.689599Z","shell.execute_reply.started":"2022-08-25T20:23:38.558807Z","shell.execute_reply":"2022-08-25T20:23:38.688465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Use the trained to create a document-term matrix from train and test sets\nx_train_dtm = vect.transform(x_train)\nx_test_dtm = vect.transform(x_test)","metadata":{"id":"de2ezWTtUujF","execution":{"iopub.status.busy":"2022-08-25T20:23:38.692811Z","iopub.execute_input":"2022-08-25T20:23:38.69316Z","iopub.status.idle":"2022-08-25T20:23:38.83673Z","shell.execute_reply.started":"2022-08-25T20:23:38.693109Z","shell.execute_reply":"2022-08-25T20:23:38.835287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vect_tunned = CountVectorizer(stop_words='english', ngram_range=(1,2), min_df=0.1, max_df=0.7, max_features=100)\nvect_tunned","metadata":{"id":"bJdBrQBJUufk","outputId":"a4737965-dba2-4bd1-c905-f426143c356c","execution":{"iopub.status.busy":"2022-08-25T20:23:38.838367Z","iopub.execute_input":"2022-08-25T20:23:38.839175Z","iopub.status.idle":"2022-08-25T20:23:38.847832Z","shell.execute_reply.started":"2022-08-25T20:23:38.839133Z","shell.execute_reply":"2022-08-25T20:23:38.846516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Model Building**","metadata":{"id":"PfblobCSYP01"}},{"cell_type":"code","source":"#training SVM model with linear kernel\n#Support Vector Classification-wrapper around SVM\nfrom sklearn.svm import SVC\nmodel = SVC(kernel='linear', random_state = 10)\nmodel.fit(x_train_dtm, y_train)\n#predicting output for test data\npred = model.predict(x_test_dtm)","metadata":{"id":"M9-7O1udXmq0","execution":{"iopub.status.busy":"2022-08-25T20:23:38.849354Z","iopub.execute_input":"2022-08-25T20:23:38.849718Z","iopub.status.idle":"2022-08-25T20:23:43.2705Z","shell.execute_reply.started":"2022-08-25T20:23:38.849684Z","shell.execute_reply":"2022-08-25T20:23:43.269184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#accuracy score\naccuracy_score(y_test,pred)","metadata":{"id":"QNQxYTxAeWXd","outputId":"c9dda2b2-2471-4b3d-d5c7-e30d137342db","execution":{"iopub.status.busy":"2022-08-25T20:23:43.272101Z","iopub.execute_input":"2022-08-25T20:23:43.273253Z","iopub.status.idle":"2022-08-25T20:23:43.288334Z","shell.execute_reply.started":"2022-08-25T20:23:43.273205Z","shell.execute_reply":"2022-08-25T20:23:43.286967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#building confusion matrix\ncm = confusion_matrix(y_test, pred)\ncm","metadata":{"id":"AWIC_yglZHiq","outputId":"67d6ed86-c790-4747-beaf-f7e6d69caf46","execution":{"iopub.status.busy":"2022-08-25T20:23:43.289934Z","iopub.execute_input":"2022-08-25T20:23:43.290371Z","iopub.status.idle":"2022-08-25T20:23:43.313192Z","shell.execute_reply.started":"2022-08-25T20:23:43.290334Z","shell.execute_reply":"2022-08-25T20:23:43.312053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#defining the size of the canvas\nplt.rcParams['figure.figsize'] = [15,8]\n#confusion matrix to DataFrame\nconf_matrix = pd.DataFrame(data = cm,columns = ['Predicted:0','Predicted:1',], index = ['Actual:0','Actual:1',])\n#plotting the confusion matrix\nsns.heatmap(conf_matrix, annot = True, fmt = 'd', cmap = 'Paired', cbar = False,linewidths = 0.1, annot_kws = {'size':25})\nplt.xticks(fontsize = 20)\nplt.yticks(fontsize = 20)\nplt.show()\n","metadata":{"id":"oqfXofeIcf5H","outputId":"2fdf39cf-f08c-462b-9bdf-904b4042eec6","execution":{"iopub.status.busy":"2022-08-25T20:23:43.314525Z","iopub.execute_input":"2022-08-25T20:23:43.315594Z","iopub.status.idle":"2022-08-25T20:23:43.486717Z","shell.execute_reply.started":"2022-08-25T20:23:43.315555Z","shell.execute_reply":"2022-08-25T20:23:43.485318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test,pred))","metadata":{"id":"zVorRH37ctkF","outputId":"058901e2-1945-4201-a2d0-9f3f8668da32","execution":{"iopub.status.busy":"2022-08-25T20:23:43.488101Z","iopub.execute_input":"2022-08-25T20:23:43.488786Z","iopub.status.idle":"2022-08-25T20:23:43.584844Z","shell.execute_reply.started":"2022-08-25T20:23:43.488726Z","shell.execute_reply":"2022-08-25T20:23:43.58369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* As we you can see above we have plotted the confusion matrix for predicted sentiments and actual sentiments (negative and positive)\n* SVM Classifier gives us the best accuracy score i.e 91% precision scores according to the classification report.\n* The confusion matrix shows the TP,TN,FP,FN for sentiments(negative, positive)","metadata":{"id":"0EWSAXAPdaGf"}},{"cell_type":"markdown","source":"# **Thank You**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}