{"metadata":{"kernelspec":{"name":"ir","display_name":"R","language":"R"},"language_info":{"name":"R","codemirror_mode":"r","pygments_lexer":"r","mimetype":"text/x-r-source","file_extension":".r","version":"4.0.5"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This R environment comes with many helpful analytics packages installed\n# It is defined by the kaggle/rstats Docker image: https://github.com/kaggle/docker-rstats\n# For example, here's a helpful package to load\n\nlibrary(tidyverse) # metapackage of all tidyverse packages\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nlist.files(path = \"../input\")\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"051d70d956493feee0c6d64651c6a088724dca2a","_execution_state":"idle","execution":{"iopub.status.busy":"2023-06-01T02:44:06.383877Z","iopub.execute_input":"2023-06-01T02:44:06.386285Z","iopub.status.idle":"2023-06-01T02:44:07.698037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Source\n- https://www.rdocumentation.org/packages/arrow/versions/3.0.0\n- https://r4ds.hadley.nz/arrow.html\n- https://github.com/brews/bayfoxr\n- http://www.sthda.com/english/wiki/be-awesome-in-ggplot2-a-practical-guide-to-be-highly-effective-r-software-and-data-visualization\n- https://www.rdocumentation.org/packages/glmnet/versions/4.1-7/topics/glmnet\n- https://bookdown.org/rdpeng/exdata/the-base-plotting-system-1.html","metadata":{}},{"cell_type":"markdown","source":"- The goal of this competition is to detect and translate American Sign Language (ASL) fingerspelling into text.\n\nThis competition requires submissions to be made in the form of TensorFlow Lite models. You are welcome to train your model using the framework of your choice as long as you convert the model checkpoint into the tflite format prior to submission. Please see the evaluation page for details.\n\n### Files - [train/supplemental_metadata].csv\n - path - The path to the landmark file.\n - file_id - A unique identifier for the data file.\n - participant_id - A unique identifier for the data contributor.\n - sequence_id - A unique identifier for the landmark sequence. Each data file may contain many sequences.\n - phrase - The labels for the landmark sequence. The train and test datasets contain randomly generated addresses, phone numbers, and urls derived from components of real addresses/phone numbers/urls. Any overlap with real addresses, phone numbers, or urls is purely accidental. The supplemental dataset consists of fingerspelled sentences. Note that some of the urls include adult content. The intent of this competition is to support the Deaf and Hard of Hearing community in engaging with technology on an equal footing with other adults.","metadata":{}},{"cell_type":"code","source":"devtools::install_github(\"brews/bayfoxr\")\n\nlibrary(bayfoxr)\n","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:07.702255Z","iopub.execute_input":"2023-06-01T02:44:07.736384Z","iopub.status.idle":"2023-06-01T02:44:29.323055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(arrow)\nlandmarks <- read_parquet(\"/kaggle/input/asl-fingerspelling/supplemental_landmarks/1118603411.parquet\")\nhead(landmarks,4)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:29.359768Z","iopub.execute_input":"2023-06-01T02:44:29.361473Z","iopub.status.idle":"2023-06-01T02:44:40.060862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"landmarks <- na.omit(landmarks)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:40.063491Z","iopub.execute_input":"2023-06-01T02:44:40.064976Z","iopub.status.idle":"2023-06-01T02:44:41.712219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"str(landmarks,8)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:41.716171Z","iopub.execute_input":"2023-06-01T02:44:41.718694Z","iopub.status.idle":"2023-06-01T02:44:41.816061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"quantile(landmarks$x_face_0, probs = c(0.05, 0.50, 0.95))","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:41.818663Z","iopub.execute_input":"2023-06-01T02:44:41.820264Z","iopub.status.idle":"2023-06-01T02:44:41.838822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(landmarks$x_face_0)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:41.841736Z","iopub.execute_input":"2023-06-01T02:44:41.843102Z","iopub.status.idle":"2023-06-01T02:44:42.149319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"qplot(landmarks$x_face_0, landmarks$x_face_30, data = landmarks, \n      geom = c(\"point\", \"smooth\"))","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:42.153529Z","iopub.execute_input":"2023-06-01T02:44:42.155899Z","iopub.status.idle":"2023-06-01T02:44:44.160825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data <- ggplot(landmarks, aes(x = landmarks$x_face_51))\ndata + geom_bar(fill = \"steelblue\", color =\"steelblue\") +\n  theme_minimal()\n","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:44.164382Z","iopub.execute_input":"2023-06-01T02:44:44.166636Z","iopub.status.idle":"2023-06-01T02:44:44.661154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"read_parquet_schema <- function (file, col_select = NULL, as_data_frame = TRUE, props = ParquetArrowReaderProperties$create(), \n                                 ...) \n{\n  require(arrow)\n  reader <- ParquetFileReader$create(file, props = props, ...)\n  schema <- reader$GetSchema()\n  names <- names(schema)\n  return(names)\n}","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:44.66536Z","iopub.execute_input":"2023-06-01T02:44:44.667375Z","iopub.status.idle":"2023-06-01T02:44:44.683111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Column Names\nfile <- \"/kaggle/input/asl-fingerspelling/train_landmarks/1098899348.parquet\"\ncol_names <- read_parquet_schema(file)\nhead(col_names)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:44.688126Z","iopub.execute_input":"2023-06-01T02:44:44.690876Z","iopub.status.idle":"2023-06-01T02:44:44.768081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(jsonlite)\nparquet_df <- read_parquet(\"/kaggle/input/asl-fingerspelling/supplemental_landmarks/1112747136.parquet\")\njson_df <- read_json(\"/kaggle/input/asl-fingerspelling/character_to_prediction_index.json\")\n\n# Merge \nmerged_df <- merge(parquet_df, json_df)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:44.770592Z","iopub.execute_input":"2023-06-01T02:44:44.771978Z","iopub.status.idle":"2023-06-01T02:44:54.866153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"head(merged_df)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:54.868688Z","iopub.execute_input":"2023-06-01T02:44:54.870196Z","iopub.status.idle":"2023-06-01T02:44:55.592725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(table(sum(is.na(merged_df))))","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:55.595317Z","iopub.execute_input":"2023-06-01T02:44:55.596695Z","iopub.status.idle":"2023-06-01T02:44:58.78249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"merged_df <- na.omit(merged_df)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:44:58.784936Z","iopub.execute_input":"2023-06-01T02:44:58.78625Z","iopub.status.idle":"2023-06-01T02:45:01.087319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ggplot(data = merged_df, aes(x =frame, y = x_face_4)) +\n  geom_point() +\n  theme(panel.background = element_rect(fill = \"green\"),\n        text = element_text(colour = \"white\"))","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:01.089811Z","iopub.execute_input":"2023-06-01T02:45:01.091117Z","iopub.status.idle":"2023-06-01T02:45:01.410748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data <- ggplot(merged_df, aes(x = frame))\ndata +  geom_bar(fill = \"steelblue\", color =\"steelblue\") +\n  theme_minimal()","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:01.413319Z","iopub.execute_input":"2023-06-01T02:45:01.414779Z","iopub.status.idle":"2023-06-01T02:45:01.75617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Empirical Cumulative Density Function\ndata +  stat_ecdf()","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:01.758699Z","iopub.execute_input":"2023-06-01T02:45:01.760217Z","iopub.status.idle":"2023-06-01T02:45:02.049611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data + stat_count()","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:02.052161Z","iopub.execute_input":"2023-06-01T02:45:02.053535Z","iopub.status.idle":"2023-06-01T02:45:02.323639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ggplot(merged_df, aes(frame, x_face_6)) +\n  geom_point() + geom_quantile() +\n  theme_minimal()","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:02.326149Z","iopub.execute_input":"2023-06-01T02:45:02.327481Z","iopub.status.idle":"2023-06-01T02:45:03.087375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data + stat_boxplot(coeff = 3.5)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:03.091385Z","iopub.execute_input":"2023-06-01T02:45:03.094782Z","iopub.status.idle":"2023-06-01T02:45:03.371932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ggplot(merged_df, aes(frame, x_face_6)) +\n  geom_jitter(aes(color = x_face_2), size = 0.5)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:03.374796Z","iopub.execute_input":"2023-06-01T02:45:03.407645Z","iopub.status.idle":"2023-06-01T02:45:03.771984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data <- ggplot(merged_df, aes(x = frame, y = x_face_6, \n                     ymin = x_face_6-x_face_5, ymax = x_face_6+x_face_5))\ndata + geom_errorbar(aes(color = x_face_1),  position = \"dodge\")","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:03.775143Z","iopub.execute_input":"2023-06-01T02:45:03.776725Z","iopub.status.idle":"2023-06-01T02:45:04.455464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(reshape2)\ncormat <- melt(merged_df)\nhead(cormat)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:45:04.459182Z","iopub.execute_input":"2023-06-01T02:45:04.46129Z","iopub.status.idle":"2023-06-01T02:45:04.573397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with(merged_df, plot(frame,x_face_1))","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:13:28.644694Z","iopub.execute_input":"2023-06-01T03:13:28.646371Z","iopub.status.idle":"2023-06-01T03:13:28.84871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lasso Regression\nlibrary(glmnet)\n#define response variable\ny <- merged_df$frame\n\n#define matrix of predictor variables\nx <- data.matrix(merged_df[, c('x_face_0', 'x_face_1', 'x_face_2', 'x_face_3')])\n","metadata":{"execution":{"iopub.status.busy":"2023-06-01T02:52:23.443516Z","iopub.execute_input":"2023-06-01T02:52:23.445091Z","iopub.status.idle":"2023-06-01T02:52:23.463144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"library(glmnet)\n\n#perform k-fold cross-validation to find optimal lambda value\nmodel <- cv.glmnet(x, y,family = c(\"gaussian\"),alpha = 1)\n# optimal lambda value\noptimal_lambda <- model$lambda.min\noptimal_lambda","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:10:27.995573Z","iopub.execute_input":"2023-06-01T03:10:28.001231Z","iopub.status.idle":"2023-06-01T03:10:28.211528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(model) \n","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:10:37.154246Z","iopub.execute_input":"2023-06-01T03:10:37.187787Z","iopub.status.idle":"2023-06-01T03:10:37.295406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Coefficients of model\nmodel_coef <- glmnet(x, y, alpha = 1,lambda = optimal_lambda)\ncoef(model_coef)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:11:07.403742Z","iopub.execute_input":"2023-06-01T03:11:07.405392Z","iopub.status.idle":"2023-06-01T03:11:07.429252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#use fitted best model to make predictions\ny_predicted <- predict(model, s = optimal_lambda, newx = x)\n\n#find SST and SSE\nsst <- sum((y - mean(y))^2)\nsse <- sum((y_predicted - y)^2)\n\n#find R-Squared\nrsq <- 1 - sse/sst\nrsq","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:11:10.343993Z","iopub.execute_input":"2023-06-01T03:11:10.346788Z","iopub.status.idle":"2023-06-01T03:11:10.379061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot(y_predicted)","metadata":{"execution":{"iopub.status.busy":"2023-06-01T03:11:17.884314Z","iopub.execute_input":"2023-06-01T03:11:17.885971Z","iopub.status.idle":"2023-06-01T03:11:18.057086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}