<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Rsquared Academy Blog</title>
<link>https://blog.rsquaredacademy.com/data-wrangling/</link>
<atom:link href="https://blog.rsquaredacademy.com/data-wrangling/index.xml" rel="self" type="application/rss+xml"/>
<description>Import, tidy, and reshape data with dplyr, tidyr, stringr, and friends.</description>
<generator>quarto-1.10.18</generator>
<lastBuildDate>Fri, 14 Jan 2022 00:00:00 GMT</lastBuildDate>
<item>
  <title>Handling Categorical Data in R - Part 4</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2022-01-14-handling-categorical-data-in-r-part-4.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<script src="../../rmarkdown-libs/header-attrs/header-attrs.js"></script>
<p>
<img src="https://blog.rsquaredacademy.com/img/forcats-part-4.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
This is part 4 of a series on “Handling Categorical Data in R” where we are learning to <strong>read</strong>, <strong>store</strong>, <strong>summarize</strong>, <strong>reshape</strong> &amp; <strong>visualize</strong> categorical data.
</p>
<p>
Below are the links to the other articles of this series:
</p>
<ul>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-1/">Part 1 - Introduction to Factor</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-2/">Part 2 - Summarize Categorical Data</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-3/">Part 3 - Reshape Categorical Data</a>
</li>
</ul>
<p>
In this article, we will explore the different ways of visualizing categorical data using <a href="https://ggplot2.tidyverse.org/">ggplot2</a>.
</p>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
You can download all the data sets, R scripts, practice questions and their solutions from our <a href="https://github.com/rsquaredacademy-education/online-courses/">GitHub</a> repository.
</p>
</section>
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In this section, we will learn to visualize categorical data. We will look at the following type of plots:
</p>
<ul>
<li>
univariate bar plot
</li>
<li>
bivariate bar plot
<ul>
<li>
grouped
</li>
<li>
stacked
</li>
<li>
proportional
</li>
</ul>
</li>
<li>
mosaic plot
</li>
<li>
pie chart
</li>
<li>
donut chart
</li>
</ul>
<p>
We will be using <a href="https://ggplot2.tidyverse.org">ggplot2</a> package throughout this article. So you should know the basics of data visualization with ggplot2. If you are new to or have never used ggplot2, do not worry. We have several <a href="https://blog.rsquaredacademy.com/data-visualization/index.html">tutorials</a> and an <a href="https://viz-ggplot2.rsquaredacademy.com/">ebook</a> on ggplot2, you can go through them first and then come back to this article. Let us read the case study data before we start our visualization journey.
</p>
<pre class="r"><code># read data
data &lt;- readRDS('analytics.rds')</code></pre>
</section>
<section id="bar-plot" class="level2">
<h2 class="anchored" data-anchor-id="bar-plot">
Bar Plot
</h2>
<p>
Bar charts provide a visual representation of categorical data. The bars can be plotted either vertically or horizontally. The categories/groups appear along the horizontal X axis and the height of the bar represents a measured value.
</p>
<pre class="r"><code>ggplot(data) +
  geom_bar(aes(x = device), fill = "blue") +
  xlab("Device") + ylab("Count")</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/bar_plot-1.png" width="672">
</p>
<p>
In the above example, the bars represent the count/frequency of the categories. If the bars represent continuous data, the value could be <em>mean</em> or <em>sum</em> of the variable being represented.
</p>
<section id="grouped-bar-plot" class="level3">
<h3 class="anchored" data-anchor-id="grouped-bar-plot">
Grouped Bar Plot
</h3>
<p>
A grouped bar chart plots values for two levels of a categorical variable instead of one. You should use grouped bar chart when making comparisons across different categories of data. Use it when you want to look at how the second category variable changes within each level of the first and vice versa.
</p>
<pre class="r"><code>ggplot(data) +
  geom_bar(aes(x = device, fill = gender), position = "dodge") +
  xlab("Device") + ylab("Count")</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/grouped_bar_plot-1.png" width="672">
</p>
</section>
<section id="stacked-bar-plot" class="level3">
<h3 class="anchored" data-anchor-id="stacked-bar-plot">
Stacked Bar Plot
</h3>
<p>
In stacked bar plots, the bars are stacked on top of each other instead of placing them next to each other. Use stacked bar plots while looking at cumulative value.
</p>
<pre class="r"><code>ggplot(data) +
  geom_bar(aes(x = device, fill = gender)) +
  xlab("Device") + ylab("Count")</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/stacked_bar_plot-1.png" width="672">
</p>
</section>
<section id="proportional-bar-plot" class="level3">
<h3 class="anchored" data-anchor-id="proportional-bar-plot">
Proportional Bar Plot
</h3>
<p>
Also known as percent stacked plot, the height of all bars in this plot are the same. The distribution of the second categorical variable is scaled to 1 or 100. The length of each bar is determined by its share in the category. Use this when you want to concurrently observe each of several variables as they fluctuate and as their percentage ratio’s change.
</p>
<pre class="r"><code>data %&gt;% 
  select(device, gender) %&gt;% 
  table() %&gt;% 
  tibble::as_tibble() %&gt;% 
  ggplot(aes(x = device, y = n, fill = gender)) +
  geom_bar(stat = "identity", position = "fill") +
  xlab("Device") + ylab("Gender")</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/prop_bar_plot-1.png" width="672">
</p>
</section>
</section>
<section id="mosaic-plot" class="level2">
<h2 class="anchored" data-anchor-id="mosaic-plot">
Mosaic Plot
</h2>
<p>
A mosaic plot is a graphical representation of a two way table or contingency table. It was introduced by Hartigan &amp; Kleiner and is divided into rectangles. Proportions on horizontal axis represents the number of observations for each level of the X variable. The vertical length of each rectangle is proportional to the proportion of Y variable in each level of X variable.
</p>
<pre class="r"><code>ggplot(data = data) +
  geom_mosaic(aes(x = product(channel, device), fill = channel)) +
  xlab("Device") + ylab("Channel")</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/mosaic_plot-1.png" width="672">
</p>
</section>
<section id="pie-chart" class="level2">
<h2 class="anchored" data-anchor-id="pie-chart">
Pie Chart
</h2>
<p>
Pie chart is a circular chart, divided into slices to show relevant sizes of data. It shows the distribution of the different levels of a categorical variable as a circle is divided into radial slices. Each level corresponds with a single slice of the circle and size indicates the proportion of the level. Use it when comparing each group’s contribution to the whole as opposed to comparing groups to each other.
</p>
<section id="base-r" class="level3">
<h3 class="anchored" data-anchor-id="base-r">
Base R
</h3>
<pre class="r"><code>data %&gt;% 
  pull(device) %&gt;%
  table() %&gt;%
  pie()</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/pie_chart-1.png" width="672">
</p>
</section>
<section id="d-pie-chart" class="level3">
<h3 class="anchored" data-anchor-id="d-pie-chart">
3D Pie Chart
</h3>
<pre class="r"><code>data %&gt;% 
  pull(device) %&gt;% 
  table() %&gt;% 
  pie3D(explode = 0.1)</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/pie_chart_3d-1.png" width="672">
</p>
</section>
<section id="ggplot2" class="level3">
<h3 class="anchored" data-anchor-id="ggplot2">
ggplot2
</h3>
<pre class="r"><code>data %&gt;% 
  pull(device) %&gt;% 
  fct_count() %&gt;% 
  rename(device = f, count = n) %&gt;% 
  ggplot() +
  geom_bar(aes(x = "", y = count, fill = device), width = 1, stat = "identity") +
  coord_polar("y", start = 0)</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/gg_pie_chart-1.png" width="672">
</p>
</section>
</section>
<section id="donut-chart" class="level2">
<h2 class="anchored" data-anchor-id="donut-chart">
Donut Chart
</h2>
<p>
Donut chart is a variation of the pie chart. It has a round hole in the middle which makes it look like a donut. The focus is on the length of the arcs and not the proportions of the slices. Blank spaces inside donut chart can be used to display information inside it.
</p>
<pre class="r"><code>data %&gt;% 
  pull(device) %&gt;% 
  fct_count() %&gt;% 
  rename(device = f, count = n) %&gt;% 
  ggdonutchart("count", label = "device", fill = "device", color = "white",
               palette = c("#00AFBB", "#E7B800", "#FC4E07"))</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/gg_donut_chart-1.png" width="672">
</p>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">
Summary
</h2>
<ul>
<li>
Bar charts provide a visual representation of categorical data.
</li>
<li>
Use grouped bar chart to make comparison against different categories of data.
</li>
<li>
Use stacked bar chart while looking at cumulative data.
</li>
<li>
Use proportional bar chart when you want to concurrently observe each of the several variables as they fluctuate.
</li>
<li>
Use mosaic plot to discover associations between two variables.
</li>
<li>
Use pie chart and donut chart when comparing each group’s contribution to the whole.
</li>
</ul>
</section>
<section id="your-turn" class="level2">
<h2 class="anchored" data-anchor-id="your-turn">
Your Turn…
</h2>
<p>
Generate all the below plots:
</p>
<ol style="list-style-type: decimal">
<li>
Bar plot of <code>channel</code>
</li>
</ol>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_1-1.png" width="672" style="display: block; margin: auto;">
</p>
<ol start="2" style="list-style-type: decimal">
<li>
Display grouped bar plot of <code>user_type</code> by <code>channel</code>
</li>
</ol>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_2-1.png" width="672" style="display: block; margin: auto;">
</p>
<ol start="3" style="list-style-type: decimal">
<li>
Display stacked bar plot of <code>channel</code> by <code>gender</code>
</li>
</ol>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_3-1.png" width="672" style="display: block; margin: auto;">
</p>
<ol start="4" style="list-style-type: decimal">
<li>
Display proportional bar plot of <code>channel</code> by <code>device</code>
</li>
</ol>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_4-1.png" width="672" style="display: block; margin: auto;">
</p>
<ol start="5" style="list-style-type: decimal">
<li>
Display mosaic plot of <code>device</code> by <code>channel</code>
</li>
</ol>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_5-1.png" width="672" style="display: block; margin: auto;">
</p>
<ol start="6" style="list-style-type: decimal">
<li>
Display pie or donut chart of <code>channel</code>
</li>
</ol>
<section id="pie-chart-1" class="level5">
<h5 class="anchored" data-anchor-id="pie-chart-1">
6.1 Pie Chart
</h5>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_6-1.png" width="672" style="display: block; margin: auto;">
</p>
</section>
<section id="d-pie-chart-1" class="level5">
<h5 class="anchored" data-anchor-id="d-pie-chart-1">
6.2 3D Pie Chart
</h5>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_7-1.png" width="672" style="display: block; margin: auto;">
</p>
</section>
<section id="pie-chart-ggplot2" class="level5">
<h5 class="anchored" data-anchor-id="pie-chart-ggplot2">
6.3 Pie Chart (ggplot2)
</h5>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_8-1.png" width="672" style="display: block; margin: auto;">
</p>
</section>
<section id="donut-chart-1" class="level5">
<h5 class="anchored" data-anchor-id="donut-chart-1">
6.4 Donut Chart
</h5>
<p>
<img src="https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/2022-01-14-handling-categorical-data-in-r-part-4_files/figure-html/plot_9-1.png" width="960" style="display: block; margin: auto;">
</p>
<p>
*As the reader of this blog, you are our most important critic and commentator. We value your opinion and want to know what we are doing right, what we could do better, what areas you would like to see us publish in, and any other words of wisdom you are willing to pass our way.
</p>
<p>
We welcome your comments. You can email to let us know what you did or did not like about our blog as well as what we can do to make our post better.*
</p>
<p>
<strong>Email: <a href="mailto:support@rsquaredacademy.com" class="email">support@rsquaredacademy.com</a></strong>
</p>
</section>
</section>



 ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-4/</guid>
  <pubDate>Fri, 14 Jan 2022 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/forcats-part-4.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>Handling Categorical Data in R - Part 3</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-3/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2022-01-13-handling-categorical-data-in-r-part-3.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<script src="../../rmarkdown-libs/header-attrs/header-attrs.js"></script>
<script src="../../rmarkdown-libs/kePrint/kePrint.js"></script>
<p><link href="../../rmarkdown-libs/lightable/lightable.css" rel="stylesheet"></p>
<p>
<img src="https://blog.rsquaredacademy.com/img/forcats-part-3.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
This is part 3 of a series on “Handling Categorical Data in R” where we are learning to <strong>read</strong>, <strong>store</strong>, <strong>summarize</strong>, <strong>reshape</strong> &amp; <strong>visualize</strong> categorical data.
</p>
<p>
Below are the links to the other articles of this series:
</p>
<ul>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-1/">Part 1 - Introduction to Factor</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-2/">Part 2 - Summarize Categorical Data</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-4/">Part 4 - Visualize Categorical Data</a>
</li>
</ul>
<p>
In this article, we will learn to manipulate/reshape categorical data by changing the value and order of levels/categories.
</p>
<section id="table-of-contents" class="level2">
<h2 class="anchored" data-anchor-id="table-of-contents">
Table of Contents
</h2>
<ul>
<li>
Resources
</li>
<li>
Introduction
</li>
<li>
How to change value of levels?
</li>
<li>
How to add or remove levels?
</li>
<li>
How to change order of levels?
</li>
<li>
Practice Questions
</li>
</ul>
</section>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
You can download all the data sets, R scripts, practice questions and their solutions from our <a href="https://github.com/rsquaredacademy-education/online-courses/">GitHub</a> repository.
</p>
</section>
<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">
Introduction
</h2>
<p>
In this section, our focus will be on handling the levels of a categorical variable, and exploring the <a href="https://forcats.tidyverse.org/">forcats</a> package for the same. We will basically look at 3 key operations or transformations we would like to do when it comes to factors which are:
</p>
<ul>
<li>
change value of levels
</li>
<li>
add or remove levels
</li>
<li>
change order of levels
</li>
</ul>
<p>
Before we start working with the value of the levels, let us read the case study data and take a quick look at some of the functions we used in the previous articles.
</p>
<pre class="r"><code># read data
data &lt;- readRDS('analytics.rds')</code></pre>
<p>
We will store the source of traffic as <code>channel</code> instead of referring to the column in the <code>data.frame</code> every time.
</p>
<pre class="r"><code>channel &lt;- data$channel</code></pre>
<p>
Let us go back to the function we used for tabulating data, <code>fct_count()</code>. If you observe the result, it is in the same order as displayed by <code>levels()</code>.
</p>
<pre class="r"><code>fct_count(channel)</code></pre>
<pre><code>## # A tibble: 8 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 (Other)          6073
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Display          3375
## 5 Organic Search 139668
## 6 Paid Search      4395
## 7 Referral        35615
## 8 Social           8031</code></pre>
<p>
If you want to sort the results by the count i.e.&nbsp;most common level comes at the top, use the <code>sort</code> argument.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_sort.png" width="90%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_count(channel, sort = TRUE)</code></pre>
<pre><code>## # A tibble: 8 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 Organic Search 139668
## 2 Direct          39853
## 3 Referral        35615
## 4 Social           8031
## 5 Affiliates       7388
## 6 (Other)          6073
## 7 Paid Search      4395
## 8 Display          3375</code></pre>
<p>
If you want to view the proportion along with the count, set the <code>prop</code> argument to <code>TRUE</code>.
</p>
<pre class="r"><code>fct_count(channel, prop = TRUE)</code></pre>
<pre><code>## # A tibble: 8 x 3
##   f                   n      p
##   &lt;fct&gt;           &lt;int&gt;  &lt;dbl&gt;
## 1 (Other)          6073 0.0248
## 2 Affiliates       7388 0.0302
## 3 Direct          39853 0.163 
## 4 Display          3375 0.0138
## 5 Organic Search 139668 0.571 
## 6 Paid Search      4395 0.0180
## 7 Referral        35615 0.146 
## 8 Social           8031 0.0329</code></pre>
<p>
One of the important steps in data preparation/sanitization is to check if the levels are valid i.e.&nbsp;only levels which should be present in the data are actually present. <code>fct_match()</code> can be used to check validity of levels. It returns a logical vector if the level is present and an error if not.
</p>
<pre class="r"><code>table(fct_match(channel, "Social"))</code></pre>
<pre><code>## 
##  FALSE   TRUE 
## 236367   8031</code></pre>
</section>
<section id="changevalue" class="level2">
<h2 class="anchored" data-anchor-id="changevalue">
Change Value of Levels
</h2>
<p>
In this section, we will learn how to change the value of the levels. In order to keep it interesting, we will state an objective from our case study and then map it into a function from the <a href="https://forcats.tidyverse.org/">forcats</a> package.
</p>
<section id="combine-both-paid-organic-search-into-a-single-level-search" class="level3">
<h3 class="anchored" data-anchor-id="combine-both-paid-organic-search-into-a-single-level-search">
Combine both Paid &amp; Organic Search into a single level, Search
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_collapse.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In this case, we want to change the value of two levels, <em>Paid Search</em> &amp; <em>Organic Search</em> and give them the common value <em>Search</em>. You can also look at it as collapsing two levels into one. There are two functions we can use here:
</p>
<ul>
<li>
<code>fct_collapse()</code>
</li>
<li>
<code>fct_recode()</code>
</li>
</ul>
<p>
Let us look at <code>fct_collapse()</code> first. After specifying the categorical variable, we specify the new value followed by a character vector of the existing values. Remember, the new value is not enclosed in quotes (single or double) but the existing values must be a <code>character</code> vector.
</p>
<pre class="r"><code>fct_count(
  fct_collapse(
    channel,
    Search = c("Paid Search", "Organic Search")
  )
)</code></pre>
<pre><code>## # A tibble: 7 x 2
##   f               n
##   &lt;fct&gt;       &lt;int&gt;
## 1 (Other)      6073
## 2 Affiliates   7388
## 3 Direct      39853
## 4 Display      3375
## 5 Search     144063
## 6 Referral    35615
## 7 Social       8031</code></pre>
<p>
In the case of <code>fct_recode()</code>, each value being changed must be specified in a new line. Similar to <code>fct_collapse()</code>, the new value is not enclosed in quotes but the existing values must be.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_recode.png" width="90%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_count(
  fct_recode(
    channel,
    Search = "Paid Search",
    Search = "Organic Search"
  )
)</code></pre>
<pre><code>## # A tibble: 7 x 2
##   f               n
##   &lt;fct&gt;       &lt;int&gt;
## 1 (Other)      6073
## 2 Affiliates   7388
## 3 Direct      39853
## 4 Display      3375
## 5 Search     144063
## 6 Referral    35615
## 7 Social       8031</code></pre>
<p>
The <a href="https://dplyr.tidyverse.org/">dplyr</a> and <a href="https://cran.r-project.org/package=car">car</a> packages also have recode functions.
</p>
</section>
<section id="retain-only-those-channels-which-have-driven-a-minimum-traffic-of-5000-to-the-website" class="level3">
<h3 class="anchored" data-anchor-id="retain-only-those-channels-which-have-driven-a-minimum-traffic-of-5000-to-the-website">
Retain only those channels which have driven a minimum traffic of 5000 to the website
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_lump_1.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
Instead of having all the channels, we desire to retain only those channels which have driven at least <em>5000</em> visits to the website. What about the rest of the channels which have driven less than <em>5000</em>? We will recategorize them as <strong>Other</strong>. Keep in mind that we already have a <strong>(Other)</strong> level in our data. <code>fct_lump_min()</code> will lump together all levels which do not have a minimum count specified. In our case study, only <strong>Display</strong> drives less than <em>5000</em> visits and it will be categorized into <strong>Other</strong>.
</p>
<pre class="r"><code>fct_count(fct_lump_min(channel, 5000))</code></pre>
<pre><code>## # A tibble: 7 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 (Other)          6073
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Organic Search 139668
## 5 Referral        35615
## 6 Social           8031
## 7 Other            7770</code></pre>
</section>
<section id="retain-only-top-3-referring-channels-and-categorize-rest-into-other" class="level3">
<h3 class="anchored" data-anchor-id="retain-only-top-3-referring-channels-and-categorize-rest-into-other">
Retain only top 3 referring channels and categorize rest into Other
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_lump_2.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
Suppose you decide to retain only the top 3 channels in terms of the traffic driven to the website. In our case study, these are <strong>Direct</strong>, <strong>Organic Search</strong> and <strong>Referral</strong>. We want to retain these 3 levels and categorize the rest as <strong>Other</strong>. <code>fct_lump_n()</code> will retain top <code>n</code> levels by count/frequency and lump the rest into <strong>Other</strong>.
</p>
<pre class="r"><code>fct_count(fct_lump_n(channel, 3))</code></pre>
<pre><code>## # A tibble: 4 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 Direct          39853
## 2 Organic Search 139668
## 3 Referral        35615
## 4 Other           29262</code></pre>
<p>
In our case, <code>n</code> is 3 and hence the top 3 channels in terms of traffic driven are retained while the rest are lumped into <strong>Other</strong>.
</p>
</section>
<section id="retain-only-those-channels-which-have-driven-at-least-2-of-the-overall-traffic" class="level3">
<h3 class="anchored" data-anchor-id="retain-only-those-channels-which-have-driven-at-least-2-of-the-overall-traffic">
Retain only those channels which have driven at least 2% of the overall traffic
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_lump_3.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In the second scenario above, we retained channels based on minimum traffic driven by them to the website. The criteria was count of visits. If you want to specify the criteria as a percentage or proportion instead of count, use <code>fct_lump_prop()</code>. The criteria is a value between <strong>0</strong> and <strong>1</strong>. In our case study, we want to retain channels that have driven at least <strong>2%</strong> of the overall traffic. Hence, we have specified the criteria as <code>0.02</code>.
</p>
<pre class="r"><code>fct_count(fct_lump_prop(channel, 0.02))</code></pre>
<pre><code>## # A tibble: 7 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 (Other)          6073
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Organic Search 139668
## 5 Referral        35615
## 6 Social           8031
## 7 Other            7770</code></pre>
<p>
As you can see, only <strong>Display</strong> drives less than <strong>2%</strong> of overall traffic and has been lumped into <strong>Other</strong>.
</p>
</section>
<section id="retain-the-following-channels-and-merge-the-rest-into-other" class="level3">
<h3 class="anchored" data-anchor-id="retain-the-following-channels-and-merge-the-rest-into-other">
Retain the following channels and merge the rest into Other
</h3>
<ul>
<li>
Organic Search
</li>
<li>
Direct
</li>
<li>
Referral
</li>
</ul>
<p>
In the previous scenarios, we have been retaining or lumping channels based on some criteria like count or percentage of traffic driven to the website. In this scenario, we want to retain certain levels by specifying their labels and combine the rest into <strong>Other</strong>. While we can use <code>fct_collapse()</code> or <code>fct_recode()</code>, a more appropriate function would be <code>fct_other()</code>. We will do a comparison of the three functions in a short while.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_others.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
<code>fct_other()</code> has two arguments, <code>keep</code> and <code>drop</code>. <code>keep</code> is used when we know the levels we want to retain and <code>drop</code> is used when we know the levels we want to drop. In this scenario, we know the levels we want to retain and hence we will use the <code>keep</code> argument and specify them. <strong>Organic Search</strong>, <strong>Direct</strong> and <strong>Referral</strong> will be retained while the rest of the channels will be lumped into <strong>Other</strong>.
</p>
<pre class="r"><code>fct_count(
  fct_other(
    channel, 
    keep = c("Organic Search", "Direct", "Referral"))
)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 Direct          39853
## 2 Organic Search 139668
## 3 Referral        35615
## 4 Other           29262</code></pre>
</section>
<section id="merge-the-following-channels-into-other-and-retain-rest-of-them" class="level3">
<h3 class="anchored" data-anchor-id="merge-the-following-channels-into-other-and-retain-rest-of-them">
Merge the following channels into Other and retain rest of them:
</h3>
<ul>
<li>
Display
</li>
<li>
Paid Search
</li>
</ul>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_others_drop.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In this scenario, we know the levels we want to drop and hence we will use the <code>drop</code> argument and specify them. <strong>Display</strong> and <strong>Paid Search</strong> will be lumped into <strong>Other</strong> while the rest of the channels will be retained.
</p>
<pre class="r"><code>fct_count(
  fct_other(
    channel, 
    drop = c("Display", "Paid Search")
  )
)</code></pre>
<pre><code>## # A tibble: 7 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 (Other)          6073
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Organic Search 139668
## 5 Referral        35615
## 6 Social           8031
## 7 Other            7770</code></pre>
<p>
In the previous scenario, we said we will compare <code>fct_other()</code> with <code>fct_collapse()</code> and <code>fct_recode()</code>. Let us use the other two functions as well and see the difference.
</p>
<pre class="r"><code># collapse
fct_count(
  fct_collapse(
  channel,
  Other = c("(Other)", "Affiliate", "Display", "Paid Search", "Social")
  )
)</code></pre>
<pre><code>## Warning: Unknown levels in `f`: Affiliate</code></pre>
<pre><code>## # A tibble: 5 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 Other           21874
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Organic Search 139668
## 5 Referral        35615</code></pre>
<pre class="r"><code># recode
fct_count(
  fct_recode(
  channel,
  Other = "(Other)", 
  Other = "Affiliate", 
  Other = "Display", 
  Other = "Paid Search", 
  Other = "Social"
  )
)</code></pre>
<pre><code>## Warning: Unknown levels in `f`: Affiliate</code></pre>
<pre><code>## # A tibble: 5 x 2
##   f                   n
##   &lt;fct&gt;           &lt;int&gt;
## 1 Other           21874
## 2 Affiliates       7388
## 3 Direct          39853
## 4 Organic Search 139668
## 5 Referral        35615</code></pre>
<p>
As you can observe, <code>fct_other()</code> requires less typing and is easier to specify.
</p>
</section>
<section id="anonymize-the-data-set-before-sharing-it-with-your-colleagues" class="level3">
<h3 class="anchored" data-anchor-id="anonymize-the-data-set-before-sharing-it-with-your-colleagues">
Anonymize the data set before sharing it with your colleagues
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_anonymize.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
Anonymizing data is extremely important when you are sharing sensitive data with others. Here, we want to anonymize the channels which drive traffic to the website so that we can share it with others without divulging the names of the channels. <code>fct_anon()</code> allows us to anonymize the levels in the data. Using the <code>prefix</code> argument, we can specify the prefix to be used while anonymizing the data.
</p>
<pre class="r"><code>fct_count(fct_anon(channel, prefix = "ch_"))</code></pre>
<pre><code>## # A tibble: 8 x 2
##   f          n
##   &lt;fct&gt;  &lt;int&gt;
## 1 ch_1    6073
## 2 ch_2    8031
## 3 ch_3    4395
## 4 ch_4   39853
## 5 ch_5  139668
## 6 ch_6    3375
## 7 ch_7   35615
## 8 ch_8    7388</code></pre>
</section>
<section id="key-functions" class="level3">
<h3 class="anchored" data-anchor-id="key-functions">
Key Functions
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Function
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>fct_collapse()</code>
</td>
<td style="text-align:left;">
Collapse factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_recode()</code>
</td>
<td style="text-align:left;">
Recode factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_lump_min()</code>
</td>
<td style="text-align:left;">
Lump factor levels with count lesser than specified value
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_lump_n()</code>
</td>
<td style="text-align:left;">
Lump all levels except the top n levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_lump_prop()</code>
</td>
<td style="text-align:left;">
Lump factor levels with count lesser than specified proportion
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_lump_lowfreq()</code>
</td>
<td style="text-align:left;">
Lump together least frequent levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_other()</code>
</td>
<td style="text-align:left;">
Replace levels with Other level
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_anon()</code>
</td>
<td style="text-align:left;">
Anonymize factor levels
</td>
</tr>
</tbody>
</table>
</section>
</section>
<section id="addremove" class="level2">
<h2 class="anchored" data-anchor-id="addremove">
Add / Remove Levels
</h2>
<p>
In this small section, we will learn to:
</p>
<ul>
<li>
add new levels
</li>
<li>
drop levels
</li>
<li>
make missing values explicit
</li>
</ul>
<section id="add-a-new-level-blog" class="level3">
<h3 class="anchored" data-anchor-id="add-a-new-level-blog">
Add a new level, Blog
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_expand.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
<code>fct_expand()</code> allows us to add new levels to the data. The label of the new level must be specified after the variable name and must be enclosed in quotes. If the level already exists, it will be ignored. Let us add a new level, <code>Blog</code>.
</p>
<pre class="r"><code>levels(fct_expand(channel, "Blog"))</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"        
## [9] "Blog"</code></pre>
</section>
<section id="drop-existing-level" class="level3">
<h3 class="anchored" data-anchor-id="drop-existing-level">
Drop existing level
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_drop.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
On the other hand, <code>fct_drop()</code> will drop levels which have no values i.e.&nbsp;unused levels. If you want to drop only specific levels, use the <code>only</code> argument and specify the name of the level in quotes. Let us drop the new level we added in the previous example.
</p>
<pre class="r"><code>levels(fct_drop(fct_expand(channel, "Blog")))</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
</section>
<section id="make-missing-values-explicit" class="level3">
<h3 class="anchored" data-anchor-id="make-missing-values-explicit">
Make missing values explicit
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_explicit_na.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In our data set, the gender column has many missing values, and in R, missing values are represented by <code>NA</code>. Suppose you are sharing the data or analysis with someone who is not an R user, and does not know what <code>NA</code> represents. In such a scenario, we can use the <code>fct_explicit_na()</code> function to make the missing values in the gender column explicit i.e.&nbsp;it will appear as <code>(Missing)</code> instead of <code>NA</code>. This will help non R users to understand that there are missing values in the data.
</p>
<pre class="r"><code>fct_count(fct_explicit_na(data$gender))</code></pre>
<pre><code>## # A tibble: 3 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 female     40565
## 2 male       61617
## 3 (Missing) 142216</code></pre>
</section>
<section id="key-functions-1" class="level3">
<h3 class="anchored" data-anchor-id="key-functions-1">
Key Functions
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Function
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>fct_expand()</code>
</td>
<td style="text-align:left;">
Add additional levels to a factor
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_drop()</code>
</td>
<td style="text-align:left;">
Drop unused factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_explicit_na()</code>
</td>
<td style="text-align:left;">
Make missing values explicit
</td>
</tr>
</tbody>
</table>
</section>
</section>
<section id="changeorder" class="level2">
<h2 class="anchored" data-anchor-id="changeorder">
Change Order of Levels
</h2>
<p>
In this last section, we will learn how to change the order of the levels. We will look at the following scenarios from our case study:
</p>
<p>
We want to make
</p>
<ul>
<li>
Organic Search the first level
</li>
<li>
Referral the third level
</li>
<li>
Display the last level
</li>
</ul>
<section id="make-organic-search-the-first-level" class="level3">
<h3 class="anchored" data-anchor-id="make-organic-search-the-first-level">
Make Organic Search the first level
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_1.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In this scenario, we want the levels to appear in a certain order. In the first case, we want <strong>Organic Search</strong> to be the first level. <code>fct_relevel()</code> allows us to manually reorder the levels. To move a level to the beginning, specify the label (it must be enclosed in quotes).
</p>
<pre class="r"><code>levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_relevel(channel, "Organic Search"))</code></pre>
<pre><code>## [1] "Organic Search" "(Other)"        "Affiliates"     "Direct"        
## [5] "Display"        "Paid Search"    "Referral"       "Social"</code></pre>
</section>
<section id="make-referral-the-third-level" class="level3">
<h3 class="anchored" data-anchor-id="make-referral-the-third-level">
Make Referral the third level
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_2.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
The <code>after</code> argument is useful when we want to move the level to the end or anywhere between the beginning and end. In the second case, we want <strong>Referral</strong> to be the third level. After specifying the label, use the <code>after</code> argument and specify the level after which <strong>Referral</strong> should appear. Since we want to move it to the third position, we will set the value of <code>after</code> to <code>2</code> i.e.&nbsp;<strong>Referral</strong> should come after the second position.
</p>
<pre class="r"><code>levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_relevel(channel, "Referral", after = 2))</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Referral"       "Direct"        
## [5] "Display"        "Organic Search" "Paid Search"    "Social"</code></pre>
</section>
<section id="make-display-the-last-level" class="level3">
<h3 class="anchored" data-anchor-id="make-display-the-last-level">
Make Display the last level
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_3.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In this last case, we want to move <strong>Display</strong> to the end. If you know the number of levels, you can specify a value here. In our data, there are eight channels i.e.&nbsp;eight levels, so we can set the value of <code>after</code> to <code>7</code>. What happens when we do not know the number of levels or if they tend to vary? In such cases, to move a level to the end, set the value of <code>after</code> to <code>Inf</code>.
</p>
<pre class="r"><code>levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_relevel(channel, "Display", after = Inf))</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Organic Search"
## [5] "Paid Search"    "Referral"       "Social"         "Display"</code></pre>
<p>
Let us now look at a scenario where we want to order the levels by
</p>
<ul>
<li>
frequency (largest to smallest)
</li>
<li>
order of appearance (in data)
</li>
</ul>
</section>
<section id="order-levels-by-frequency" class="level3">
<h3 class="anchored" data-anchor-id="order-levels-by-frequency">
Order levels by frequency
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_infreq.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In the first case, the levels with the most frequency should appear at the top. <code>fct_infreq()</code> will order the levels by their frequency.
</p>
<pre class="r"><code># reorder levels
levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_infreq(channel))</code></pre>
<pre><code>## [1] "Organic Search" "Direct"         "Referral"       "Social"        
## [5] "Affiliates"     "(Other)"        "Paid Search"    "Display"</code></pre>
</section>
<section id="order-levels-by-appearance" class="level3">
<h3 class="anchored" data-anchor-id="order-levels-by-appearance">
Order levels by appearance
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_inorder.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
In the second case, the order of the levels should be the same as the order of their appearance in the data. <code>fct_inorder()</code> will order the levels according to the order in which they appear in the data.
</p>
<pre class="r"><code># reorder levels
levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_inorder(channel))</code></pre>
<pre><code>## [1] "Organic Search" "Direct"         "Referral"       "Affiliates"    
## [5] "(Other)"        "Social"         "Display"        "Paid Search"</code></pre>
</section>
<section id="reverse-the-order-of-the-levels" class="level3">
<h3 class="anchored" data-anchor-id="reverse-the-order-of-the-levels">
Reverse the order of the levels
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_rev.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
The order of the levels can be reversed using <code>fct_rev()</code>.
</p>
<pre class="r"><code># reorder levels
levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_rev(channel))</code></pre>
<pre><code>## [1] "Social"         "Referral"       "Paid Search"    "Organic Search"
## [5] "Display"        "Direct"         "Affiliates"     "(Other)"</code></pre>
</section>
<section id="randomly-shuffle-the-order-of-the-levels" class="level3">
<h3 class="anchored" data-anchor-id="randomly-shuffle-the-order-of-the-levels">
Randomly shuffle the order of the levels
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_shuffle.png" width="90%" style="display: block; margin: auto;">
</p>
<p>
The order of the levels can be randomly shuffled using <code>fct_shuffle()</code>.
</p>
<pre class="r"><code># reorder levels
levels(channel)</code></pre>
<pre><code>## [1] "(Other)"        "Affiliates"     "Direct"         "Display"       
## [5] "Organic Search" "Paid Search"    "Referral"       "Social"</code></pre>
<pre class="r"><code>levels(fct_shuffle(channel))</code></pre>
<pre><code>## [1] "Organic Search" "Referral"       "Display"        "(Other)"       
## [5] "Direct"         "Paid Search"    "Affiliates"     "Social"</code></pre>
</section>
<section id="key-functions-2" class="level3">
<h3 class="anchored" data-anchor-id="key-functions-2">
Key Functions
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Function
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>fct_relevel()</code>
</td>
<td style="text-align:left;">
Reorder factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_shift()</code>
</td>
<td style="text-align:left;">
Shift factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_infreq()</code>
</td>
<td style="text-align:left;">
Reorder factor levels by frequency
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_rev()</code>
</td>
<td style="text-align:left;">
Reverse order of factor levels
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_inorder()</code>
</td>
<td style="text-align:left;">
Reorder factor levels by first appearance
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>fct_shuffle()</code>
</td>
<td style="text-align:left;">
Randomly shuffle factor levels
</td>
</tr>
</tbody>
</table>
</section>
</section>
<section id="practice" class="level2">
<h2 class="anchored" data-anchor-id="practice">
Your Turn…
</h2>
<ol style="list-style-type: decimal">
<li>
<p>
Display the count/frequency of the following variables in the descending order
</p>
<ul>
<li>
<code>device</code>
</li>
<li>
<code>landing_page</code>
</li>
<li>
<code>exit_page</code>
</li>
</ul>
</li>
<li>
<p>
Check if <code>laptop</code> is a level in the <code>device</code> column.
</p>
</li>
<li>
<p>
Combine the following levels in <code>landing_page</code> into <code>Account</code>
</p>
<ul>
<li>
<code>My Account</code>
</li>
<li>
<code>Register</code>
</li>
<li>
<code>Sign In</code>
</li>
<li>
<code>Your Info</code>
</li>
</ul>
</li>
<li>
<p>
Combine levels in <code>landing_page</code> that drive less than 1000 visits.
</p>
</li>
<li>
<p>
Get top 10 landing and exit pages.
</p>
</li>
<li>
<p>
Get landing pages that drive at least 5% of the total traffic to the website.
</p>
</li>
<li>
<p>
Retain only the following levels in the <code>browser</code> column:
</p>
<ul>
<li>
<code>Chrome</code>
</li>
<li>
<code>Firefox</code>
</li>
<li>
<code>Safari</code>
</li>
<li>
<code>Edge</code>
</li>
</ul>
</li>
<li>
<p>
Anonymize landing and exit page levels.
</p>
</li>
<li>
<p>
Make <code>Home</code> first level in the <code>landing_page</code> column.
</p>
</li>
<li>
<p>
Make <code>Apparel</code> second level in the <code>landing_page</code> column.
</p>
</li>
<li>
<p>
Make <code>Specials</code> last level in the <code>landing_page</code> column.
</p>
</li>
<li>
<p>
Order the levels in the browser by frequency:
</p>
</li>
<li>
<p>
Order the levels in landing page by appearance:
</p>
</li>
<li>
<p>
Shuffle the levels in os
</p>
</li>
<li>
<p>
Reverse the levels in browser
</p>
</li>
</ol>
<p>
*As the reader of this blog, you are our most important critic and commentator. We value your opinion and want to know what we are doing right, what we could do better, what areas you would like to see us publish in, and any other words of wisdom you are willing to pass our way.
</p>
<p>
We welcome your comments. You can email to let us know what you did or did not like about our blog as well as what we can do to make our post better.*
</p>
<p>
<strong>Email: <a href="mailto:support@rsquaredacademy.com" class="email">support@rsquaredacademy.com</a></strong>
</p>
</section>



 ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-3/</guid>
  <pubDate>Thu, 13 Jan 2022 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/forcats-part-3.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>Handling Categorical Data in R - Part 2</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-2/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2022-01-12-handling-categorical-data-in-r-part-2.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<script src="../../rmarkdown-libs/header-attrs/header-attrs.js"></script>
<p>
<img src="https://blog.rsquaredacademy.com/img/forcats-part-2.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
This is part 2 of a series on “Handling Categorical Data in R” where we are learning to <strong>read</strong>, <strong>store</strong>, <strong>summarize</strong>, <strong>reshape</strong> &amp; <strong>visualize</strong> categorical data.
</p>
<p>
Below are the links to the other articles of this series:
</p>
<ul>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-1/">Part 1 - Introduction to Factor</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-3/">Part 3 - Reshape Categorical Data</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-4/">Part 4 - Visualize Categorical Data</a>
</li>
</ul>
<p>
In this article, we will learn to summarize categorical data. In the process, we will do a deep dive on working with tables in R and explore a diverse set of packages.
</p>
<section id="table-of-contents" class="level2">
<h2 class="anchored" data-anchor-id="table-of-contents">
Table of Contents
</h2>
<ul>
<li>
Resources
</li>
<li>
Introduction
</li>
<li>
Tables in R
</li>
<li>
Practice Questions
</li>
</ul>
</section>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
You can download all the data sets, R scripts, practice questions and their solutions from our <a href="https://github.com/rsquaredacademy-education/online-courses/">GitHub</a> repository.
</p>
</section>
<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">
Introduction
</h2>
<p>
Categorical data cannot be summarized in the same way as numeric data. It does not make sense to look at range, standard deviation etc. since data consists of a few distinct values only. So how do we summarize such data? We can look at
</p>
<ul>
<li>
count/frequency
</li>
<li>
proportion
</li>
<li>
cumulative frequency
</li>
<li>
cross table
</li>
<li>
contingency table etc.
</li>
</ul>
<p>
In this section, we will explore the above ways of summarizing categorical data. We will also spend some time learning about tables as you will be using them extensively while working with categorical data. R has many packages for tabulating data and we list and explore all of them in the R scripts shared in the GitHub repository.
</p>
<section id="number-of-categories" class="level3">
<h3 class="anchored" data-anchor-id="number-of-categories">
Number of Categories
</h3>
<p>
From our case study, we want to know the number of devices used to browse the website, the name of the devices and the proportion of traffic they drive to our website. Let us read the case study data set before we analyze the website traffic.
</p>
<pre class="r"><code># read data
data &lt;- readRDS('analytics.rds')</code></pre>
<p>
Let us begin with the number of devices. To view the number of groups/categories in a categorical variable, use <code>nlevels()</code>.
</p>
<pre class="r"><code>nlevels(data$device)</code></pre>
<pre><code>## [1] 3</code></pre>
<p>
There are 3 categories of devices used by the visitors to browse the website. This can also be used for data sanitization i.e.&nbsp;as an analyst you know that there are only 3 valid categories of device into which any visitor can be classified into. If you see more than 3 categories, you might want to check if there are any issues in data collection or processing. Now that we know there are 3 categories of devices, let us check if they are valid. The <code>levels()</code> function will return the labels of the groups.
</p>
</section>
<section id="category-names" class="level3">
<h3 class="anchored" data-anchor-id="category-names">
Category Names
</h3>
<p>
Knowing the number of levels is useful but not sufficient. <code>levels()</code> is one of the most useful functions when it comes to dealing with categorical data.
</p>
<pre class="r"><code>levels(data$device)</code></pre>
<pre><code>## [1] "Desktop" "Mobile"  "Tablet"</code></pre>
<p>
Other functions that you can use include <code>unique()</code> and <code>fct_unique()</code>. Both these functions will return the unique names/labels along with the levels while <code>levels()</code> returns the labels of the levels.
</p>
<pre class="r"><code>unique(data$device)</code></pre>
<pre><code>## [1] Desktop Mobile  Tablet 
## Levels: Desktop Mobile Tablet</code></pre>
<pre class="r"><code>fct_unique(data$device)</code></pre>
<pre><code>## [1] Desktop Mobile  Tablet 
## Levels: Desktop Mobile Tablet</code></pre>
</section>
<section id="names-counts" class="level3">
<h3 class="anchored" data-anchor-id="names-counts">
Names &amp; Counts
</h3>
<p>
So we have checked the number of devices and their names. Let us now examine their distribution i.e.&nbsp;count/frequency. <code>table()</code> and <code>summary()</code> will display the levels and their counts while <code>fct_count()</code> will return a tibble with 2 columns (level &amp; count). It is extremely useful for further data processing or visualization (using ggplot2).
</p>
<pre class="r"><code>table(data$device)</code></pre>
<pre><code>## 
## Desktop  Mobile  Tablet 
##  177282   63482    3634</code></pre>
<pre class="r"><code>fct_count(data$device)</code></pre>
<pre><code>## # A tibble: 3 x 2
##   f            n
##   &lt;fct&gt;    &lt;int&gt;
## 1 Desktop 177282
## 2 Mobile   63482
## 3 Tablet    3634</code></pre>
<pre class="r"><code>summary(data$device)</code></pre>
<pre><code>## Desktop  Mobile  Tablet 
##  177282   63482    3634</code></pre>
</section>
</section>
<section id="tables" class="level2">
<h2 class="anchored" data-anchor-id="tables">
Tables
</h2>
<p>
In the previous section, we used the <code>table()</code> function to tabulate categorical data. We will recreate the tabulation for device and store it in a new variable tab.
</p>
<pre class="r"><code>tab &lt;- table(data$device)
tab</code></pre>
<pre><code>## 
## Desktop  Mobile  Tablet 
##  177282   63482    3634</code></pre>
<p>
What does this function return? It is not a <code>vector</code>, <code>list</code>, <code>data.frame</code> or <code>matrix</code>. Let us use the <code>class()</code> function to check the class of the object returned by <code>table()</code>. It returns an object of the class <strong>table</strong>. This is a new type of object. Let us spend some time understanding tables as they are useful for organizing and summarizing categorical data. <strong>table</strong> is also the most used object when it comes to dealing with categorical data.
</p>
<p>
The <code>table()</code> function returns the counts of the categories but let us say we want to view the proportion or percentage instead of counts i.e.&nbsp;the proportion or percentage of traffic driven to our website by the different devices. The <code>proportions()</code> or <code>prop.table()</code> function comes in handy in such cases. It takes a table object as input (tab in our case).
</p>
<pre class="r"><code>prop.table(tab)</code></pre>
<pre><code>## 
##    Desktop     Mobile     Tablet 
## 0.72538237 0.25974844 0.01486919</code></pre>
<pre class="r"><code>proportions(tab)</code></pre>
<pre><code>## 
##    Desktop     Mobile     Tablet 
## 0.72538237 0.25974844 0.01486919</code></pre>
<p>
To get the percentages, multiply the output by 100. Use the <code>round()</code> function to round the decimal places according to your requirements.
</p>
<pre class="r"><code>proportions(tab) * 100</code></pre>
<pre><code>## 
##   Desktop    Mobile    Tablet 
## 72.538237 25.974844  1.486919</code></pre>
<pre class="r"><code>round(proportions(tab) * 100, 2)</code></pre>
<pre><code>## 
## Desktop  Mobile  Tablet 
##   72.54   25.97    1.49</code></pre>
<p>
So far, we have used <code>table()</code> to tabulate a single categorical variable. It can be used for a lot more than just tabulating data. We can examine the relationship between two categorical variables as well as create multidimensional tables. Let us look at the relationship between <strong>gender</strong> and <strong>device</strong> in our case study. Does gender affect the type of device used? To answer this, we will create a two way or cross table. In the <code>table()</code> function, we can specify multiple variables by separating them with a comma.
</p>
<pre class="r"><code>tab2 &lt;- table(data$gender, data$device)
tab2</code></pre>
<pre><code>##         
##          Desktop Mobile Tablet
##   female   32803   7268    494
##   male     46418  14503    696
##   &lt;NA&gt;     98061  41711   2444</code></pre>
<p>
Keep in mind that the order of the variables matter. Rows represent the first variable while column represents the second.
</p>
<pre class="r"><code>table(data$device, data$gender)</code></pre>
<pre><code>##          
##           female  male  &lt;NA&gt;
##   Desktop  32803 46418 98061
##   Mobile    7268 14503 41711
##   Tablet     494   696  2444</code></pre>
<p>
The <code>proportions()</code> function works with two way tables as well.
</p>
<pre class="r"><code>proportions(tab2)</code></pre>
<pre><code>##         
##              Desktop      Mobile      Tablet
##   female 0.134219593 0.029738378 0.002021293
##   male   0.189927904 0.059341729 0.002847814
##   &lt;NA&gt;   0.401234871 0.170668336 0.010000082</code></pre>
<pre class="r"><code>proportions(tab2) * 100</code></pre>
<pre><code>##         
##             Desktop     Mobile     Tablet
##   female 13.4219593  2.9738378  0.2021293
##   male   18.9927904  5.9341729  0.2847814
##   &lt;NA&gt;   40.1234871 17.0668336  1.0000082</code></pre>
<p>
We would like to introduce another function at this point of time, <code>margin.table()</code>. What does this function do? It computes the marginal frequencies i.e.&nbsp;the sum of the rows or columns. It takes a <strong>table</strong> object as input. The <code>margin</code> argument allows us to specify whether we want the sum of rows or columns. <code>1</code> indicates rows and <code>2</code> indicates columns.
</p>
<pre class="r"><code>margin.table(tab2, 1) # sum of rows</code></pre>
<pre><code>## 
## female   male   &lt;NA&gt; 
##  40565  61617 142216</code></pre>
<pre class="r"><code>margin.table(tab2, 2) # sum of columns</code></pre>
<pre><code>## 
## Desktop  Mobile  Tablet 
##  177282   63482    3634</code></pre>
<p>
If the margin argument is NULL (which it is by default), the function returns the sum of all cells of the table.
</p>
<pre class="r"><code>margin.table(tab2)</code></pre>
<pre><code>## [1] 244398</code></pre>
<p>
<code>table()</code> does not display row or column labels. It does display the group labels though. Let us revisit the output from <code>tab2</code>. You can observe that while it includes the group labels, the row and column labels are missing. The output from the <code>dimnames()</code> function shows the group labels of the variables but the row &amp; column labels are absent.
</p>
<pre class="r"><code>dimnames(tab2)</code></pre>
<pre><code>## [[1]]
## [1] "female" "male"   NA      
## 
## [[2]]
## [1] "Desktop" "Mobile"  "Tablet"</code></pre>
<pre class="r"><code>names(tab2)</code></pre>
<pre><code>## NULL</code></pre>
<pre class="r"><code>names(dimnames(tab2)) </code></pre>
<pre><code>## [1] "" ""</code></pre>
<p>
The output from <code>names(dimnames(tab2))</code> is also empty. Let us add the variable names as the row &amp; column labels to <code>tab2</code>.
</p>
<pre class="r"><code>names(dimnames(tab2)) &lt;- c("Gender", "Device")
tab2</code></pre>
<pre><code>##         Device
## Gender   Desktop Mobile Tablet
##   female   32803   7268    494
##   male     46418  14503    696
##   &lt;NA&gt;     98061  41711   2444</code></pre>
<p>
Now look at the output from <code>tab2</code> and you can observe the difference. The same is also visible when we run <code>dimnames(tab2)</code>.
</p>
<pre class="r"><code>dimnames(tab2)</code></pre>
<pre><code>## $Gender
## [1] "female" "male"   NA      
## 
## $Device
## [1] "Desktop" "Mobile"  "Tablet"</code></pre>
<p>
To add margin totals to the table, use <code>addmargins()</code>. Like <code>proportions()</code> and <code>margin.table()</code>, it also takes a <code>table</code> object as the input.
</p>
<pre class="r"><code>addmargins(tab2)</code></pre>
<pre><code>##         Device
## Gender   Desktop Mobile Tablet    Sum
##   female   32803   7268    494  40565
##   male     46418  14503    696  61617
##   &lt;NA&gt;     98061  41711   2444 142216
##   Sum     177282  63482   3634 244398</code></pre>
<p>
<code>rowSums()</code> returns the row total while <code>colSums()</code> returns the column total. They are similar to <code>margin.table()</code>.
</p>
<pre class="r"><code>rowSums(tab2)</code></pre>
<pre><code>## female   male   &lt;NA&gt; 
##  40565  61617 142216</code></pre>
<pre class="r"><code>colSums(tab2)</code></pre>
<pre><code>## Desktop  Mobile  Tablet 
##  177282   63482    3634</code></pre>
<p>
<code>xtabs()</code> is another way of creating multidimensional tables in R. In comparison to <code>table()</code>, it
</p>
<ul>
<li>
uses <code>formula</code> notation for input
</li>
<li>
the data argument ensures variable names are referenced instead of using $ i.e.&nbsp;<code>data<img src="https://latex.codecogs.com/png.latex?variable%3C/code%3E%3C/li%3E%0A%3Cli%3Edisplays%20row%20&amp;amp;%20column%20labels%20by%20default%3C/li%3E%0A%3C/ul%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Etabx%20&amp;lt;-%20xtabs(~gender+device,%20data%20=%20data)%0Atabx%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%2032803%20%20%207268%20%20%20%20494%0A##%20%20%20male%20%20%20%20%2046418%20%2014503%20%20%20%20696%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%2098061%20%2041711%20%20%202444%3C/code%3E%3C/pre%3E%0A%3Cp%3EThe%20following%20functions%20work%20with%20%3Ccode%3Extabs()%3C/code%3E%20as%20well%3C/p%3E%0A%3Cul%3E%0A%3Cli%3E%3Ccode%3Eproportions()%3C/code%3E%3C/li%3E%0A%3Cli%3E%3Ccode%3Emargin.table()%3C/code%3E%3C/li%3E%0A%3Cli%3E%3Ccode%3Eaddmargins()%3C/code%3E%3C/li%3E%0A%3C/ul%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Eproportions(tabx)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20%20%20%20%20Desktop%20%20%20%20%20%20Mobile%20%20%20%20%20%20Tablet%0A##%20%20%20female%200.134219593%200.029738378%200.002021293%0A##%20%20%20male%20%20%200.189927904%200.059341729%200.002847814%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%200.401234871%200.170668336%200.010000082%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Emargin.table(tabx,%201)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20gender%0A##%20female%20%20%20male%20%20%20&amp;lt;NA&amp;gt;%0A##%20%2040565%20%2061617%20142216%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Emargin.table(tabx,%202)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20device%0A##%20Desktop%20%20Mobile%20%20Tablet%0A##%20%20177282%20%20%2063482%20%20%20%203634%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Eaddmargins(tabx)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%20%20%20%20Sum%0A##%20%20%20female%20%20%2032803%20%20%207268%20%20%20%20494%20%2040565%0A##%20%20%20male%20%20%20%20%2046418%20%2014503%20%20%20%20696%20%2061617%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%2098061%20%2041711%20%20%202444%20142216%0A##%20%20%20Sum%20%20%20%20%20177282%20%2063482%20%20%203634%20244398%3C/code%3E%3C/pre%3E%0A%3Cp%3ESo%20far,%20we%20have%20been%20working%20with%20one%20or%20two%20dimensional%20tables.%20Both%20the%20%3Ccode%3Etable()%3C/code%3E%20and%20%3Ccode%3Extabs()%3C/code%3E%20functions%20are%20capable%20of%20creating%20multidimensional%20tables.%20Keep%20in%20mind%20that%20multidimensional%20tables%20are%20complex%20and%20it%20becomes%20increasingly%20difficult%20to%20understand%20or%20interpret%20them.%3C/p%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Etab3%20&amp;lt;-%20xtabs(~gender+device+channel,%20data%20=%20data)%0Atab3%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20,%20,%20channel%20=%20(Other)%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%20%20786%20%20%20%20258%20%20%20%20%20%200%0A##%20%20%20male%20%20%20%20%20%201063%20%20%20%20507%20%20%20%20%2019%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%20%202173%20%20%201186%20%20%20%20%2081%0A##%0A##%20,%20,%20channel%20=%20Affiliates%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%201314%20%20%20%20%2060%20%20%20%20%20%200%0A##%20%20%20male%20%20%20%20%20%201714%20%20%20%20169%20%20%20%20%20%200%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%20%203518%20%20%20%20548%20%20%20%20%2065%0A##%0A##%20,%20,%20channel%20=%20Direct%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%204785%20%20%20%20977%20%20%20%20%2059%0A##%20%20%20male%20%20%20%20%20%207010%20%20%202381%20%20%20%20%2095%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%2015824%20%20%208292%20%20%20%20430%0A##%0A##%20,%20,%20channel%20=%20Display%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%20%20123%20%20%20%20753%20%20%20%20104%0A##%20%20%20male%20%20%20%20%20%20%20210%20%20%20%20491%20%20%20%20%2073%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%20%20%20554%20%20%20%20911%20%20%20%20156%0A##%0A##%20,%20,%20channel%20=%20Organic%20Search%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%2017109%20%20%204480%20%20%20%20282%0A##%20%20%20male%20%20%20%20%2025016%20%20%209563%20%20%20%20448%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%2054071%20%2027223%20%20%201476%0A##%0A##%20,%20,%20channel%20=%20Paid%20Search%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%20%20645%20%20%20%20230%20%20%20%20%2022%0A##%20%20%20male%20%20%20%20%20%20%20887%20%20%20%20478%20%20%20%20%2026%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%20%201274%20%20%20%20782%20%20%20%20%2051%0A##%0A##%20,%20,%20channel%20=%20Referral%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%207387%20%20%20%20%2074%20%20%20%20%20%200%0A##%20%20%20male%20%20%20%20%20%209251%20%20%20%20185%20%20%20%20%20%200%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%2018052%20%20%20%20615%20%20%20%20%2051%0A##%0A##%20,%20,%20channel%20=%20Social%0A##%0A##%20%20%20%20%20%20%20%20%20device%0A##%20gender%20%20%20Desktop%20Mobile%20Tablet%0A##%20%20%20female%20%20%20%20%20654%20%20%20%20436%20%20%20%20%2027%0A##%20%20%20male%20%20%20%20%20%201267%20%20%20%20729%20%20%20%20%2035%0A##%20%20%20&amp;lt;NA&amp;gt;%20%20%20%20%20%202595%20%20%202154%20%20%20%20134%3C/code%3E%3C/pre%3E%0A%3Cp%3E%3Cstrong%3Eftable%3C/strong%3E%20stands%20for%20flat%20tables%20and%20is%20useful%20for%20printing%20attractive%20tables.%20It%20makes%20it%20easy%20to%20read%20and%20interpret%20multidimensional%20tables.%20In%20the%20next%20example,%20we%20will%20use%20%3Ccode%3Eftable()%3C/code%3E%20to%20print%20the%20tables%20we%20have%20created%20in%20the%20previous%20examples%20and%20compare%20the%20outputs.%3C/p%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Eftable(tabx)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20device%20Desktop%20Mobile%20Tablet%0A##%20gender%0A##%20female%20%20%20%20%20%20%20%20%20%2032803%20%20%207268%20%20%20%20494%0A##%20male%20%20%20%20%20%20%20%20%20%20%20%2046418%20%2014503%20%20%20%20696%0A##%20NA%20%20%20%20%20%20%20%20%20%20%20%20%20%2098061%20%2041711%20%20%202444%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Eftable(tab2)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20Device%20Desktop%20Mobile%20Tablet%0A##%20Gender%0A##%20female%20%20%20%20%20%20%20%20%20%2032803%20%20%207268%20%20%20%20494%0A##%20male%20%20%20%20%20%20%20%20%20%20%20%2046418%20%2014503%20%20%20%20696%0A##%20NA%20%20%20%20%20%20%20%20%20%20%20%20%20%2098061%20%2041711%20%20%202444%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Eftable(tab3)%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20channel%20(Other)%20Affiliates%20Direct%20Display%20Organic%20Search%20Paid%20Search%20Referral%20Social%0A##%20gender%20device%0A##%20female%20Desktop%20%20%20%20%20%20%20%20%20%20%20%20%20786%20%20%20%20%20%20%201314%20%20%204785%20%20%20%20%20123%20%20%20%20%20%20%20%20%20%2017109%20%20%20%20%20%20%20%20%20645%20%20%20%20%207387%20%20%20%20654%0A##%20%20%20%20%20%20%20%20Mobile%20%20%20%20%20%20%20%20%20%20%20%20%20%20258%20%20%20%20%20%20%20%20%2060%20%20%20%20977%20%20%20%20%20753%20%20%20%20%20%20%20%20%20%20%204480%20%20%20%20%20%20%20%20%20230%20%20%20%20%20%20%2074%20%20%20%20436%0A##%20%20%20%20%20%20%20%20Tablet%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%200%20%20%20%20%20%20%20%20%20%200%20%20%20%20%2059%20%20%20%20%20104%20%20%20%20%20%20%20%20%20%20%20%20282%20%20%20%20%20%20%20%20%20%2022%20%20%20%20%20%20%20%200%20%20%20%20%2027%0A##%20male%20%20%20Desktop%20%20%20%20%20%20%20%20%20%20%20%201063%20%20%20%20%20%20%201714%20%20%207010%20%20%20%20%20210%20%20%20%20%20%20%20%20%20%2025016%20%20%20%20%20%20%20%20%20887%20%20%20%20%209251%20%20%201267%0A##%20%20%20%20%20%20%20%20Mobile%20%20%20%20%20%20%20%20%20%20%20%20%20%20507%20%20%20%20%20%20%20%20169%20%20%202381%20%20%20%20%20491%20%20%20%20%20%20%20%20%20%20%209563%20%20%20%20%20%20%20%20%20478%20%20%20%20%20%20185%20%20%20%20729%0A##%20%20%20%20%20%20%20%20Tablet%20%20%20%20%20%20%20%20%20%20%20%20%20%20%2019%20%20%20%20%20%20%20%20%20%200%20%20%20%20%2095%20%20%20%20%20%2073%20%20%20%20%20%20%20%20%20%20%20%20448%20%20%20%20%20%20%20%20%20%2026%20%20%20%20%20%20%20%200%20%20%20%20%2035%0A##%20NA%20%20%20%20%20Desktop%20%20%20%20%20%20%20%20%20%20%20%202173%20%20%20%20%20%20%203518%20%2015824%20%20%20%20%20554%20%20%20%20%20%20%20%20%20%2054071%20%20%20%20%20%20%20%201274%20%20%20%2018052%20%20%202595%0A##%20%20%20%20%20%20%20%20Mobile%20%20%20%20%20%20%20%20%20%20%20%20%201186%20%20%20%20%20%20%20%20548%20%20%208292%20%20%20%20%20911%20%20%20%20%20%20%20%20%20%2027223%20%20%20%20%20%20%20%20%20782%20%20%20%20%20%20615%20%20%202154%0A##%20%20%20%20%20%20%20%20Tablet%20%20%20%20%20%20%20%20%20%20%20%20%20%20%2081%20%20%20%20%20%20%20%20%2065%20%20%20%20430%20%20%20%20%20156%20%20%20%20%20%20%20%20%20%20%201476%20%20%20%20%20%20%20%20%20%2051%20%20%20%20%20%20%2051%20%20%20%20134%3C/code%3E%3C/pre%3E%0A%3Cp%3EBy%20default,%20missing%20values%20(NAs)%20are%20excluded%20from%20tables.%20Let%20us%20modify%20the%20gender%20data%20from%20our%20case%20study%20a%20bit%20and%20see%20how%20the%20%3Ccode%3Etable()%3C/code%3E%20function%20deals%20with%20missing%20values.%20We%20won%E2%80%99t%20explicitly%20specify%20%3Ccode%3ENA%3C/code%3E%20as%20a%20level%20while%20recreating%20the%20gender%20data.%3C/p%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Egen%20&amp;lt;-%20as.factor(as.character(data">gender)) table(gen)</code>

<pre><code>## gen
## female   male 
##  40565  61617</code></pre>
<p>
As you can see, <code>table()</code> excludes missing values while tabulating the data. In order to ensure that missing values are also counted, we can use the <code>useNA</code> argument. It can take two values:
</p>
<ul>
<li>
ifany
</li>
<li>
always
</li>
</ul>
<p>
In the first case, it will show <code>NA</code> as a level and the count only if there are missing values in the data. In the second case, it will always show <code>NA</code> as a level irrespective of whether there are missing values in the data or not.
</p>
<pre class="r"><code>table(gen, useNA = "ifany")</code></pre>
<pre><code>## gen
## female   male   &lt;NA&gt; 
##  40565  61617 142216</code></pre>
<pre class="r"><code>table(data$device, useNA = "always")</code></pre>
<pre><code>## 
## Desktop  Mobile  Tablet    &lt;NA&gt; 
##  177282   63482    3634       0</code></pre>
<p>
In this final section on tables, we will learn how to select/access the different parts of a <code>table</code>. We will use <code>[</code> operator to select rows and columns of a table (it is similar to selecting data from a <code>data.frame</code>). Below are a few examples:
</p>
<ul>
<li>
select first row
</li>
</ul>
<pre class="r"><code>tab2[1, ]           </code></pre>
<pre><code>## Desktop  Mobile  Tablet 
##   32803    7268     494</code></pre>
<ul>
<li>
select first column
</li>
</ul>
<pre class="r"><code>tab2[, 1]           </code></pre>
<pre><code>## female   male   &lt;NA&gt; 
##  32803  46418  98061</code></pre>
<ul>
<li>
select first two rows
</li>
</ul>
<pre class="r"><code>tab2[1:2, ]         </code></pre>
<pre><code>##         Device
## Gender   Desktop Mobile Tablet
##   female   32803   7268    494
##   male     46418  14503    696</code></pre>
<ul>
<li>
select first two columns
</li>
</ul>
<pre class="r"><code>tab2[, 1:2]        </code></pre>
<pre><code>##         Device
## Gender   Desktop Mobile
##   female   32803   7268
##   male     46418  14503
##   &lt;NA&gt;     98061  41711</code></pre>
<ul>
<li>
select nth row
</li>
</ul>
<pre class="r"><code>tab2[2, ]          </code></pre>
<pre><code>## Desktop  Mobile  Tablet 
##   46418   14503     696</code></pre>
<ul>
<li>
select nth column
</li>
</ul>
<pre class="r"><code>tab2[, 2]          </code></pre>
<pre><code>## female   male   &lt;NA&gt; 
##   7268  14503  41711</code></pre>
<ul>
<li>
select row by group label
</li>
</ul>
<pre class="r"><code>tab2["female", ]   </code></pre>
<pre><code>## Desktop  Mobile  Tablet 
##   32803    7268     494</code></pre>
<ul>
<li>
select column by group label
</li>
</ul>
<pre class="r"><code>tab2[, "Mobile"]   </code></pre>
<pre><code>## female   male   &lt;NA&gt; 
##   7268  14503  41711</code></pre>
<p>
Before we end this section, let us learn how to test if an object is of class table using <code>is.table()</code>.
</p>
<pre class="r"><code>is.table(tab2)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
Next, we will look at different R packages for two way/contingency tables.
</p>
<section id="contingency-table" class="level3">
<h3 class="anchored" data-anchor-id="contingency-table">
Contingency Table
</h3>
<p>
For cross tables with output similar to SAS or SPSS, use any of the below:
</p>
<ul>
<li>
<code>CrossTable()</code> from the <a href="cran.r-project.org/package=gmodels">gmodels</a> package
</li>
<li>
<code>ds_cross_table()</code> from the <a href="https://descriptr.rsquaredacademy.com">descriptr</a> package
</li>
</ul>
<pre class="r"><code>gmodels::CrossTable(data$device, data$gender)</code></pre>
<pre><code>## 
##  
##    Cell Contents
## |-------------------------|
## |                       N |
## | Chi-square contribution |
## |           N / Row Total |
## |           N / Col Total |
## |         N / Table Total |
## |-------------------------|
## 
##  
## Total Observations in Table:  102182 
## 
##  
##              | data$gender 
##  data$device |    female |      male | Row Total | 
## -------------|-----------|-----------|-----------|
##      Desktop |     32803 |     46418 |     79221 | 
##              |    58.228 |    38.334 |           | 
##              |     0.414 |     0.586 |     0.775 | 
##              |     0.809 |     0.753 |           | 
##              |     0.321 |     0.454 |           | 
## -------------|-----------|-----------|-----------|
##       Mobile |      7268 |     14503 |     21771 | 
##              |   218.694 |   143.975 |           | 
##              |     0.334 |     0.666 |     0.213 | 
##              |     0.179 |     0.235 |           | 
##              |     0.071 |     0.142 |           | 
## -------------|-----------|-----------|-----------|
##       Tablet |       494 |       696 |      1190 | 
##              |     0.986 |     0.649 |           | 
##              |     0.415 |     0.585 |     0.012 | 
##              |     0.012 |     0.011 |           | 
##              |     0.005 |     0.007 |           | 
## -------------|-----------|-----------|-----------|
## Column Total |     40565 |     61617 |    102182 | 
##              |     0.397 |     0.603 |           | 
## -------------|-----------|-----------|-----------|
## 
## </code></pre>
<pre class="r"><code>descriptr::ds_cross_table(data, device, gender)</code></pre>
<pre><code>##     Cell Contents
##  |---------------|
##  |     Frequency |
##  |       Percent |
##  |       Row Pct |
##  |       Col Pct |
##  |---------------|
## 
##  Total Observations:  244398 
## 
## ----------------------------------------------------------------------------
## |              |                          gender                           |
## ----------------------------------------------------------------------------
## |       device |       female |         male |           NA |    Row Total |
## ----------------------------------------------------------------------------
## |      Desktop |        32803 |        46418 |        98061 |       177282 |
## |              |        0.134 |         0.19 |        0.401 |              |
## |              |         0.19 |         0.26 |         0.55 |         0.73 |
## |              |         0.81 |         0.75 |         0.69 |              |
## ----------------------------------------------------------------------------
## |       Mobile |         7268 |        14503 |        41711 |        63482 |
## |              |         0.03 |        0.059 |        0.171 |              |
## |              |         0.11 |         0.23 |         0.66 |         0.26 |
## |              |         0.18 |         0.24 |         0.29 |              |
## ----------------------------------------------------------------------------
## |       Tablet |          494 |          696 |         2444 |         3634 |
## |              |        0.002 |        0.003 |         0.01 |              |
## |              |         0.14 |         0.19 |         0.67 |         0.01 |
## |              |         0.01 |         0.01 |         0.02 |              |
## ----------------------------------------------------------------------------
## | Column Total |        40565 |        61617 |       142216 |       244398 |
## |              |        0.166 |        0.252 |        0.582 |              |
## ----------------------------------------------------------------------------</code></pre>
<p>
We list and explore different R packages for summarizing categorical data in our GitHub repository.
</p>
</section>
<section id="key-functions" class="level3">
<h3 class="anchored" data-anchor-id="key-functions">
Key Functions
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_module_4_functions.png" width="70%" style="display: block; margin: auto;">
</p>
</section>
</li></ul></section> ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-2/</guid>
  <pubDate>Wed, 12 Jan 2022 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/forcats-part-2.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>Handling Categorical Data in R - Part 1</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-1/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2022-01-07-handling-categorical-data-in-r-part-1.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<script src="../../rmarkdown-libs/header-attrs/header-attrs.js"></script>
<script src="../../rmarkdown-libs/kePrint/kePrint.js"></script>
<p><link href="../../rmarkdown-libs/lightable/lightable.css" rel="stylesheet"></p>
<p>
<img src="https://blog.rsquaredacademy.com/img/forcats-part-1.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
This is part 1 of a series on “Handling Categorical Data in R.” Almost every data science project involves working with categorical data, and we should know how to <strong>read</strong>, <strong>store</strong>, <strong>summarize</strong>, <strong>reshape</strong> &amp; <strong>visualize</strong> such data. Working with categorical data is different from working with other data types such as numbers or text. In this article, we will understand what categorical data is, how R stores it using <strong>factor</strong>, and explore the rich set of functions (built-in &amp; through packages) provided by R for working with such data. Throughout the series, we will also <strong>work through a case study</strong> to better understand the concepts we learn. Happy learning!
</p>
<p>
Below are the links to the other articles of this series:
</p>
<ul>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-2/">Part 2 - Summarize Categorical Data</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-3/">Part 3 - Reshape Categorical Data</a>
</li>
<li>
<a href="https://blog.rsquaredacademy.com/handling-categorical-data-in-r-part-4/">Part 4 - Visualize Categorical Data</a>
</li>
</ul>
<section id="table-of-contents" class="level2">
<h2 class="anchored" data-anchor-id="table-of-contents">
Table of Contents
</h2>
<ul>
<li>
Resources
</li>
<li>
Introduction
</li>
<li>
Case Study
</li>
<li>
Factors
</li>
</ul>
</section>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
You can download all the data sets, R scripts, practice questions and their solutions from our <a href="https://github.com/rsquaredacademy-education/online-courses/">GitHub</a> repository.
</p>
</section>
<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">
Introduction
</h2>
<p>
Before we begin our deep dive on categorical data, let us get a quick overview of different data types. Feel free to skip this section if you know the difference between nominal and ordinal data.
</p>
<section id="data-types" class="level3">
<h3 class="anchored" data-anchor-id="data-types">
Data Types
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_data_types.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
In the chart above, we can see that data can be primarily classified into <strong>qualitative</strong> or <strong>quantitative</strong>. (<strong>The word categorical is used interchangeably with qualitative throughout the series</strong>). Qualitative data consists of labels or names. Quantitative data, on the other hand, consists of numbers and indicate how much or how many. This brings us to the next level of classification:
</p>
<ul>
<li>
discrete
</li>
<li>
continuous
</li>
</ul>
<p>
In the chart, we can observe that qualitative data is always <strong>discrete</strong> where as quantitative data may be <strong>discrete</strong> or <strong>continuous</strong>. Qualitative data is further classified into
</p>
<ul>
<li>
nominal
</li>
<li>
ordinal
</li>
</ul>
<p>
First, we will understand discrete and continuous data, and then proceed to explore nominal and ordinal data.
</p>
<p>
<strong>Discrete Data</strong>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_discrete.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
Discrete data arises in situations where counting is involved. It can take on only a finite number of values and cannot be divided into smaller parts. For example, let us consider the number of students in a class. We can have 5 0r 10 students but not 5.5 students (we can’t have half a student).
</p>
<p>
<strong>Continuous Data</strong>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_measure.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
Continuous data arises in situations where measuring is involved. It can take any numeric value in a specified range and can be divided into smaller parts and still have meaning. Examples include money, temperature, length, volume etc.
</p>
</section>
<section id="categorical-data" class="level3">
<h3 class="anchored" data-anchor-id="categorical-data">
Categorical Data
</h3>
<p>
Since our interest is in categorical data, we will spend more time understanding the different types of categorical data through various examples. Let us begin by formally defining categorical data:
</p>
<ul>
<li>
it is always discrete
</li>
<li>
it may be divided into groups
</li>
<li>
consists of names or labels
</li>
<li>
takes on limited &amp; fixed number of possible values
</li>
<li>
arises in situation when counting is involved
</li>
<li>
analysis generally involves the use of data tables
</li>
</ul>
</section>
<section id="dichotomous" class="level3">
<h3 class="anchored" data-anchor-id="dichotomous">
Dichotomous
</h3>
<p>
A categorical variable that can take on exactly two values is termed as <strong>binary</strong> or <strong>dichotomous</strong> variable.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_yes_no.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="polychotomous" class="level3">
<h3 class="anchored" data-anchor-id="polychotomous">
Polychotomous
</h3>
<p>
Categorical variables with more than two possible values are called <strong>polychotomous</strong> variables.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_yes_no_maybe.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="ordinal" class="level3">
<h3 class="anchored" data-anchor-id="ordinal">
Ordinal
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_satisfaction_rating.jpg" width="80%" style="display: block; margin: auto;">
</p>
<p>
In ordinal data, the categories can be ordered or ranked. Examples include
</p>
<ul>
<li>
socio-economic status
</li>
<li>
education level
</li>
<li>
income level
</li>
<li>
satisfaction rating
</li>
</ul>
<p>
While we can rank the categories, we cannot assign a value to them. For example, in satisfaction ranking, we cannot say that like is twice as positive as dislike i.e.&nbsp;we are unable to say how much they differ from each other. While the order or rank of data is meaningful, the difference between two pieces of data cannot be measured/determined or are meaningless. Ordinal data provide information about relative comparisons, but not the magnitude of the differences.
</p>
</section>
<section id="nominal" class="level3">
<h3 class="anchored" data-anchor-id="nominal">
Nominal
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_transportation_mode.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
Nominal data do not have an intrinsic order and cannot be ordered or measured. Examples include
</p>
<ul>
<li>
blood group
</li>
<li>
gender
</li>
<li>
religion
</li>
<li>
color
</li>
</ul>
<p>
Categorical data are sometimes coded with numbers, with those numbers replacing names. Although such numbers might appear to be quantitative, they are actually categorical data. When they do take numerical values, those numbers do not have any mathematical meaning. Examples include months expressed in numbers.
</p>
</section>
</section>
<section id="casestudy" class="level2">
<h2 class="anchored" data-anchor-id="casestudy">
Case Study
</h2>
<p>
As is the practice, throughout this series, we will work on a case study related to an e-commerce firm. As most of you would already be aware, a lot of data is captured when you go on the internet by the websites you browse as well as by third party cookies. Data collected is then used to display ads as well as to feed to recommendation algorithms.
</p>
<p>
The data used in the case study represents the basic information that is captured when users visit any website. It closely resembles real world data for an e-commerce store. We will try to generate insights about the visitors to be used by an imaginary marketing team for better targeting and promotion. The case study data set can be imported using the RStudio IDE or R code.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_web_analytics.png" width="80%" style="display: block; margin: auto;">
</p>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<p>
The data set is available in both <strong>CSV</strong> &amp; <strong>RDS</strong> formats.
</p>
<p>
<strong>CSV</strong>
</p>
<p>
If you want to specify the data types while reading the data, use the <a href="https://readr.tidyverse.org/">readr</a> package. We have explored how to import data into R in a previous <a href="https://blog.rsquaredacademy.com/import-data-into-r-part-1/">article</a>. We will read a subset of columns from the data set (it has 20 columns) which will cover both nominal and ordinal data types. To import the data, we will use the <code>read_csv()</code> function. The first input is the name of the data set, <code>analytics.csv</code>. Ensure that the name is enclosed in single/double quotes.
</p>
<pre class="r"><code>read_csv("analytics_raw.csv", 
         col_types = cols_only(device = col_factor(levels = c("Desktop", "Tablet", "Mobile")), 
                               gender = col_factor(levels = c("female", "male", "NA")), 
                               user_rating = col_factor(levels = c("1", "2", "3", "4", "5"),
                                                        ordered = TRUE)))</code></pre>
<pre><code>## New names:
## * `` -&gt; ...1</code></pre>
<pre><code>## # A tibble: 244,398 x 3
##    device  gender user_rating
##    &lt;fct&gt;   &lt;fct&gt;  &lt;ord&gt;      
##  1 Desktop female 4          
##  2 Mobile  NA     5          
##  3 Desktop NA     4          
##  4 Desktop NA     5          
##  5 Desktop NA     4          
##  6 Mobile  NA     4          
##  7 Desktop NA     4          
##  8 Desktop NA     4          
##  9 Desktop female 5          
## 10 Desktop NA     4          
## # ... with 244,388 more rows</code></pre>
<p>
Since we are specifying the column data types while importing the data, we will use the <code>col_types</code> argument to list out the data types. As we are reading in a subset of the columns and not all of them, we will use the <code>cols_only()</code> function indicating that only the columns specified must be read in and not all of them.
</p>
<p>
Categorical data and the levels/groups are specified using the <code>col_factor()</code> function. Use the <code>levels</code> argument to specify the levels/groups and the <code>ordered</code> argument to indicate if the data is ordinal. By default, it is set to <code>FALSE</code>, change this to <code>TRUE</code> if the column is ordinal.
</p>
<p>
<strong>RDS</strong>
</p>
<p>
The <code>.rds</code> file can be read using <code>readRDS()</code>.
</p>
<pre class="r"><code>data &lt;- readRDS('analytics.rds')
head(data)</code></pre>
<pre><code>## # A tibble: 6 x 19
##   device  os        browser user_type channel gender frequency recency page_depth
##   &lt;fct&gt;   &lt;fct&gt;     &lt;fct&gt;   &lt;fct&gt;     &lt;fct&gt;   &lt;fct&gt;      &lt;dbl&gt;   &lt;dbl&gt;      &lt;dbl&gt;
## 1 Desktop Windows   Chrome  New Visi~ Organi~ female         1       0          1
## 2 Mobile  iOS       Safari  Returnin~ Organi~ &lt;NA&gt;           3       1          1
## 3 Desktop Chrome OS Chrome  New Visi~ Direct  &lt;NA&gt;           1       0          5
## 4 Desktop Macintosh Chrome  Returnin~ Organi~ &lt;NA&gt;           2       0          1
## 5 Desktop Macintosh Chrome  Returnin~ Referr~ &lt;NA&gt;           5       8          1
## 6 Mobile  Android   Chrome  New Visi~ Organi~ &lt;NA&gt;           1       0          5
## # ... with 10 more variables: hour_of_day &lt;chr&gt;, age &lt;dbl&gt;, duration &lt;dbl&gt;,
## #   landing_page &lt;fct&gt;, exit_page &lt;fct&gt;, country &lt;fct&gt;, quantity &lt;dbl&gt;,
## #   revenue &lt;dbl&gt;, purchase_flag &lt;lgl&gt;, user_rating &lt;dbl&gt;</code></pre>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Column
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
device
</td>
<td style="text-align:left;">
Device used to browse the website
</td>
</tr>
<tr>
<td style="text-align:left;">
os
</td>
<td style="text-align:left;">
Operating system of the device
</td>
</tr>
<tr>
<td style="text-align:left;">
browser
</td>
<td style="text-align:left;">
Browser used to visit the website
</td>
</tr>
<tr>
<td style="text-align:left;">
user_type
</td>
<td style="text-align:left;">
New or returning visitor
</td>
</tr>
<tr>
<td style="text-align:left;">
channel
</td>
<td style="text-align:left;">
Source of traffic
</td>
</tr>
<tr>
<td style="text-align:left;">
gender
</td>
<td style="text-align:left;">
Gender of the visitor
</td>
</tr>
<tr>
<td style="text-align:left;">
frequency
</td>
<td style="text-align:left;">
Count of visits to the website
</td>
</tr>
<tr>
<td style="text-align:left;">
recency
</td>
<td style="text-align:left;">
Number of days since last visit
</td>
</tr>
<tr>
<td style="text-align:left;">
page_depth
</td>
<td style="text-align:left;">
Number of website pages browsed
</td>
</tr>
<tr>
<td style="text-align:left;">
hour_of_day
</td>
<td style="text-align:left;">
Hour of day
</td>
</tr>
<tr>
<td style="text-align:left;">
age
</td>
<td style="text-align:left;">
Age of the visitor
</td>
</tr>
<tr>
<td style="text-align:left;">
duration
</td>
<td style="text-align:left;">
Time spent on the website (in seconds)
</td>
</tr>
<tr>
<td style="text-align:left;">
landing_page
</td>
<td style="text-align:left;">
Page on which visitor landed
</td>
</tr>
<tr>
<td style="text-align:left;">
exit_page
</td>
<td style="text-align:left;">
Page on which visitor exited
</td>
</tr>
<tr>
<td style="text-align:left;">
country
</td>
<td style="text-align:left;">
Country of origin
</td>
</tr>
<tr>
<td style="text-align:left;">
city
</td>
<td style="text-align:left;">
City of the visitor
</td>
</tr>
<tr>
<td style="text-align:left;">
quantity
</td>
<td style="text-align:left;">
Number of units purchased
</td>
</tr>
<tr>
<td style="text-align:left;">
revenue
</td>
<td style="text-align:left;">
Total revenue
</td>
</tr>
<tr>
<td style="text-align:left;">
purchase_flag
</td>
<td style="text-align:left;">
Whether the visitor checked out?
</td>
</tr>
<tr>
<td style="text-align:left;">
user_rating
</td>
<td style="text-align:left;">
Website UI rating given by visitor
</td>
</tr>
</tbody>
</table>
<p>
Now that we have an overview of the case study, let us move on to the next section where we explore how R stores categorical data using factors.
</p>
</section>
</section>
<section id="factors" class="level2">
<h2 class="anchored" data-anchor-id="factors">
Factors
</h2>
<p>
In this very important section, we will learn how R
</p>
<ul>
<li>
stores categorical data
</li>
<li>
checks if given data is categorical
</li>
<li>
converts other data types to factor
</li>
<li>
handles missing values in categorical data
</li>
<li>
specifies the orders of the categories/levels
</li>
<li>
stores ordinal data
</li>
</ul>
<p>
In R, categorical data is stored as <code>factor</code>. Before we explore the <code>factor</code> family of functions, let us generate the sample data we will use in this module. We will generate the <code>device</code> column from the case study data set using the <code>sample()</code> function. We provide the following inputs to generate the data:
</p>
<ul>
<li>
values from which the data must be generated
</li>
<li>
the size of the sample
</li>
<li>
indicate if the values must be repeated (TRUE/FALSE)
</li>
</ul>
<pre class="r"><code>device &lt;- sample(c("Desktop", "Mobile", "Tablet"), size = 25, replace = TRUE)
device</code></pre>
<pre><code>##  [1] "Mobile"  "Mobile"  "Tablet"  "Mobile"  "Mobile"  "Mobile"  "Tablet" 
##  [8] "Tablet"  "Desktop" "Mobile"  "Mobile"  "Mobile"  "Tablet"  "Mobile" 
## [15] "Desktop" "Tablet"  "Desktop" "Mobile"  "Mobile"  "Tablet"  "Desktop"
## [22] "Mobile"  "Tablet"  "Mobile"  "Mobile"</code></pre>
<section id="membership-testing" class="level3">
<h3 class="anchored" data-anchor-id="membership-testing">
Membership Testing
</h3>
<p>
Great! We have successfully generated the sample data and along the way learnt a new R function for sampling. First, let us check if the sample is a <code>factor</code> using the membership function <code>is.factor()</code>.
</p>
<pre class="r"><code>is.factor(device)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<p>
Membership testing functions always have the prefix <code>is_</code> and return only logical values. If the object is a member of the specified class, they return <code>TRUE</code> else <code>FALSE</code>. Since our sample data is not stored as a factor, R has returned <code>FALSE</code>.
</p>
</section>
<section id="coercion" class="level3">
<h3 class="anchored" data-anchor-id="coercion">
Coercion
</h3>
<p>
Let us try to coerce it into <code>factor</code> using the coercion function <code>as.factor()</code>.
</p>
<pre class="r"><code>as.factor(device)</code></pre>
<pre><code>##  [1] Mobile  Mobile  Tablet  Mobile  Mobile  Mobile  Tablet  Tablet  Desktop
## [10] Mobile  Mobile  Mobile  Tablet  Mobile  Desktop Tablet  Desktop Mobile 
## [19] Mobile  Tablet  Desktop Mobile  Tablet  Mobile  Mobile 
## Levels: Desktop Mobile Tablet</code></pre>
<p>
Do you spot any difference in the output? In the last line, it displays the levels or categories of the variable. Don’t worry if you didn’t spot it. We are just getting started and you will pick it up by the end of this section. Another function that can be used to coerce data into factor is <code>as_factor()</code> from the <a href="https://forcats.tidyverse.org/">forcats</a> package.
</p>
<pre class="r"><code>as_factor(device)</code></pre>
<pre><code>##  [1] Mobile  Mobile  Tablet  Mobile  Mobile  Mobile  Tablet  Tablet  Desktop
## [10] Mobile  Mobile  Mobile  Tablet  Mobile  Desktop Tablet  Desktop Mobile 
## [19] Mobile  Tablet  Desktop Mobile  Tablet  Mobile  Mobile 
## Levels: Mobile Tablet Desktop</code></pre>
<p>
Did you notice any difference between these two functions? Focus on the last line of the output where the levels are displayed. Now observe the order of the levels. <code>as.factor()</code> displays levels in the alphabetical order whereas <code>as_factor()</code> displays them in order of appearance in the data. <strong>Mobile</strong>, followed by <strong>Tablet</strong>, and then <strong>Desktop</strong>. If you look at the data, they appear in the same order.
</p>
</section>
<section id="factor-function" class="level3">
<h3 class="anchored" data-anchor-id="factor-function">
Factor Function
</h3>
<p>
If you want finer control while creating factors, use the <code>factor()</code> function. <code>as.factor()</code> should suffice in most cases but use <code>factor()</code> when you want to:
</p>
<ul>
<li>
specify levels
</li>
<li>
modify labels
</li>
<li>
include <code>NA</code> as a level/category
</li>
<li>
create ordered factors
</li>
<li>
specify order of levels
</li>
</ul>
<p>
The first input is a vector, usually a numeric or character vector with a small number of unique values. In our example, it is a character vector of length 25 (i.e.&nbsp;25 values) but 3 unique values.
</p>
<pre class="r"><code>factor(device)</code></pre>
<pre><code>##  [1] Mobile  Mobile  Tablet  Mobile  Mobile  Mobile  Tablet  Tablet  Desktop
## [10] Mobile  Mobile  Mobile  Tablet  Mobile  Desktop Tablet  Desktop Mobile 
## [19] Mobile  Tablet  Desktop Mobile  Tablet  Mobile  Mobile 
## Levels: Desktop Mobile Tablet</code></pre>
<p>
If you want to specify the levels or categories, use the <code>levels</code> argument.
</p>
<pre class="r"><code>factor(device, levels = c("Desktop", "Mobile", "Tablet"))</code></pre>
<pre><code>##  [1] Mobile  Mobile  Tablet  Mobile  Mobile  Mobile  Tablet  Tablet  Desktop
## [10] Mobile  Mobile  Mobile  Tablet  Mobile  Desktop Tablet  Desktop Mobile 
## [19] Mobile  Tablet  Desktop Mobile  Tablet  Mobile  Mobile 
## Levels: Desktop Mobile Tablet</code></pre>
<p>
Levels not specified will be replaced by <code>NA</code>. Let us specify only <strong>Desktop</strong> and <strong>Mobile</strong> as the levels in the device column and see what happens.
</p>
<pre class="r"><code>factor(device, levels = c("Desktop", "Mobile"))</code></pre>
<pre><code>##  [1] Mobile  Mobile  &lt;NA&gt;    Mobile  Mobile  Mobile  &lt;NA&gt;    &lt;NA&gt;    Desktop
## [10] Mobile  Mobile  Mobile  &lt;NA&gt;    Mobile  Desktop &lt;NA&gt;    Desktop Mobile 
## [19] Mobile  &lt;NA&gt;    Desktop Mobile  &lt;NA&gt;    Mobile  Mobile 
## Levels: Desktop Mobile</code></pre>
<p>
As you can see, <strong>Tablet</strong> has been replaced by <code>NA</code>.
</p>
</section>
<section id="modify-labels" class="level3">
<h3 class="anchored" data-anchor-id="modify-labels">
Modify Labels
</h3>
<p>
You can change the labels of the levels using the <code>labels</code> argument. The labels must be in the same order as the levels. We will modify the labels to <strong>Desk</strong>, <strong>Mob</strong> &amp; <strong>Tab</strong> for <strong>Desktop</strong>, <strong>Mobile</strong> &amp; <strong>Tablet</strong> respectively.
</p>
<pre class="r"><code>factor(device, 
       levels = c("Desktop", "Mobile", "Tablet"),
       labels = c("Desk", "Mob", "Tab"))</code></pre>
<pre><code>##  [1] Mob  Mob  Tab  Mob  Mob  Mob  Tab  Tab  Desk Mob  Mob  Mob  Tab  Mob  Desk
## [16] Tab  Desk Mob  Mob  Tab  Desk Mob  Tab  Mob  Mob 
## Levels: Desk Mob Tab</code></pre>
<p>
You can see that not only the values but the levels are also modified.
</p>
</section>
<section id="missing-values" class="level3">
<h3 class="anchored" data-anchor-id="missing-values">
Missing Values
</h3>
<p>
Let us regenerate the device column but include some missing values (<code>NA</code>) deliberately to see how <code>factor()</code> handles them.
</p>
<pre class="r"><code># sample with missing values
device &lt;- sample(c("Desktop", "Mobile", "Tablet", NA), size = 25, replace = TRUE)
device</code></pre>
<pre><code>##  [1] "Desktop" "Tablet"  "Tablet"  NA        "Mobile"  NA        "Desktop"
##  [8] "Desktop" "Tablet"  "Mobile"  "Tablet"  "Desktop" "Desktop" "Mobile" 
## [15] NA        "Desktop" "Mobile"  "Tablet"  "Mobile"  "Mobile"  "Tablet" 
## [22] NA        "Desktop" "Tablet"  "Mobile"</code></pre>
<pre class="r"><code># store as categorical data
factor(device)</code></pre>
<pre><code>##  [1] Desktop Tablet  Tablet  &lt;NA&gt;    Mobile  &lt;NA&gt;    Desktop Desktop Tablet 
## [10] Mobile  Tablet  Desktop Desktop Mobile  &lt;NA&gt;    Desktop Mobile  Tablet 
## [19] Mobile  Mobile  Tablet  &lt;NA&gt;    Desktop Tablet  Mobile 
## Levels: Desktop Mobile Tablet</code></pre>
<p>
<code>NA</code> is not shown as one of the levels. Why does this happen? By default, it will ignore them. If you look at the arguments of the <code>factor()</code> function, the <code>exclude</code> argument is set to <code>NA</code> by default i.e.&nbsp;<code>NA</code> is excluded automatically. What should we do to ensure that <code>NA</code> is also treated as a level? In order to treat <code>NA</code> as a level, set the <code>exclude</code> argument to <code>NULL</code>.
</p>
<pre class="r"><code>factor(device, exclude = NULL)</code></pre>
<pre><code>##  [1] Desktop Tablet  Tablet  &lt;NA&gt;    Mobile  &lt;NA&gt;    Desktop Desktop Tablet 
## [10] Mobile  Tablet  Desktop Desktop Mobile  &lt;NA&gt;    Desktop Mobile  Tablet 
## [19] Mobile  Mobile  Tablet  &lt;NA&gt;    Desktop Tablet  Mobile 
## Levels: Desktop Mobile Tablet &lt;NA&gt;</code></pre>
<p>
As you can see, <code>NA</code> is displayed as one of the levels in the data.
</p>
</section>
<section id="ordered-factors" class="level3">
<h3 class="anchored" data-anchor-id="ordered-factors">
Ordered Factors
</h3>
<p>
So far, we have been looking at nominal data. Let us now explore how R handles ordered data. We will generate a new data set of satisfaction ratings to use in this section. Satisfaction ratings are widely used to measure a customer’s satisfaction with an organization, service or a product.
</p>
<pre class="r"><code>rating &lt;- sample(c("Dislike", "Neutral", "Like"), size = 25, replace = TRUE)
rating</code></pre>
<pre><code>##  [1] "Like"    "Dislike" "Neutral" "Dislike" "Neutral" "Like"    "Neutral"
##  [8] "Dislike" "Dislike" "Dislike" "Dislike" "Neutral" "Dislike" "Neutral"
## [15] "Dislike" "Dislike" "Like"    "Neutral" "Dislike" "Neutral" "Like"   
## [22] "Neutral" "Dislike" "Like"    "Dislike"</code></pre>
<p>
It consists of three values Dislike, Neutral &amp; Like in that order. You can see that there is an intrinsic order here. Like is better than neutral which in turn is better than dislike. While we can order them, we can’t quantify the difference between them. We can’t say neutral is so many times better than dislike.
</p>
<p>
<strong>Membership Testing</strong>
</p>
<p>
As we did earlier, let us check if the data is ordered using the membership function <code>is.ordered()</code>.
</p>
<pre class="r"><code>is.ordered(rating)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<p>
R returns <code>FALSE</code> as the variable <code>rating</code> is not ordered. Let us use <code>as.ordered()</code> to coerce it into an ordered factor.
</p>
<pre class="r"><code>as.ordered(rating)</code></pre>
<pre><code>##  [1] Like    Dislike Neutral Dislike Neutral Like    Neutral Dislike Dislike
## [10] Dislike Dislike Neutral Dislike Neutral Dislike Dislike Like    Neutral
## [19] Dislike Neutral Like    Neutral Dislike Like    Dislike
## Levels: Dislike &lt; Like &lt; Neutral</code></pre>
<p>
Look at the last line where the levels are displayed. In case of ordered factors, you will see a <code>&lt;</code> between the labels. This is used to indicate the order of the levels. Now <code>rating</code> is both an ordered but the order of the levels is not correct. It should be <code>Dislike &lt; Neutral &lt; Like</code> but is displayed in order of appearance in the data. Let us use the <code>factor()</code> function since we need more control over how the levels are ranked and set the <code>ordered</code> argument to <code>TRUE</code>.
</p>
<pre class="r"><code>factor(rating, ordered = TRUE)</code></pre>
<pre><code>##  [1] Like    Dislike Neutral Dislike Neutral Like    Neutral Dislike Dislike
## [10] Dislike Dislike Neutral Dislike Neutral Dislike Dislike Like    Neutral
## [19] Dislike Neutral Like    Neutral Dislike Like    Dislike
## Levels: Dislike &lt; Like &lt; Neutral</code></pre>
<p>
The ranking of the levels has not changed and is still the same. Why is this happening? If you observe carefully, the ranking follows the alphabetical order (<strong>D</strong>esktop, <strong>M</strong>obile, <strong>T</strong>able). The <code>factor()</code> function uses the same order for the levels.
</p>
</section>
<section id="modify-order-of-levels" class="level3">
<h3 class="anchored" data-anchor-id="modify-order-of-levels">
Modify Order of Levels
</h3>
<p>
To change the order/ranking of the levels, we need to specify it using the <code>levels</code> argument. Let us do that in the next example.
</p>
<pre class="r"><code>factor(rating, levels = c("Dislike", "Neutral", "Like"), ordered = TRUE)</code></pre>
<pre><code>##  [1] Like    Dislike Neutral Dislike Neutral Like    Neutral Dislike Dislike
## [10] Dislike Dislike Neutral Dislike Neutral Dislike Dislike Like    Neutral
## [19] Dislike Neutral Like    Neutral Dislike Like    Dislike
## Levels: Dislike &lt; Neutral &lt; Like</code></pre>
<p>
Now, you can see that the levels are ranked correctly. The <code>ordered()</code> function can also be used to create ordered factors. Let us recreate the previous example using the <code>ordered()</code> function.
</p>
<pre class="r"><code>ordered(rating, levels = c("Dislike", "Neutral", "Like"))</code></pre>
<pre><code>##  [1] Like    Dislike Neutral Dislike Neutral Like    Neutral Dislike Dislike
## [10] Dislike Dislike Neutral Dislike Neutral Dislike Dislike Like    Neutral
## [19] Dislike Neutral Like    Neutral Dislike Like    Dislike
## Levels: Dislike &lt; Neutral &lt; Like</code></pre>
<p>
You can specify levels, modify labels and handle missing values using the <code>ordered()</code> function as well.
</p>
</section>
<section id="key-functions" class="level3">
<h3 class="anchored" data-anchor-id="key-functions">
Key Functions
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_module_3_functions.png" width="70%" style="display: block; margin: auto;">
</p>
</section>
<section id="summary" class="level3">
<h3 class="anchored" data-anchor-id="summary">
Summary
</h3>
<ul>
<li>
R uses <strong>factor</strong> to handle categorical data.
</li>
<li>
Use <code>as.factor()</code> or <code>as_factor()</code> to coerce other data types to <strong>factor</strong>.
</li>
<li>
Use <code>is.factor()</code> or <code>is.ordered()</code> to identify factor &amp; ordered factor respectively.
</li>
<li>
Use <code>factor()</code> to
<ul>
<li>
specify labels
</li>
<li>
modify labels
</li>
<li>
handle missing data
</li>
<li>
create ordered factors
</li>
<li>
specify order of levels
</li>
</ul>
</li>
<li>
Use <code>ordered()</code> to create ordered factors.
</li>
</ul>
</section>
<section id="your-turn" class="level3">
<h3 class="anchored" data-anchor-id="your-turn">
Your Turn…
</h3>
<p>
Use <code>analytics_raw.rds</code> data set to answer the below questions.
</p>
<ol style="list-style-type: decimal">
<li>
<p>
Check whether the below variables are factor
</p>
<ul>
<li>
<code>device</code>
</li>
<li>
<code>page_depth</code>
</li>
<li>
<code>landing_page</code>
</li>
</ul>
</li>
<li>
<p>
Coerce the following variables to type factor
</p>
<ul>
<li>
<code>device</code>
</li>
<li>
<code>os</code>
</li>
<li>
<code>browser</code>
</li>
<li>
<code>user_type</code>
</li>
<li>
<code>channel</code>
</li>
<li>
<code>gender</code>
</li>
<li>
<code>landing_page</code>
</li>
<li>
<code>exit_page</code>
</li>
<li>
<code>city</code>
</li>
<li>
<code>country</code>
</li>
<li>
<code>user_type</code>
</li>
</ul>
</li>
<li>
<p>
Use only the following levels in the <code>gender</code> column:
</p>
<ul>
<li>
<code>male</code>
</li>
<li>
<code>female</code>
</li>
</ul>
</li>
<li>
<p>
Include <code>NA</code> as a level in the gender column.
</p>
</li>
<li>
<p>
Change label of <code>NA</code> to <code>missing</code> in the <code>gender</code> column.
</p>
</li>
<li>
<p>
Change the labels of the levels in <code>user_type</code> column to
</p>
<ul>
<li>
<code>New</code>
</li>
<li>
<code>Returning</code>
</li>
</ul>
</li>
<li>
<p>
Check if the <code>user_rating</code> column is ordered. If not, coerce it to type ordered factor.
</p>
</li>
</ol>
<p>
*As the reader of this blog, you are our most important critic and commentator. We value your opinion and want to know what we are doing right, what we could do better, what areas you would like to see us publish in, and any other words of wisdom you are willing to pass our way.
</p>
<p>
We welcome your comments. You can email to let us know what you did or did not like about our blog as well as what we can do to make our post better.*
</p>
<p>
<strong>Email: <a href="mailto:support@rsquaredacademy.com" class="email">support@rsquaredacademy.com</a></strong>
</p>
</section>
</section>



 ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-categorical-data-in-r-part-1/</guid>
  <pubDate>Fri, 07 Jan 2022 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/forcats-part-1.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>A Comprehensive Introduction to Handling Date &amp; Time in R</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-date-and-time-in-r/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2020-04-17-handling-date-and-time-in-r.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
<span class="screen-reader-only">Note</span>Updated September 2026
</div>
</div>
<div class="callout-body-container callout-body">
<p>All code re-verified on R 4.5.2 + lubridate 1.9.4 — no API changes needed. <code>clock</code> deep-dive ships as a standalone follow-up.</p>
</div>
</div>
<script src="../../rmarkdown-libs/kePrint/kePrint.js"></script>
<p>
<img src="https://blog.rsquaredacademy.com/img/handling-date-and-time-in-r.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
In this tutorial, we will learn to handle date &amp; time in R. We will start off by learning how to <strong>get current date &amp; time</strong> before moving on to understand <strong>how R handles date/time internally</strong> and the different classes such as <code>Date</code> &amp; <code>POSIXct/lt</code>. We will spend some time exploring <strong>time zones, daylight savings and ISO 8001 standard</strong> for representing date/time. We will look at all the <strong>weird formats in which date/time come in real world</strong> and learn to <strong>parse them using conversion specifications</strong>. After this, we will also <strong>learn how to handle date/time columns while reading external data into R</strong>. We will learn to <strong>extract and update different date/time components</strong> such as year, month, day, hour, minute etc., <strong>create sequence of dates</strong> in different ways and explore intervals, durations and period. We will end the tutorial by learning how to <strong>round/rollback dates</strong>. Throughout the tutorial, we will also <strong>work through a case study</strong> to better understand the concepts we learn. Happy learning!
</p>
<section id="table-of-contents" class="level2">
<h2 class="anchored" data-anchor-id="table-of-contents">
Table of Contents
</h2>
<ul>
<li>
Resources
</li>
<li>
Introduction
</li>
<li>
Case Study
</li>
<li>
Date &amp; Time Classes
</li>
<li>
Date Arithmetic
</li>
<li>
Timezones &amp; Daylight Savings
</li>
<li>
Date &amp; Time Formats
</li>
<li>
Parse/Read Date &amp; Time
</li>
<li>
Date &amp; Time Components
</li>
<li>
Create, Update &amp; Verify
</li>
<li>
Intervals, Durations &amp; Period
</li>
<li>
Round &amp; Rollback
</li>
<li>
References
</li>
</ul>
</section>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
Below are the links to all the resources related to this tutorial:
</p>
<ul>
<li>
<a href="https://slides.rsquaredacademy.com/handling-date-and-time-in-r.pdf" target="_blank">Slides</a>
</li>
<li>
<a href="https://github.com/rsquaredacademy-education/online-courses/" target="_blank">Code &amp; Data</a>
</li>
<li>
<a href="https://rstudio.cloud/project/1072419" target="_blank">RStudio Cloud Project</a>
</li>
<li>
<a href="https://wrangle-r.rsquaredacademy.com/lubridate.html" target="_blank">ebook</a>
</li>
</ul>
<p align="center">
<a href="https://rsquared-academy.thinkific.com/courses/handling-date-and-time-in-R" target="_blank"><img src="https://blog.rsquaredacademy.com/img/lubirdate-blog-course-ad.png" width="100%" alt="new courses ad" style="text-decoration: none;"></a>
</p>
</section>
<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">
Introduction
</h2>
{{% youtube “322IcnZiYx4” %}}
<section id="date" class="level3">
<h3 class="anchored" data-anchor-id="date">
Date
</h3>
<p>
Let us begin by looking at the current date and time. <code>Sys.Date()</code> and <code>today()</code> will return the current date.
</p>
<pre class="r"><code>Sys.Date()</code></pre>
<pre><code>## [1] "2020-06-10"</code></pre>
<pre class="r"><code>lubridate::today()</code></pre>
<pre><code>## [1] "2020-06-10"</code></pre>
</section>
<section id="time" class="level3">
<h3 class="anchored" data-anchor-id="time">
Time
</h3>
<p>
<code>Sys.time()</code> and <code>now()</code> return the date, time and the timezone. In <code>now()</code>, we can specify the timezone using the <code>tzone</code> argument.
</p>
<pre class="r"><code>Sys.time()</code></pre>
<pre><code>## [1] "2020-06-10 22:47:50 IST"</code></pre>
<pre class="r"><code>lubridate::now()</code></pre>
<pre><code>## [1] "2020-06-10 22:47:50 IST"</code></pre>
<pre class="r"><code>lubridate::now(tzone = "UTC")</code></pre>
<pre><code>## [1] "2020-06-10 17:17:50 UTC"</code></pre>
</section>
<section id="am-or-pm" class="level3">
<h3 class="anchored" data-anchor-id="am-or-pm">
AM or PM?
</h3>
<p>
<code>am()</code> and <code>pm()</code> allow us to check whether date/time occur in the <code>AM</code> or <code>PM</code>? They return a logical value i.e.&nbsp;<code>TRUE</code> or <code>FALSE</code>
</p>
<pre class="r"><code>lubridate::am(now())</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>lubridate::pm(now())</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="leap-year" class="level3">
<h3 class="anchored" data-anchor-id="leap-year">
Leap Year
</h3>
<p>
We can also check if the current year is a leap year using <code>leap_year()</code>.
</p>
<pre class="r"><code>Sys.Date()</code></pre>
<pre><code>## [1] "2020-06-10"</code></pre>
<pre class="r"><code>lubridate::leap_year(Sys.Date())</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="summary" class="level3">
<h3 class="anchored" data-anchor-id="summary">
Summary
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Function
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>Sys.Date()</code>
</td>
<td style="text-align:left;">
Current Date
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::today()</code>
</td>
<td style="text-align:left;">
Current Date
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>Sys.time()</code>
</td>
<td style="text-align:left;">
Current Time
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::now()</code>
</td>
<td style="text-align:left;">
Current Time
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::am()</code>
</td>
<td style="text-align:left;">
Whether time occurs in am?
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::pm()</code>
</td>
<td style="text-align:left;">
Whether time occurs in pm?
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::leap_year()</code>
</td>
<td style="text-align:left;">
Check if the year is a leap year?
</td>
</tr>
</tbody>
</table>
</section>
<section id="your-turn" class="level3">
<h3 class="anchored" data-anchor-id="your-turn">
Your Turn
</h3>
<ul>
<li>
get current date
</li>
<li>
get current time
</li>
<li>
check whether the time occurs in am or pm?
</li>
<li>
check whether the following years were leap years
<ul>
<li>
2018
</li>
<li>
2016
</li>
</ul>
</li>
</ul>
</section>
</section>
<section id="casestudy" class="level2">
<h2 class="anchored" data-anchor-id="casestudy">
Case Study
</h2>
<p>
Throughout the tutorial, we will work on a case study related to transactions of an imaginary trading company. The data set includes information about invoice and payment dates.
</p>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>transact &lt;- readr::read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/transact.csv')</code></pre>
<pre><code>## # A tibble: 2,466 x 3
##    Invoice    Due        Payment   
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;    
##  1 2013-01-02 2013-02-01 2013-01-15
##  2 2013-01-26 2013-02-25 2013-03-03
##  3 2013-07-03 2013-08-02 2013-07-08
##  4 2013-02-10 2013-03-12 2013-03-17
##  5 2012-10-25 2012-11-24 2012-11-28
##  6 2012-01-27 2012-02-26 2012-02-22
##  7 2013-08-13 2013-09-12 2013-09-09
##  8 2012-12-16 2013-01-15 2013-01-12
##  9 2012-05-14 2012-06-13 2012-07-01
## 10 2013-07-01 2013-07-31 2013-07-26
## # ... with 2,456 more rows</code></pre>
<p>
We will explore more about reading data sets with date/time columns after learning how to parse date/time. We have shared the code for reading the data sets used in the practice questions both in the Learning Management System as well as in our GitHub repo.
</p>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<p>
The data set has 3 columns. All the dates are in the format (yyyy-mm-dd).
</p>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Column
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
Invoice
</td>
<td style="text-align:left;">
Invoice Date
</td>
</tr>
<tr>
<td style="text-align:left;">
Due
</td>
<td style="text-align:left;">
Due Date
</td>
</tr>
<tr>
<td style="text-align:left;">
Payment
</td>
<td style="text-align:left;">
Payment Date
</td>
</tr>
</tbody>
</table>
<p>
In the case study, we will try to answer a few questions we have about the <code>transact</code> data.
</p>
<ul>
<li>
extract date, month and year from Due
</li>
<li>
compute the number of days to settle invoice
</li>
<li>
compute days over due
</li>
<li>
check if due year is a leap year
</li>
<li>
check when due day in february is 29, whether it is a leap year
</li>
<li>
how many invoices were settled within due date
</li>
<li>
how many invoices are due in each quarter
</li>
</ul>
</section>
</section>
<section id="classes" class="level2">
<h2 class="anchored" data-anchor-id="classes">
Date &amp; Time Classes
</h2>
{{% youtube “IEX49t8sSgw” %}}
<p>
In this section, we will look at two things. First, how to create date/time data in R, and second, how to convert other data types to date/time. Let us begin by creating the release date of R 3.6.2.
</p>
<pre class="r"><code>release_date &lt;- 2019-12-12
release_date</code></pre>
<pre><code>## [1] 1995</code></pre>
<p>
Okay! Why do we see <code>1995</code> when we call the date? What is happening here? Let us quickly check the data type of <code>release_date</code>.
</p>
<pre class="r"><code>class(release_date)</code></pre>
<pre><code>## [1] "numeric"</code></pre>
<p>
The data type is <code>numeric</code> i.e.&nbsp;R has subtracted <code>12</code> twice from <code>2019</code> to return <code>1995</code>. Clearly, the above method is not the right way to store date/time. Let us see if we can get some hints from the built-in R functions we used in the previous section. If you observe the output, all of them returned date/time wrapped in quotes. Hmmm… let us wrap our date in quotes and see what happens.
</p>
<pre class="r"><code>release_date &lt;- "2019-12-12"
release_date</code></pre>
<pre><code>## [1] "2019-12-12"</code></pre>
<p>
Alright, now R does not do any arithmetic and returns the date as we specified. Great! Is this the right format to store date/time then? No.&nbsp;Why? What is the problem if date/time is saved as character/string? The problem is the nature or type of operations done on date/time are different when compared to string/character, number or logical values.
</p>
<ul>
<li>
how do we add/subtract dates?
</li>
<li>
how do we extract components such as year, month, day etc.
</li>
</ul>
<p>
To answer the above questions, we will first check the data type of <code>Sys.Date()</code> and <code>now()</code>.
</p>
<pre class="r"><code>class(Sys.Date())</code></pre>
<pre><code>## [1] "Date"</code></pre>
<pre class="r"><code>class(lubridate::now())</code></pre>
<pre><code>## [1] "POSIXct" "POSIXt"</code></pre>
<pre class="r"><code>class(release_date)</code></pre>
<pre><code>## [1] "character"</code></pre>
<p>
As you can see from the above output, there are 3 different classes for storing date/time in R
</p>
<ul>
<li>
<code>Date</code>
</li>
<li>
<code>POSIXct</code>
</li>
<li>
<code>POSIXlt</code>
</li>
</ul>
<p>
Let us explore each of the above classes one by one.
</p>
<section id="date-1" class="level3">
<h3 class="anchored" data-anchor-id="date-1">
Date
</h3>
<section id="introduction" class="level4">
<h4 class="anchored" data-anchor-id="introduction">
Introduction
</h4>
<p>
The <code>Date</code> class represents calendar dates. Let us go back to <code>Sys.Date()</code>. If you check the class of <code>Sys.Date()</code>, it is <code>Date</code>. Internally, this date is a number i.e.&nbsp;an integer. The <code>unclass()</code> function will show how dates are stored internally.
</p>
<pre class="r"><code>Sys.Date()</code></pre>
<pre><code>## [1] "2020-06-10"</code></pre>
<pre class="r"><code>unclass(Sys.Date())</code></pre>
<pre><code>## [1] 18423</code></pre>
<p>
What does this integer represent? Why has R stored the date as an integer? In R, dates are represented as the number of days since <code>1970-01-01</code>. All the dates in R are internally stored in this way. Before we explore this concept further, let us learn to create <code>Date</code> objects in R. We will continue to use the release date of R 3.6.2, <code>2019-12-12</code>.
</p>
<p>
Until now, we have stored the above date as character/string but now we will use <code>as.Date()</code> to save it as a <code>Date</code> object. <code>as.Date()</code> is the easiest and simplest way to create dates in R.
</p>
<pre class="r"><code>release_date &lt;- as.Date("2019-12-12")
release_date</code></pre>
<pre><code>## [1] "2019-12-12"</code></pre>
<p>
The <code>as_date()</code> function from the lubridate package is similar to <code>as.Date()</code>.
</p>
<pre class="r"><code>release_date &lt;- lubridate::as_date("2019-12-12")
release_date</code></pre>
<pre><code>## [1] "2019-12-12"</code></pre>
<p>
If you look at the difference between <code>release_date</code> and <code>1970-01-01</code>, it will be the same as <code>unclass(release_date)</code>.
</p>
<pre class="r"><code>release_date - as.Date("1970-01-01")</code></pre>
<pre><code>## Time difference of 18242 days</code></pre>
<pre class="r"><code>unclass(release_date)</code></pre>
<pre><code>## [1] 18242</code></pre>
<p>
Let us come back to <code>1970-01-01</code> i.e.&nbsp;the origin for dates in R.
</p>
<pre class="r"><code>lubridate::origin</code></pre>
<pre><code>## [1] "1970-01-01 UTC"</code></pre>
<p>
From the previous examples, we know that dates are internally stored as number of days since <code>1970-01-01</code>. How about dates older than the origin? How are they stored? Let us look at that briefly.
</p>
<pre class="r"><code>unclass(as.Date("1963-08-28"))</code></pre>
<pre><code>## [1] -2318</code></pre>
<p>
Dates older than the origin are stored as negative integers. For those who are not aware, Martin Luther King, Jr.&nbsp;delivered his famous <strong>I Have a Dream</strong> speech on <code>1963-08-28</code>. Let us move on and learn how to convert numbers into dates.
</p>
</section>
<section id="convert-numeric" class="level4">
<h4 class="anchored" data-anchor-id="convert-numeric">
Convert Numeric
</h4>
<p>
The <code>as.Date()</code> function can be used to convert any of the following to a <code>Date</code> object
</p>
<ul>
<li>
character/string
</li>
<li>
number
</li>
<li>
factor (categorical/qualitative)
</li>
</ul>
<p>
We have explored how to convert strings to date. How about converting numbers to date? Sure, we can create date from numbers by specifying the origin and number of days since it.
</p>
<pre class="r"><code>as.Date(18242, origin = "1970-01-01")</code></pre>
<pre><code>## [1] "2019-12-12"</code></pre>
<p>
The origin can be changed to another date (while changing the number as well.)
</p>
<pre class="r"><code>as.Date(7285, origin = "2000-01-01")</code></pre>
<pre><code>## [1] "2019-12-12"</code></pre>
</section>
</section>
<section id="iso-8601" class="level3">
<h3 class="anchored" data-anchor-id="iso-8601">
ISO 8601
</h3>
<p>
<img src="https://blog.rsquaredacademy.com/img/iso.png" width="70%" style="display: block; margin: auto;">
</p>
<p>
If you have carefully observed, the format in which we have been specifying the dates as well as of those returned by functions such as <code>Sys.Date()</code> or <code>Sys.time()</code> is the same i.e.&nbsp;<code>YYYY-MM-DD</code>. It includes
</p>
<ul>
<li>
the year including the century
</li>
<li>
the month
</li>
<li>
the date
</li>
</ul>
<p>
The month and date separated by <code>-</code>. This default format used in R is the ISO 8601 standard for date/time. ISO 8601 is the internationally accepted way to represent dates and times and uses the 24 hour clock system. Let us create the release date using another function <code>ISOdate()</code>.
</p>
<pre class="r"><code>ISOdate(year  = 2019,
        month = 12,
        day   = 12,
        hour  = 8,
        min   = 5, 
        sec   = 3,
        tz    = "UTC")</code></pre>
<pre><code>## [1] "2019-12-12 08:05:03 UTC"</code></pre>
<p>
We will look at all the different weird ways in which date/time are specified in the real world in the Date &amp; Time Formats section. For the time being, let us continue exploring date/time classes in R. The next class we are going to look at is <code>POSIXct/POSIXlt</code>.
</p>
</section>
<section id="posix" class="level3">
<h3 class="anchored" data-anchor-id="posix">
POSIX
</h3>
<p>
You might be wondering what is this POSIX thing? POSIX stands for <strong>P</strong>ortable <strong>O</strong>perating <strong>S</strong>ystem <strong>I</strong>nterface. It is a family of standards specified for maintaining compatibility between different operating systems. Before we learn to create POSIX objects, let us look at <code>now()</code> from lubridate.
</p>
<pre class="r"><code>class(lubridate::now())</code></pre>
<pre><code>## [1] "POSIXct" "POSIXt"</code></pre>
<p>
<code>now()</code> returns current date/time as a POSIXct object. Let us look at its internal representation using <code>unclass()</code>
</p>
<pre class="r"><code>unclass(lubridate::now())</code></pre>
<pre><code>## [1] 1591809475
## attr(,"tzone")
## [1] ""</code></pre>
<p>
The output you see is the number of seconds since January 1, 1970.
</p>
<section id="posixct" class="level4">
<h4 class="anchored" data-anchor-id="posixct">
POSIXct
</h4>
<p>
<code>POSIXct</code> represents the number of seconds since the beginning of 1970 (UTC) and <code>ct</code> stands for calendar time. To store date/time as <code>POSIXct</code> objects, use <code>as.POSIXct()</code>. Let us now store the release date of R 3.6.2 as <code>POSIXct</code> as shown below
</p>
<pre class="r"><code>release_date &lt;- as.POSIXct("2019-12-12 08:05:03")
class(release_date)</code></pre>
<pre><code>## [1] "POSIXct" "POSIXt"</code></pre>
<pre class="r"><code>unclass(release_date) </code></pre>
<pre><code>## [1] 1576118103
## attr(,"tzone")
## [1] ""</code></pre>
</section>
<section id="posixlt" class="level4">
<h4 class="anchored" data-anchor-id="posixlt">
POSIXlt
</h4>
<p>
<code>POSIXlt</code> represents the following information in a list
</p>
<ul>
<li>
seconds
</li>
<li>
minutes
</li>
<li>
hour
</li>
<li>
day of the month
</li>
<li>
month
</li>
<li>
year
</li>
<li>
day of week
</li>
<li>
day of year
</li>
<li>
daylight saving time flag
</li>
<li>
time zone
</li>
<li>
offset in seconds from GMT
</li>
</ul>
<p>
The <code>lt</code> in <code>POSIXlt</code> stands for local time. Use <code>as.POSIXlt()</code> to store date/time as <code>POSIXlt</code> objects. Let us store the release date as a <code>POSIXlt</code> object as shown below
</p>
<pre class="r"><code>release_date &lt;- as.POSIXlt("2019-12-12 08:05:03")
release_date</code></pre>
<pre><code>## [1] "2019-12-12 08:05:03 IST"</code></pre>
<p>
As we said earlier, <code>POSIXlt</code> stores date/time components in a list and these can be extracted. Let us look at the date/time components returned by <code>POSIXlt</code> using <code>unclass()</code>.
</p>
<pre class="r"><code>release_date &lt;- as.POSIXlt("2019-12-12 08:05:03")
unclass(release_date)</code></pre>
<pre><code>## $sec
## [1] 3
## 
## $min
## [1] 5
## 
## $hour
## [1] 8
## 
## $mday
## [1] 12
## 
## $mon
## [1] 11
## 
## $year
## [1] 119
## 
## $wday
## [1] 4
## 
## $yday
## [1] 345
## 
## $isdst
## [1] 0
## 
## $zone
## [1] "IST"
## 
## $gmtoff
## [1] NA</code></pre>
<p>
Use <code>unlist()</code> if you want the components returned as a <code>vector</code>.
</p>
<pre class="r"><code>release_date &lt;- as.POSIXlt("2019-12-12 08:05:03")
unlist(release_date)</code></pre>
<pre><code>##    sec    min   hour   mday    mon   year   wday   yday  isdst   zone gmtoff 
##    "3"    "5"    "8"   "12"   "11"  "119"    "4"  "345"    "0"  "IST"     NA</code></pre>
<p>
To extract specific components, use <code><img src="https://latex.codecogs.com/png.latex?%3C/code%3E.%3C/p%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3Erelease_date%20&amp;lt;-%20as.POSIXlt(&amp;quot;2019-12-12%2008:05:03&amp;quot;)%0Arelease_date">hour</code>

</p><pre><code>## [1] 8</code></pre>
<pre class="r"><code>release_date$mon</code></pre>
<pre><code>## [1] 11</code></pre>
<pre class="r"><code>release_date$zone</code></pre>
<pre><code>## [1] "IST"</code></pre>
<p>
Now, let us look at the components returned by <code>POSIXlt</code>. Some of them are intuitive
</p>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Component
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>sec</code>
</td>
<td style="text-align:left;">
Second
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>min</code>
</td>
<td style="text-align:left;">
Minute
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>hour</code>
</td>
<td style="text-align:left;">
Hour of the day
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>mon</code>
</td>
<td style="text-align:left;">
Month of the year (0-11
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>zone</code>
</td>
<td style="text-align:left;">
Timezone
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>wday</code>
</td>
<td style="text-align:left;">
Day of week
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>mday</code>
</td>
<td style="text-align:left;">
Day of month
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>year</code>
</td>
<td style="text-align:left;">
Years since 1900
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>yday</code>
</td>
<td style="text-align:left;">
Day of year
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>isdst</code>
</td>
<td style="text-align:left;">
Daylight saving flag
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>gmtoff</code>
</td>
<td style="text-align:left;">
Offset is seconds from GMT
</td>
</tr>
</tbody>
</table>
<p>
Great! We will end this section with a few tips/suggestions on when to use <code>Date</code> or <code>POSIXct/POSIXlt</code>.
</p>
<ul>
<li>
use <code>Date</code> when there is no time component
</li>
<li>
use <code>POSIX</code> when dealing with time and timezones
</li>
<li>
use <code>POSIXlt</code> when you want to access/extract the different components
</li>
</ul>
</section></section></section> ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-date-and-time-in-r/</guid>
  <pubDate>Fri, 17 Apr 2020 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/handling-date-and-time-in-r.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>Date &amp; Time in R - Introduction</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/handling-date-and-time-in-r-part-1/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2020-04-17-handling-date-and-time-in-r-part-1.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<script src="../../rmarkdown-libs/header-attrs/header-attrs.js"></script>
<script src="../../rmarkdown-libs/kePrint/kePrint.js"></script>
<p><link href="../../rmarkdown-libs/lightable/lightable.css" rel="stylesheet"></p>
<p>
<img src="https://blog.rsquaredacademy.com/img/handling-date-and-time-in-r.png" width="80%" style="display: block; margin: auto;">
</p>
<p>
In this new series <a href="https://tutorials.rsquaredacademy.com/date-time/index.html">Handling Date &amp; Time in R</a>, we will learn to handle date &amp; time in R. We will start off by learning how to <strong>get current date &amp; time</strong> before moving on to understand <strong>how R handles date/time internally</strong> and the different classes such as <code>Date</code> &amp; <code>POSIXct/lt</code>. We will spend some time exploring <strong>time zones, daylight savings and ISO 8001 standard</strong> for representing date/time. We will look at all the <strong>weird formats in which date/time come in real world</strong> and learn to <strong>parse them using conversion specifications</strong>. After this, we will also <strong>learn how to handle date/time columns while reading external data into R</strong>. We will learn to <strong>extract and update different date/time components</strong> such as year, month, day, hour, minute etc., <strong>create sequence of dates</strong> in different ways and explore intervals, durations and period. We will end the tutorial by learning how to <strong>round/rollback dates</strong>. Throughout the series, we will also <strong>work through a case study</strong> to better understand the concepts we learn. Happy learning!
</p>
<section id="resources" class="level2">
<h2 class="anchored" data-anchor-id="resources">
Resources
</h2>
<p>
Below are the links to all the resources related to this tutorial:
</p>
<ul>
<li>
<a href="https://slides.rsquaredacademy.com/handling-date-and-time-in-r.pdf" target="_blank">Slides</a>
</li>
<li>
<a href="https://github.com/rsquaredacademy-education/online-courses/" target="_blank">Code &amp; Data</a>
</li>
<li>
<a href="https://rstudio.cloud/project/1072419" target="_blank">RStudio Cloud Project</a>
</li>
<li>
<a href="https://wrangle-r.rsquaredacademy.com/lubridate.html" target="_blank">ebook</a>
</li>
</ul>
<p align="center">
<a href="https://rsquared-academy.thinkific.com/courses/handling-date-and-time-in-R" target="_blank"><img src="https://blog.rsquaredacademy.com/img/lubirdate-blog-course-ad.png" width="100%" alt="new courses ad" style="text-decoration: none;"></a>
</p>
</section>
<section id="intro" class="level2">
<h2 class="anchored" data-anchor-id="intro">
Introduction
</h2>
{{% youtube “322IcnZiYx4” %}}
<section id="date" class="level3">
<h3 class="anchored" data-anchor-id="date">
Date
</h3>
<p>
Let us begin by looking at the current date and time. <code>Sys.Date()</code> and <code>today()</code> will return the current date.
</p>
<pre class="r"><code>Sys.Date()</code></pre>
<pre><code>## [1] "2021-02-03"</code></pre>
<pre class="r"><code>lubridate::today()</code></pre>
<pre><code>## [1] "2021-02-03"</code></pre>
</section>
<section id="time" class="level3">
<h3 class="anchored" data-anchor-id="time">
Time
</h3>
<p>
<code>Sys.time()</code> and <code>now()</code> return the date, time and the timezone. In <code>now()</code>, we can specify the timezone using the <code>tzone</code> argument.
</p>
<pre class="r"><code>Sys.time()</code></pre>
<pre><code>## [1] "2021-02-03 18:54:23 IST"</code></pre>
<pre class="r"><code>lubridate::now()</code></pre>
<pre><code>## [1] "2021-02-03 18:54:23 IST"</code></pre>
<pre class="r"><code>lubridate::now(tzone = "UTC")</code></pre>
<pre><code>## [1] "2021-02-03 13:24:23 UTC"</code></pre>
</section>
<section id="am-or-pm" class="level3">
<h3 class="anchored" data-anchor-id="am-or-pm">
AM or PM?
</h3>
<p>
<code>am()</code> and <code>pm()</code> allow us to check whether date/time occur in the <code>AM</code> or <code>PM</code>? They return a logical value i.e.&nbsp;<code>TRUE</code> or <code>FALSE</code>
</p>
<pre class="r"><code>lubridate::am(now())</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>lubridate::pm(now())</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="leap-year" class="level3">
<h3 class="anchored" data-anchor-id="leap-year">
Leap Year
</h3>
<p>
We can also check if the current year is a leap year using <code>leap_year()</code>.
</p>
<pre class="r"><code>Sys.Date()</code></pre>
<pre><code>## [1] "2021-02-03"</code></pre>
<pre class="r"><code>lubridate::leap_year(Sys.Date())</code></pre>
<pre><code>## [1] FALSE</code></pre>
</section>
<section id="summary" class="level3">
<h3 class="anchored" data-anchor-id="summary">
Summary
</h3>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Function
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
<code>Sys.Date()</code>
</td>
<td style="text-align:left;">
Current Date
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::today()</code>
</td>
<td style="text-align:left;">
Current Date
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>Sys.time()</code>
</td>
<td style="text-align:left;">
Current Time
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::now()</code>
</td>
<td style="text-align:left;">
Current Time
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::am()</code>
</td>
<td style="text-align:left;">
Whether time occurs in am?
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::pm()</code>
</td>
<td style="text-align:left;">
Whether time occurs in pm?
</td>
</tr>
<tr>
<td style="text-align:left;">
<code>lubridate::leap_year()</code>
</td>
<td style="text-align:left;">
Check if the year is a leap year?
</td>
</tr>
</tbody>
</table>
</section>
<section id="your-turn" class="level3">
<h3 class="anchored" data-anchor-id="your-turn">
Your Turn
</h3>
<ul>
<li>
get current date
</li>
<li>
get current time
</li>
<li>
check whether the time occurs in am or pm?
</li>
<li>
check whether the following years were leap years
<ul>
<li>
2018
</li>
<li>
2016
</li>
</ul>
</li>
</ul>
</section>
</section>
<section id="casestudy" class="level2">
<h2 class="anchored" data-anchor-id="casestudy">
Case Study
</h2>
<p>
Throughout the tutorial, we will work on a case study related to transactions of an imaginary trading company. The data set includes information about invoice and payment dates.
</p>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>transact &lt;- readr::read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/transact.csv')</code></pre>
<pre><code>## # A tibble: 2,466 x 3
##    Invoice    Due        Payment   
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;    
##  1 2013-01-02 2013-02-01 2013-01-15
##  2 2013-01-26 2013-02-25 2013-03-03
##  3 2013-07-03 2013-08-02 2013-07-08
##  4 2013-02-10 2013-03-12 2013-03-17
##  5 2012-10-25 2012-11-24 2012-11-28
##  6 2012-01-27 2012-02-26 2012-02-22
##  7 2013-08-13 2013-09-12 2013-09-09
##  8 2012-12-16 2013-01-15 2013-01-12
##  9 2012-05-14 2012-06-13 2012-07-01
## 10 2013-07-01 2013-07-31 2013-07-26
## # ... with 2,456 more rows</code></pre>
<p>
We will explore more about reading data sets with date/time columns after learning how to parse date/time. We have shared the code for reading the data sets used in the practice questions both in the Learning Management System as well as in our GitHub repo.
</p>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<p>
The data set has 3 columns. All the dates are in the format (yyyy-mm-dd).
</p>
<table class="table table-striped table-hover table-condensed table-responsive" style="margin-left: auto; margin-right: auto;">
<thead>
<tr>
<th style="text-align:left;">
Column
</th>
<th style="text-align:left;">
Description
</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left;">
Invoice
</td>
<td style="text-align:left;">
Invoice Date
</td>
</tr>
<tr>
<td style="text-align:left;">
Due
</td>
<td style="text-align:left;">
Due Date
</td>
</tr>
<tr>
<td style="text-align:left;">
Payment
</td>
<td style="text-align:left;">
Payment Date
</td>
</tr>
</tbody>
</table>
<p>
In the case study, we will try to answer a few questions we have about the <code>transact</code> data.
</p>
<ul>
<li>
extract date, month and year from Due
</li>
<li>
compute the number of days to settle invoice
</li>
<li>
compute days over due
</li>
<li>
check if due year is a leap year
</li>
<li>
check when due day in february is 29, whether it is a leap year
</li>
<li>
how many invoices were settled within due date
</li>
<li>
how many invoices are due in each quarter
</li>
</ul>
<p>
*As the reader of this blog, you are our most important critic and commentator. We value your opinion and want to know what we are doing right, what we could do better, what areas you would like to see us publish in, and any other words of wisdom you are willing to pass our way.
</p>
<p>
We welcome your comments. You can email to let us know what you did or did not like about our blog as well as what we can do to make our post better.*
</p>
<p>
<strong>Email: <a href="mailto:support@rsquaredacademy.com" class="email">support@rsquaredacademy.com</a></strong>
</p>
</section>
</section>



 ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/handling-date-and-time-in-r-part-1/</guid>
  <pubDate>Thu, 16 Apr 2020 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/handling-date-and-time-in-r.png" medium="image" type="image/png" height="72" width="144"/>
</item>
<item>
  <title>Data Wrangling with dbplyr</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dbplyr/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-12-09-data-wrangling-with-dbplyr.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level3">
<h3 class="anchored" data-anchor-id="introduction">
Introduction
</h3>
<p>
This is the second post in the series <strong>R &amp; Databases</strong>. You can find the links to the first post of this series below:
</p>
<ul>
<li>
<a href="https://blog.rsquaredacademy.com/quick-guide-r-sqlite/">Quick Guide: R &amp; SQLite</a>
</li>
</ul>
<p>
In this post, we will learn to query data from a database using dplyr.
</p>
</section>
<section id="libraries-code-data" class="level3">
<h3 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h3>
<p>
We will use the following libraries in this post:
</p>
<ul>
<li>
<a href="http://readr.tidyverse.org/">DBI</a>
</li>
<li>
<a href="https://rstats-db.github.io/RSQLite/">RSQLite</a>
</li>
<li>
<a href="http://dbplyr.tidyverse.org/">dbplyr</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/">dplyr</a>
</li>
</ul>
<p>
All the data sets used in this post can be found <a href="https://github.com/rsquaredacademy/datasets">here</a> and code can be downloaded from <a href="https://gist.github.com/rsquaredacademy/f5ee72cee9ab3256230cdedecd3ef25b">here</a>.
</p>
<section id="connect-to-database" class="level4">
<h4 class="anchored" data-anchor-id="connect-to-database">
Connect to Database
</h4>
<p>
Let us connect to an in memory SQLite database using <code>dbConnect()</code>.
</p>
<pre class="r"><code>con &lt;- dbConnect(RSQLite::SQLite(), ":memory:")</code></pre>
<p>
We will copy the <code>mtcars</code> data to the database so that we can use it for running dplyr statements.
</p>
<pre class="r"><code>dplyr::copy_to(con, mtcars)</code></pre>
</section>
<section id="reference-copied-data-frame" class="level4">
<h4 class="anchored" data-anchor-id="reference-copied-data-frame">
Reference Copied Data Frame
</h4>
<p>
In order to use dplyr functions, we need to reference the table in the database using <code>tbl()</code>.
</p>
<pre class="r"><code>mtcars2 &lt;- dplyr::tbl(con, "mtcars")
mtcars2</code></pre>
<pre><code>## # Source:   table&lt;mtcars&gt; [?? x 11]
## # Database: sqlite 3.30.1 [:memory:]
##      mpg   cyl  disp    hp  drat    wt  qsec    vs    am  gear  carb
##    &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt;
##  1  21       6  160    110  3.9   2.62  16.5     0     1     4     4
##  2  21       6  160    110  3.9   2.88  17.0     0     1     4     4
##  3  22.8     4  108     93  3.85  2.32  18.6     1     1     4     1
##  4  21.4     6  258    110  3.08  3.22  19.4     1     0     3     1
##  5  18.7     8  360    175  3.15  3.44  17.0     0     0     3     2
##  6  18.1     6  225    105  2.76  3.46  20.2     1     0     3     1
##  7  14.3     8  360    245  3.21  3.57  15.8     0     0     3     4
##  8  24.4     4  147.    62  3.69  3.19  20       1     0     4     2
##  9  22.8     4  141.    95  3.92  3.15  22.9     1     0     4     2
## 10  19.2     6  168.   123  3.92  3.44  18.3     1     0     4     4
## # ... with more rows</code></pre>
</section>
<section id="query-data" class="level4">
<h4 class="anchored" data-anchor-id="query-data">
Query Data
</h4>
<p>
We will look at some simple examples. Let us start by selecting <code>mpg</code>, <code>cyl</code> and <code>drat</code> columns from <code>mtcars2</code>.
</p>
<pre class="r"><code>select(mtcars2, mpg, cyl, drat)</code></pre>
<pre><code>## # Source:   lazy query [?? x 3]
## # Database: sqlite 3.30.1 [:memory:]
##      mpg   cyl  drat
##    &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt;
##  1  21       6  3.9 
##  2  21       6  3.9 
##  3  22.8     4  3.85
##  4  21.4     6  3.08
##  5  18.7     8  3.15
##  6  18.1     6  2.76
##  7  14.3     8  3.21
##  8  24.4     4  3.69
##  9  22.8     4  3.92
## 10  19.2     6  3.92
## # ... with more rows</code></pre>
<p>
We can filter data as well. Filter all the rows from <code>mtcars2</code> where <code>mpg</code> is greater than 25.
</p>
<pre class="r"><code>filter(mtcars2, mpg &gt; 25)</code></pre>
<pre><code>## # Source:   lazy query [?? x 11]
## # Database: sqlite 3.30.1 [:memory:]
##     mpg   cyl  disp    hp  drat    wt  qsec    vs    am  gear  carb
##   &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt; &lt;dbl&gt;
## 1  32.4     4  78.7    66  4.08  2.2   19.5     1     1     4     1
## 2  30.4     4  75.7    52  4.93  1.62  18.5     1     1     4     2
## 3  33.9     4  71.1    65  4.22  1.84  19.9     1     1     4     1
## 4  27.3     4  79      66  4.08  1.94  18.9     1     1     4     1
## 5  26       4 120.     91  4.43  2.14  16.7     0     1     5     2
## 6  30.4     4  95.1   113  3.77  1.51  16.9     1     1     5     2</code></pre>
<p>
Time to do some grouping and summarizing. Let us compute the average mileage for different types of cylinders.
</p>
<pre class="r"><code>mtcars2 %&gt;%
  group_by(cyl) %&gt;%
  summarise(mileage = mean(mpg))</code></pre>
<pre><code>## Warning: Missing values are always removed in SQL.
## Use `mean(x, na.rm = TRUE)` to silence this warning
## This warning is displayed only once per session.</code></pre>
<pre><code>## # Source:   lazy query [?? x 2]
## # Database: sqlite 3.30.1 [:memory:]
##     cyl mileage
##   &lt;dbl&gt;   &lt;dbl&gt;
## 1     4    26.7
## 2     6    19.7
## 3     8    15.1</code></pre>
</section>
<section id="show-query" class="level4">
<h4 class="anchored" data-anchor-id="show-query">
Show Query
</h4>
<p>
If you want to view the SQL query generated in the above step, use <code>show_query()</code> or <code>explain()</code>.
</p>
<pre class="r"><code>mileages &lt;- 
  mtcars2 %&gt;%
  group_by(cyl) %&gt;%
  summarise(mileage = mean(mpg))

dplyr::show_query(mileages)
## &lt;SQL&gt;
## SELECT `cyl`, AVG(`mpg`) AS `mileage`
## FROM `mtcars`
## GROUP BY `cyl`

dplyr::explain(mileages)
## &lt;SQL&gt;
## SELECT `cyl`, AVG(`mpg`) AS `mileage`
## FROM `mtcars`
## GROUP BY `cyl`
## 
## &lt;PLAN&gt;
##   id parent notused                       detail
## 1  6      0       0            SCAN TABLE mtcars
## 2  8      0       0 USE TEMP B-TREE FOR GROUP BY</code></pre>
</section>
<section id="collect-data" class="level4">
<h4 class="anchored" data-anchor-id="collect-data">
Collect Data
</h4>
<p>
Now, some interesting facts. When working with databases, <strong>dplyr</strong> never pulls data into R unless you explicitly ask for it. In the previous example, dplyr will not do anything until you ask for the mileages data. It generates the SQL and only pulls down a few rows when you try to print <code>mileages</code>. So how do we pull all the data and store it for further analysis? <code>collect()</code> will pull all the data and store it in a tibble and you can use it for any further analysis.
</p>
<pre class="r"><code>dplyr::collect(mileages)</code></pre>
<pre><code>## # A tibble: 3 x 2
##     cyl mileage
##   &lt;dbl&gt;   &lt;dbl&gt;
## 1     4    26.7
## 2     6    19.7
## 3     8    15.1</code></pre>
</section>
</section>
<section id="references" class="level3">
<h3 class="anchored" data-anchor-id="references">
References
</h3>
<ul>
<li>
<a href="https://dbplyr.tidyverse.org/" class="uri">https://dbplyr.tidyverse.org/</a>
</li>
</ul>
</section>
<section id="up-next.." class="level3">
<h3 class="anchored" data-anchor-id="up-next..">
Up Next..
</h3>
<p>
In the next <a href="">post</a>, we will learn basic SQL commands.
</p>
</section>



 ]]></description>
  <category>database</category>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dbplyr/</guid>
  <pubDate>Sun, 09 Dec 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/database.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Quick Guide: R &amp; SQLite</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/quick-guide-r-sqlite/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-11-27-quick-guide-r-sqlite.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level3">
<h3 class="anchored" data-anchor-id="introduction">
Introduction
</h3>
<p>
This is the first post in the series <strong>R &amp; Databases</strong>. You can find the links to the other two posts of this series below:
</p>
<ul>
<li>
<a href="https://rsquaredacademy.github.io/blog/post/data-wrangling-with-dbplyr">Data Wrangling with dbplyr</a>
</li>
<li>
<a href="https://rsquaredacademy.github.io/blog/post/sql-for-data-science-part-1">SQL for Data Science - Part 1</a>
</li>
<li>
<a href="https://rsquaredacademy.github.io/blog/post/sql-for-data-science-part-2">SQL for Data Science - Part 2</a>
</li>
</ul>
<p>
In this post, we will learn to:
</p>
<ul>
<li>
connect to a SQLite database from R
</li>
<li>
display database information
</li>
<li>
list tables in the database
</li>
<li>
query data
<ul>
<li>
read entire table
</li>
<li>
read few rows
</li>
<li>
read data in batches
</li>
</ul>
</li>
<li>
create table in database
</li>
<li>
overwrite table in database
</li>
<li>
append data to table in database
</li>
<li>
remove table from database
</li>
<li>
generate SQL query
</li>
<li>
close database connection
</li>
</ul>
</section>
<section id="libraries-code-data" class="level3">
<h3 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h3>
<p>
We will use the following libraries in this post:
</p>
<ul>
<li>
<a href="http://rstats-db.github.io/DBI/">DBI</a>
</li>
<li>
<a href="https://rstats-db.github.io/RSQLite/">RSQLite</a>
</li>
</ul>
<p>
All the data sets used in this post can be found <a href="https://github.com/rsquaredacademy/datasets">here</a> and code can be downloaded from <a href="https://gist.github.com/rsquaredacademy/7d99eb52a0e44cd1f87c8689cf1a307d">here</a>.
</p>
<section id="connection" class="level4">
<h4 class="anchored" data-anchor-id="connection">
Connection
</h4>
<p>
The first step is to connect to a database. In this post, we will connect to an in memory SQLite databse using <code>dbConnect()</code>.
</p>
<pre class="r"><code>con &lt;- dbConnect(RSQLite::SQLite(), ":memory:")</code></pre>
</section>
<section id="connection-summary" class="level4">
<h4 class="anchored" data-anchor-id="connection-summary">
Connection Summary
</h4>
<p>
We can get the more information about the connection using <code>summary()</code>.
</p>
<pre class="r"><code>summary(con)</code></pre>
<pre><code>##           Length            Class             Mode 
##                1 SQLiteConnection               S4</code></pre>
</section>
<section id="list-tables" class="level4">
<h4 class="anchored" data-anchor-id="list-tables">
List Tables
</h4>
<p>
Now that we are connected to a database, let us list all the tables present in it using <code>dbListTables()</code>.
</p>
<pre class="r"><code>dbListTables(con)</code></pre>
<pre><code>## [1] "ecom"         "sqlite_stat1" "sqlite_stat4"</code></pre>
</section>
<section id="list-fields" class="level4">
<h4 class="anchored" data-anchor-id="list-fields">
List Fields
</h4>
<p>
Time to explore the <code>ecom</code> table in the database. Use <code>dbListFields()</code> to list all the fields in the table.
</p>
<pre class="r"><code>dbListFields(con, "ecom")</code></pre>
<pre><code>## [1] "referrer" "device"   "bouncers" "n_visit"  "n_pages"  "duration"</code></pre>
</section>
</section>
<section id="querying-data" class="level3">
<h3 class="anchored" data-anchor-id="querying-data">
Querying Data
</h3>
<p>
The main objectives of connecting to a database are to:
</p>
<ul>
<li>
query data from the tables already present
</li>
<li>
create new tables
</li>
<li>
overwrite existing tables
</li>
<li>
delete existing tables
</li>
</ul>
<p>
Let us begin with querying data. We can do this in the following ways:
</p>
<ul>
<li>
read an entire table at once
</li>
<li>
read few rows from a table
</li>
<li>
read data in batches
</li>
</ul>
<section id="entire-table" class="level4">
<h4 class="anchored" data-anchor-id="entire-table">
Entire Table
</h4>
<p>
We can read an entire table from a database using <code>dbReadTable()</code>.
</p>
<pre class="r"><code>dbReadTable(con, 'ecom')</code></pre>
<pre><code>##    referrer device bouncers n_visit n_pages duration
## 1    google laptop        1      10       1      693
## 2     yahoo tablet        1       9       1      459
## 3    direct laptop        1       0       1      996
## 4      bing tablet        0       3      18      468
## 5     yahoo mobile        1       9       1      955
## 6     yahoo laptop        0       5       5      135
## 7     yahoo mobile        1      10       1       75
## 8    direct mobile        1      10       1      908
## 9      bing mobile        0       3      19      209
## 10   google mobile        1       6       1      208
## 11   direct laptop        1       9       1      738
## 12   direct tablet        0       6      12      132
## 13   direct mobile        0       9      14      406
## 14    yahoo tablet        0       5       8       80
## 15    yahoo mobile        0       7       1       19
## 16     bing laptop        1       1       1      995
## 17     bing tablet        0       5      16      368
## 18   google tablet        1       7       1      406
## 19   social tablet        0       7      10      290
## 20   social tablet        0       2       1       28</code></pre>
<p>
In some cases, we may not need the entire table but only a specific number of rows. Use <code>dbGetQuery()</code> and supply a SQL statement specifying the number of rows of data to be read from the table. In the below example, we read ten rows of data from the <code>ecom</code> table.
</p>
</section>
<section id="few-rows" class="level4">
<h4 class="anchored" data-anchor-id="few-rows">
Few Rows
</h4>
<pre class="r"><code>dbGetQuery(con, "select * from ecom limit 10")</code></pre>
<pre><code>##    referrer device bouncers n_visit n_pages duration
## 1    google laptop        1      10       1      693
## 2     yahoo tablet        1       9       1      459
## 3    direct laptop        1       0       1      996
## 4      bing tablet        0       3      18      468
## 5     yahoo mobile        1       9       1      955
## 6     yahoo laptop        0       5       5      135
## 7     yahoo mobile        1      10       1       75
## 8    direct mobile        1      10       1      908
## 9      bing mobile        0       3      19      209
## 10   google mobile        1       6       1      208</code></pre>
<p>
In case of very large table, we can read data in batches using <code>dbSendQuery()</code> and <code>dbFetch()</code>. We can mention the number of rows of data to be read while fetching the data using the query generated by <code>dbGetQuery()</code>.
</p>
</section>
<section id="read-data-in-batches" class="level4">
<h4 class="anchored" data-anchor-id="read-data-in-batches">
Read Data in Batches
</h4>
<pre class="r"><code>query &lt;- dbSendQuery(con, 'select * from ecom')
result &lt;- dbFetch(query, n = 15)
result</code></pre>
<pre><code>##    referrer device bouncers n_visit n_pages duration
## 1    google laptop        1      10       1      693
## 2     yahoo tablet        1       9       1      459
## 3    direct laptop        1       0       1      996
## 4      bing tablet        0       3      18      468
## 5     yahoo mobile        1       9       1      955
## 6     yahoo laptop        0       5       5      135
## 7     yahoo mobile        1      10       1       75
## 8    direct mobile        1      10       1      908
## 9      bing mobile        0       3      19      209
## 10   google mobile        1       6       1      208
## 11   direct laptop        1       9       1      738
## 12   direct tablet        0       6      12      132
## 13   direct mobile        0       9      14      406
## 14    yahoo tablet        0       5       8       80
## 15    yahoo mobile        0       7       1       19</code></pre>
</section>
</section>
<section id="query" class="level3">
<h3 class="anchored" data-anchor-id="query">
Query
</h3>
<section id="query-status" class="level4">
<h4 class="anchored" data-anchor-id="query-status">
Query Status
</h4>
<p>
To know the status of a query, use <code>dbHasCompleted()</code>. It is very useful in cases of queries that might take a long time to complete.
</p>
<pre class="r"><code>dbHasCompleted(query)</code></pre>
<pre><code>## [1] FALSE</code></pre>
</section>
<section id="query-info" class="level4">
<h4 class="anchored" data-anchor-id="query-info">
Query Info
</h4>
<p>
<code>dbGetInfo()</code> will return the following:
</p>
<ul>
<li>
the sql staement
</li>
<li>
number of rows fetched
</li>
<li>
number of rows modified/affected
</li>
<li>
status of the query
</li>
</ul>
<pre class="r"><code>dbGetInfo(query)</code></pre>
<pre><code>## $statement
## [1] "select * from ecom"
## 
## $row.count
## [1] 15
## 
## $rows.affected
## [1] 0
## 
## $has.completed
## [1] FALSE</code></pre>
</section>
<section id="latest-query" class="level4">
<h4 class="anchored" data-anchor-id="latest-query">
Latest Query
</h4>
<p>
To get the latest query, use <code>dbGetStatement()</code>.
</p>
<pre class="r"><code>dbGetStatement(query)</code></pre>
<pre><code>## [1] "select * from ecom"</code></pre>
</section>
<section id="rows-fetched" class="level4">
<h4 class="anchored" data-anchor-id="rows-fetched">
Rows Fetched
</h4>
<p>
To check the number of rows of data returned by a query, use <code>dbGetRowCount()</code>.
</p>
<pre class="r"><code>dbGetRowCount(query)</code></pre>
<pre><code>## [1] 15</code></pre>
</section>
<section id="rows-affected" class="level4">
<h4 class="anchored" data-anchor-id="rows-affected">
Rows Affected
</h4>
<p>
To know the number of rows modified or affected in the table, use <code>dbGetRowsAffected()</code>.
</p>
<pre class="r"><code>dbGetRowsAffected(query)</code></pre>
<pre><code>## [1] 0</code></pre>
</section>
<section id="column-info" class="level4">
<h4 class="anchored" data-anchor-id="column-info">
Column Info
</h4>
<p>
To know the name of the columns and their data types, use <code>dbColumnInfo()</code>.
</p>
<pre class="r"><code>dbColumnInfo(query)</code></pre>
<pre><code>##       name      type
## 1 referrer character
## 2   device character
## 3 bouncers   integer
## 4  n_visit    double
## 5  n_pages    double
## 6 duration    double</code></pre>
</section>
</section>
<section id="create-table" class="level3">
<h3 class="anchored" data-anchor-id="create-table">
Create Table
</h3>
<p>
So far we have explored querying data from an existing table. Now, let us turn our attention to creating new tables in the database.
</p>
<section id="introduction-1" class="level4">
<h4 class="anchored" data-anchor-id="introduction-1">
Introduction
</h4>
<p>
To create a new table, use <code>dbWriteTable()</code>. It takes the following 3 arguments:
</p>
<ul>
<li>
connection name
</li>
<li>
name of the new table
</li>
<li>
data for the new table
</li>
</ul>
<pre class="r"><code>x &lt;- 1:10
y &lt;- letters[1:10]
trial &lt;- tibble::tibble(x, y)
dbWriteTable(con, "trial", trial)</code></pre>
<pre><code>## Warning: Closing open result set, pending rows</code></pre>
<p>
Let us check if the new table has been created.
</p>
<pre class="r"><code>dbListTables(con)</code></pre>
<pre><code>## [1] "ecom"         "sqlite_stat1" "sqlite_stat4" "trial"</code></pre>
<pre class="r"><code>dbExistsTable(con, "trial")</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
Let us query data from the new table.
</p>
<pre class="r"><code>dbGetQuery(con, "select * from trial limit 5")</code></pre>
<pre><code>##   x y
## 1 1 a
## 2 2 b
## 3 3 c
## 4 4 d
## 5 5 e</code></pre>
</section>
<section id="overwrite-table" class="level4">
<h4 class="anchored" data-anchor-id="overwrite-table">
Overwrite Table
</h4>
<p>
In some cases, you may want to overwrite the data in an existing table. Use the <code>overwrite</code> argument in <code>dbWriteTable()</code> and set it to <code>TRUE</code>.
</p>
<pre class="r"><code>x &lt;- sample(100, 10)
y &lt;- letters[11:20]
trial2 &lt;- tibble::tibble(x, y)
dbWriteTable(con, "trial", trial2, overwrite = TRUE)</code></pre>
<p>
Let us see if the <strong>trial</strong> table has been overwritten.
</p>
<pre class="r"><code>dbGetQuery(con, "select * from trial limit 5")</code></pre>
<pre><code>##    x y
## 1 48 k
## 2 58 l
## 3 85 m
## 4 99 n
## 5 78 o</code></pre>
</section>
<section id="append-data" class="level4">
<h4 class="anchored" data-anchor-id="append-data">
Append Data
</h4>
<p>
You can append data to an existing table by setting the <code>append</code> argument in <code>dbWriteTable()</code> to <code>TRUE</code>.
</p>
<pre class="r"><code>x &lt;- sample(100, 10)
y &lt;- letters[5:14]
trial3 &lt;- tibble::tibble(x, y)
dbWriteTable(con, "trial", trial3, append = TRUE)</code></pre>
<p>
Let us quickly check if the new data has been appended to the <strong>trial</strong> table.
</p>
<pre class="r"><code>dbReadTable(con, "trial")</code></pre>
<pre><code>##     x y
## 1  48 k
## 2  58 l
## 3  85 m
## 4  99 n
## 5  78 o
## 6   9 p
## 7   6 q
## 8  59 r
## 9  11 s
## 10 38 t
## 11 39 e
## 12 69 f
## 13 43 g
## 14 71 h
## 15 99 i
## 16 56 j
## 17 45 k
## 18 81 l
## 19 93 m
## 20 47 n</code></pre>
<p>
We can also use <code>sqlAppendTable()</code> to append data to an existing table.
</p>
<pre class="r"><code>sqlAppendTable(con, "ecom", head(ecom))</code></pre>
<pre><code>## Warning: Do not rely on the default value of the row.names argument for
## sqlAppendTable(), it will change in the future.</code></pre>
<pre><code>## &lt;SQL&gt; INSERT INTO `ecom`
##   (`referrer`, `device`, `bouncers`, `n_visit`, `n_pages`, `duration`)
## VALUES
##   ('google', 'laptop', TRUE, 10, 1, 693),
##   ('yahoo', 'tablet', TRUE, 9, 1, 459),
##   ('direct', 'laptop', TRUE, 0, 1, 996),
##   ('bing', 'tablet', FALSE, 3, 18, 468),
##   ('yahoo', 'mobile', TRUE, 9, 1, 955),
##   ('yahoo', 'laptop', FALSE, 5, 5, 135)</code></pre>
</section>
</section>
<section id="insert-rows" class="level3">
<h3 class="anchored" data-anchor-id="insert-rows">
Insert Rows
</h3>
<section id="introduction-2" class="level4">
<h4 class="anchored" data-anchor-id="introduction-2">
Introduction
</h4>
<p>
We can insert new rows into existing tables using:
</p>
<ul>
<li>
<code>dbExecute()</code>
</li>
<li>
<code>dbSendStatement()</code>
</li>
</ul>
<p>
Both the function take 2 arguments:
</p>
<ul>
<li>
connection name
</li>
<li>
sql statement
</li>
</ul>
<pre class="r"><code># use dbExecute
dbExecute(con,
  "INSERT into trial (x, y) VALUES (32, 'c'), (45, 'k'), (61, 'h')"
)
## [1] 3

# use dbSendStatement
dbSendStatement(con,
  "INSERT into trial (x, y) VALUES (25, 'm'), (54, 'l'), (16, 'y')"
)
## &lt;SQLiteResult&gt;
##   SQL  INSERT into trial (x, y) VALUES (25, 'm'), (54, 'l'), (16, 'y')
##   ROWS Fetched: 0 [complete]
##        Changed: 3</code></pre>
</section>
<section id="remove-table" class="level4">
<h4 class="anchored" data-anchor-id="remove-table">
Remove Table
</h4>
<p>
If you want to delete/remove a table from the database, use <code>dbRemoveTable()</code>.
</p>
<pre class="r"><code>dbRemoveTable(con, "trial")</code></pre>
<pre><code>## Warning: Closing open result set, pending rows</code></pre>
</section>
</section>
<section id="sqlite-data-type" class="level3">
<h3 class="anchored" data-anchor-id="sqlite-data-type">
SQLite Data Type
</h3>
<p>
If you want to know the data type, use <code>dbDataType()</code>.
</p>
<pre class="r"><code>dbDataType(RSQLite::SQLite(), "a")</code></pre>
<pre><code>## [1] "TEXT"</code></pre>
<pre class="r"><code>dbDataType(RSQLite::SQLite(), 1:5)</code></pre>
<pre><code>## [1] "INTEGER"</code></pre>
<pre class="r"><code>dbDataType(RSQLite::SQLite(), 1.5)</code></pre>
<pre><code>## [1] "REAL"</code></pre>
<section id="close-connection" class="level4">
<h4 class="anchored" data-anchor-id="close-connection">
Close Connection
</h4>
<p>
It is a good practice to close connection to a database when you no longer need to read/write data from/to it. Use <code>dbDisconnect()</code> to close the database connection.
</p>
<pre class="r"><code>dbDisconnect(con)</code></pre>
</section>
</section>
<section id="references" class="level3">
<h3 class="anchored" data-anchor-id="references">
References
</h3>
<ul>
<li>
<a href="https://dbi.r-dbi.org/" class="uri">https://dbi.r-dbi.org/</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data-wrangling</category>
  <guid>https://blog.rsquaredacademy.com/posts/quick-guide-r-sqlite/</guid>
  <pubDate>Tue, 27 Nov 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/database.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Categorical Data Analysis using forcats</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-11-15-working-with-categorical-data-using-forcats.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In this post, we will learn to work with categorical/qualitative data in R using <a href="https://forcats.tidyverse.org">forcats</a>. Let us begin by installing and loading forcats and a set of other pacakges we will be using.
</p>
</section>
<section id="libraries-code" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code">
Libraries &amp; Code
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://forcats.tidyverse.org/index.html">forcats</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
<a href="http://magrittr.tidyverse.org/index.html">magrittr</a>
</li>
<li>
<a href="http://ggplot2.tidyverse.org/index.html">ggplot2</a>
</li>
<li>
<a href="http://tibble.tidyverse.org/index.html">tibbe</a>
</li>
<li>
<a href="http://purrr.tidyverse.org/index.html">purrr</a>
</li>
<li>
and <a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The codes from <a href="https://gist.github.com/aravindhebbali/85fac536f563ae3fd8e2605fd56a7086">here</a>.
</p>
<pre class="r"><code>library(forcats)
library(tibble)
library(magrittr)
library(purrr)
library(dplyr)
library(ggplot2)
library(readr)</code></pre>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<p>
We will use a case study to explore the various features of the forcats package. You can download the data for the case study from <a href="https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web.csv">here</a> or directly import the data using the readr package. We will do the following in this case study:
</p>
<ul>
<li>
compute the frequency of different referrers
</li>
<li>
plot average number of pages browsed for different referrers
</li>
<li>
collapse referrers with low sample size into a single group
</li>
<li>
club traffic from social media websites into a new category
</li>
<li>
group referrers with traffic below a threshold into a single category
</li>
</ul>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>ecom &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web.csv',
    col_types = cols_only(
      referrer = col_factor(levels = c("bing", "direct", "social", "yahoo", "google")),
      n_pages = col_double(), duration = col_double()
    )
  )

ecom</code></pre>
<pre><code>## # A tibble: 1,000 x 3
##    referrer n_pages duration
##    &lt;fct&gt;      &lt;dbl&gt;    &lt;dbl&gt;
##  1 google         1      693
##  2 yahoo          1      459
##  3 direct         1      996
##  4 bing          18      468
##  5 yahoo          1      955
##  6 yahoo          5      135
##  7 yahoo          1       75
##  8 direct         1      908
##  9 bing          19      209
## 10 google         1      208
## # ... with 990 more rows</code></pre>
<p>
Let us extract the <code>referrer</code> column from the above data using <code>use_series</code> and save it in a new variable <code>referrers</code>. Instead of using ecom which is a tibble, we will use <code>referrers</code> which is a vector. We do this to avoid extracting the <code>referrer</code> column from the above data in later examples.
</p>
<pre class="r"><code>referrers &lt;- use_series(ecom, referrer)</code></pre>
</section>
</section>
<section id="tabulate-referrers" class="level2">
<h2 class="anchored" data-anchor-id="tabulate-referrers">
Tabulate Referrers
</h2>
<p>
Let us look at the traffic driven by different referrer types.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_count.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>fct_count(referrers)</code></pre>
<pre><code>## # A tibble: 5 x 2
##   f          n
##   &lt;fct&gt;  &lt;int&gt;
## 1 bing     194
## 2 direct   191
## 3 social   200
## 4 yahoo    207
## 5 google   208</code></pre>
<p>
If you want to sort the output in descending order, use <code>sort</code> and set it to <code>TRUE</code>.
</p>
<pre class="r"><code>fct_count(referrers, sort = TRUE)</code></pre>
<pre><code>## # A tibble: 5 x 2
##   f          n
##   &lt;fct&gt;  &lt;int&gt;
## 1 google   208
## 2 yahoo    207
## 3 social   200
## 4 bing     194
## 5 direct   191</code></pre>
<p>
Use <code>fct_unique</code> to view the categories or levels of the referrer variable.
</p>
<pre class="r"><code>fct_unique(referrers)</code></pre>
<pre><code>## [1] bing   direct social yahoo  google
## Levels: bing direct social yahoo google</code></pre>
</section>
<section id="reorder-referrers" class="level2">
<h2 class="anchored" data-anchor-id="reorder-referrers">
Reorder Referrers
</h2>
<p>
We want to examine the average number of pages visited by each referrer type.
</p>
<pre class="r"><code>refer_summary &lt;- 
  ecom %&gt;%
  group_by(referrer) %&gt;%
  summarise(
    page = mean(n_pages),
    tos = mean(duration),
    n = n()
  )</code></pre>
<pre><code>## `summarise()` ungrouping (override with `.groups` argument)</code></pre>
<pre class="r"><code>refer_summary</code></pre>
<pre><code>## # A tibble: 5 x 4
##   referrer  page   tos     n
## * &lt;fct&gt;    &lt;dbl&gt; &lt;dbl&gt; &lt;int&gt;
## 1 bing      6.13  368.   194
## 2 direct    6.38  358.   191
## 3 social    5.42  355.   200
## 4 yahoo     5.99  336.   207
## 5 google    5.73  360.   208</code></pre>
<p>
Let us plot the average number of pages visited by each referrer type.
</p>
<pre class="r"><code>refer_summary %&gt;%
  ggplot() +
  geom_point(aes(page, referrer))</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/2018-11-15-working-with-categorical-data-using-forcats_files/figure-html/cat10-1.png" width="576" style="display: block; margin: auto;">
</p>
<p>
Use <code>fct_reorder</code> to reorder the referrer types by the average number of pages visited.
</p>
<pre class="r"><code>refer_summary %&gt;%
  ggplot() +
  geom_point(aes(page, fct_reorder(referrer, page)))</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/2018-11-15-working-with-categorical-data-using-forcats_files/figure-html/cat11-1.png" width="576" style="display: block; margin: auto;">
</p>
</section>
<section id="plot-referrer-frequency-descending-order" class="level2">
<h2 class="anchored" data-anchor-id="plot-referrer-frequency-descending-order">
Plot Referrer Frequency (Descending Order)
</h2>
<p>
Since we want to plot the referrers in descending order of frequency, we will use <code>fct_infreq()</code> to reorder by frequency.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_infreq.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>referrers %&gt;%
  fct_infreq() %&gt;%
  fct_unique()</code></pre>
<pre><code>## [1] google yahoo  social bing   direct
## Levels: google yahoo social bing direct</code></pre>
<p>
Now that we know how to reorder categories/levels by frequency, let us reorder the referrers by frequency and plot them.
</p>
<pre class="r"><code>ecom %&gt;%
  mutate(
    ref = referrer %&gt;% 
      fct_infreq()
  ) %&gt;%
  ggplot(aes(ref)) +
  geom_bar()</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/2018-11-15-working-with-categorical-data-using-forcats_files/figure-html/cat4-1.png" width="576" style="display: block; margin: auto;">
</p>
</section>
<section id="plot-referrer-frequency-ascending-order" class="level2">
<h2 class="anchored" data-anchor-id="plot-referrer-frequency-ascending-order">
Plot Referrer Frequency (Ascending Order)
</h2>
<p>
Let us look at the categories of the referrer variable.
</p>
<pre class="r"><code>fct_unique(referrers)</code></pre>
<pre><code>## [1] bing   direct social yahoo  google
## Levels: bing direct social yahoo google</code></pre>
<p>
Since we want to plot the referrers in ascending order of frequency, we will use <code>fct_rev()</code> to reverse the order.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_rev.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>referrers %&gt;%
  fct_rev() %&gt;%
  fct_unique()</code></pre>
<pre><code>## [1] google yahoo  social direct bing  
## Levels: google yahoo social direct bing</code></pre>
<p>
Let us reorder the referrers by frequency first and then reverse the order before plotting their frequencies.
</p>
<pre class="r"><code>ecom %&gt;%
  mutate(
    ref = referrer %&gt;% 
      fct_infreq() %&gt;% 
      fct_rev()
  ) %&gt;%
  ggplot(aes(ref)) +
  geom_bar()</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/2018-11-15-working-with-categorical-data-using-forcats_files/figure-html/cat5-1.png" width="576" style="display: block; margin: auto;">
</p>
</section>
<section id="case-study-2" class="level2">
<h2 class="anchored" data-anchor-id="case-study-2">
Case Study 2
</h2>
<p>
In this case study, we will learn to:
</p>
<ul>
<li>
combine categories
</li>
<li>
recategorize
</li>
</ul>
<p>
The data set we will use has just one column <code>traffics</code> i.e.&nbsp;the source of traffic for a imaginary website.
</p>
<section id="data-1" class="level3">
<h3 class="anchored" data-anchor-id="data-1">
Data
</h3>
<pre class="r"><code>traffic &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web_traffic.csv',
    col_types = list(
      col_factor(levels = c("affiliates", "bing", "direct", "facebook", 
        "yahoo", "google", "instagram", "twitter", "unknown")
      )
    )
  )

traffic</code></pre>
<pre><code>## # A tibble: 48,232 x 1
##    traffics
##    &lt;fct&gt;   
##  1 google  
##  2 google  
##  3 google  
##  4 google  
##  5 google  
##  6 google  
##  7 google  
##  8 google  
##  9 google  
## 10 google  
## # ... with 48,222 more rows</code></pre>
<p>
Let us extract the <code>traffics</code> column from the above data using <code>use_series</code> and save it in a new variable <code>traffics</code>. Instead of using traffic which is a tibble, we will use <code>traffics</code> which is a vector. We do this to avoid extracting the <code>traffics</code> column from the above data in all the examples shown below.
</p>
<pre class="r"><code>traffics &lt;- use_series(traffic, traffics)</code></pre>
</section>
</section>
<section id="tabulate-referrer" class="level2">
<h2 class="anchored" data-anchor-id="tabulate-referrer">
Tabulate Referrer
</h2>
<p>
Let us compute the traffic driven by different referrers using <code>fct_count</code>.
</p>
<pre class="r"><code>fct_count(traffics)   </code></pre>
<pre><code>## # A tibble: 9 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 affiliates  7641
## 2 bing        5893
## 3 direct      1350
## 4 facebook    8135
## 5 yahoo       4899
## 6 google      9229
## 7 instagram   3907
## 8 twitter     4521
## 9 unknown     2657</code></pre>
</section>
<section id="collapse-referrer-categories" class="level2">
<h2 class="anchored" data-anchor-id="collapse-referrer-categories">
Collapse Referrer Categories
</h2>
<p>
We want to group some of the referrers into 2 categories:
</p>
<ul>
<li>
social
</li>
<li>
search
</li>
</ul>
<p>
To group categories/levels, we will use <code>fct_collapse()</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_collapse.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>traffics %&gt;%
  fct_collapse(
    social = c("facebook", "twitter", "instagram"),
    search = c("google", "bing", "yahoo")
  ) %&gt;% 
  fct_count()</code></pre>
<pre><code>## # A tibble: 5 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 affiliates  7641
## 2 search     20021
## 3 direct      1350
## 4 social     16563
## 5 unknown     2657</code></pre>
<p>
The above result can be achieved using <code>fct_recode()</code> as shown below:
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_recode.png" width="100%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_recode(traffics, 
  search = "bing", 
  search = "yahoo", 
  search = "google",
  social = "facebook", 
  social = "twitter", 
  social = "instagram") %&gt;%
  levels()</code></pre>
<pre><code>## [1] "affiliates" "search"     "direct"     "social"     "unknown"</code></pre>
</section>
<section id="lump-infrequent-referrer-types" class="level2">
<h2 class="anchored" data-anchor-id="lump-infrequent-referrer-types">
Lump Infrequent Referrer Types
</h2>
<p>
Let us group together referrer types that drive low traffic to the website. Use <code>fct_lump()</code> to lump together categories.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_lump_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>fct_count(traffics)
## # A tibble: 9 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 affiliates  7641
## 2 bing        5893
## 3 direct      1350
## 4 facebook    8135
## 5 yahoo       4899
## 6 google      9229
## 7 instagram   3907
## 8 twitter     4521
## 9 unknown     2657

traffics %&gt;% 
  fct_lump() %&gt;% 
  table()
## .
## affiliates       bing   facebook      yahoo     google  instagram    twitter 
##       7641       5893       8135       4899       9229       3907       4521 
##    unknown      Other 
##       2657       1350</code></pre>
</section>
<section id="retain-top-3-referrers" class="level2">
<h2 class="anchored" data-anchor-id="retain-top-3-referrers">
Retain top 3 referrers
</h2>
<p>
We want to retain the top 3 referrers and combine the rest of them into a single category.
</p>
<pre><code>## # A tibble: 9 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 google      9229
## 2 facebook    8135
## 3 affiliates  7641
## 4 bing        5893
## 5 yahoo       4899
## 6 twitter     4521
## 7 instagram   3907
## 8 unknown     2657
## 9 direct      1350</code></pre>
<p>
Use <code>fct_lump()</code> and set the argument <code>n</code> to <code>3</code> indicating we want to retain top 3 categories and combine the rest.
</p>
<pre class="r"><code>traffics %&gt;% 
  fct_lump(n = 3) %&gt;% 
  table()</code></pre>
<pre><code>## .
## affiliates   facebook     google      Other 
##       7641       8135       9229      23227</code></pre>
</section>
<section id="lump-referrer-types-with-less-than-10-traffic" class="level2">
<h2 class="anchored" data-anchor-id="lump-referrer-types-with-less-than-10-traffic">
Lump Referrer Types with less than 10% traffic
</h2>
<p>
Let us combine referrers that drive less than 10% traffic to the website.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_lump_2.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre><code>## # A tibble: 9 x 3
##   f              n percent
##   &lt;fct&gt;      &lt;int&gt;   &lt;dbl&gt;
## 1 affiliates  7641   15.8 
## 2 bing        5893   12.2 
## 3 direct      1350    2.8 
## 4 facebook    8135   16.9 
## 5 yahoo       4899   10.2 
## 6 google      9229   19.1 
## 7 instagram   3907    8.1 
## 8 twitter     4521    9.37
## 9 unknown     2657    5.51</code></pre>
<p>
Since we are looking at proportion of traffic driven to the website and not the actual numbers, we use the <code>prop</code> argument and set it to <code>0.1</code>, indicating that we want to retain only those categories which have a proportion of more than 10% and combine the rest.
</p>
<pre class="r"><code>traffics %&gt;%
  fct_lump(prop = 0.1) %&gt;% 
  table()</code></pre>
<pre><code>## .
## affiliates       bing   facebook      yahoo     google      Other 
##       7641       5893       8135       4899       9229      12435</code></pre>
</section>
<section id="retain-3-referrer-types-with-lowest-traffic" class="level2">
<h2 class="anchored" data-anchor-id="retain-3-referrer-types-with-lowest-traffic">
Retain 3 Referrer Types with lowest traffic
</h2>
<p>
What if we want to retain 3 referrers which drive the lowest traffic to the website and combine the rest?
</p>
<pre><code>## # A tibble: 9 x 2
##   f              n
##   &lt;fct&gt;      &lt;int&gt;
## 1 direct      1350
## 2 unknown     2657
## 3 instagram   3907
## 4 twitter     4521
## 5 yahoo       4899
## 6 bing        5893
## 7 affiliates  7641
## 8 facebook    8135
## 9 google      9229</code></pre>
<p>
We will still use the <code>n</code> argument but instead of specifying <code>3</code>, we now specify <code>-3</code>.
</p>
<pre class="r"><code>traffics %&gt;% 
  fct_lump(n = -3) %&gt;% 
  table()</code></pre>
<pre><code>## .
##    direct instagram   unknown     Other 
##      1350      3907      2657     40318</code></pre>
</section>
<section id="retain-3-referrer-types-with-less-than-10-traffic" class="level2">
<h2 class="anchored" data-anchor-id="retain-3-referrer-types-with-less-than-10-traffic">
Retain 3 Referrer Types with less than 10% traffic
</h2>
<p>
Let us see how to retain referrers that drive less than 10 % traffic to the website and combine the rest into a single group.
</p>
<pre><code>## # A tibble: 9 x 3
##   f              n percent
##   &lt;fct&gt;      &lt;int&gt;   &lt;dbl&gt;
## 1 affiliates  7641   15.8 
## 2 bing        5893   12.2 
## 3 direct      1350    2.8 
## 4 facebook    8135   16.9 
## 5 yahoo       4899   10.2 
## 6 google      9229   19.1 
## 7 instagram   3907    8.1 
## 8 twitter     4521    9.37
## 9 unknown     2657    5.51</code></pre>
<p>
Instead of setting <code>prop</code> to <code>0.1</code>, we will set it to <code>-0.1</code>.
</p>
<pre class="r"><code>traffics %&gt;% 
  fct_lump(prop = -0.1) %&gt;% 
  table()</code></pre>
<pre><code>## .
##    direct instagram   twitter   unknown     Other 
##      1350      3907      4521      2657     35797</code></pre>
</section>
<section id="replace-levels" class="level2">
<h2 class="anchored" data-anchor-id="replace-levels">
Replace Levels
</h2>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_others_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
Let us assume we want to retain a couple of important categories and group the rest into a single category. In the below example, we retain <em>google</em> and <em>yahoo</em> while grouping the rest as others using <code>fct_other()</code>.
</p>
<pre class="r"><code>fct_other(traffics, keep = c("google", "yahoo")) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "yahoo"  "google" "Other"</code></pre>
</section>
<section id="drop-levels" class="level2">
<h2 class="anchored" data-anchor-id="drop-levels">
Drop Levels
</h2>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_others_2.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
What if you want to drop a couple of categories instead of grouping them? Use the <code>drop</code> argument in <code>fct_other()</code> and specify the categories to be dropped. In the below example, we drop the following referrer categories:
</p>
<ul>
<li>
instagram
</li>
<li>
twitter
</li>
</ul>
<pre class="r"><code>fct_other(traffics, drop = c("instagram", "twitter")) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "affiliates" "bing"       "direct"     "facebook"   "yahoo"     
## [6] "google"     "unknown"    "Other"</code></pre>
</section>
<section id="reorder-levels" class="level2">
<h2 class="anchored" data-anchor-id="reorder-levels">
Reorder Levels
</h2>
<p>
The categories can be reordered using <code>fct_relevel()</code>. In the above example, we reorder the categories to ensure <em>google</em> appears first. Similarly in the below example, we reorder the levels to ensure <em>twitter</em> appears first irrespective of its frequency or order of appearance in the data.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_1.png" width="100%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_relevel(traffics, "twitter") %&gt;%
  levels()</code></pre>
<pre><code>## [1] "twitter"    "affiliates" "bing"       "direct"     "facebook"  
## [6] "yahoo"      "google"     "instagram"  "unknown"</code></pre>
<p>
If the category needs to appear at a particular position, use the <code>after</code> argument and specify the position after which it should appear. For example, if <em>google</em> should be the third category, we would specify <code>after = 2</code> i.e. <em>google</em> should come after the 2nd position (i.e.&nbsp;third position).
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_2.png" width="100%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_relevel(traffics, "google", after = 2) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "affiliates" "bing"       "google"     "direct"     "facebook"  
## [6] "yahoo"      "instagram"  "twitter"    "unknown"</code></pre>
<p>
If the category should appear last, supply the value <code>Inf</code> (infinity) to the <code>after</code> argument as shown below.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_relevel_3.png" width="100%" style="display: block; margin: auto;">
</p>
<pre class="r"><code>fct_relevel(traffics, "facebook", after = Inf) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "affiliates" "bing"       "direct"     "yahoo"      "google"    
## [6] "instagram"  "twitter"    "unknown"    "facebook"</code></pre>
</section>
<section id="case-study-3" class="level2">
<h2 class="anchored" data-anchor-id="case-study-3">
Case Study 3
</h2>
<p>
In this case study, we deal with categorical data which is ordered and cyclical. It contains response to an imaginary survey.
</p>
<section id="data-2" class="level3">
<h3 class="anchored" data-anchor-id="data-2">
Data
</h3>
<pre class="r"><code>response_data &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/response.csv',
    col_types = list(col_factor(levels = c("like", "like somewhat", "neutral", 
      "dislike somewhat", "dislike"), ordered = TRUE)
    )
  )</code></pre>
<p>
Since we will be using only one column from the above data set, let us extract it using <code>use_series()</code> and save it as <code>responses</code>.
</p>
<pre class="r"><code>responses &lt;- use_series(response_data, response)
levels(responses)</code></pre>
<pre><code>## [1] "like"             "like somewhat"    "neutral"          "dislike somewhat"
## [5] "dislike"</code></pre>
</section>
</section>
<section id="shift-levels" class="level2">
<h2 class="anchored" data-anchor-id="shift-levels">
Shift Levels
</h2>
<p>
To shift the levels, we use <code>fct_shift()</code>. Use the <code>n</code> argument to indicate the direction of the shift. If <code>n</code> is positive, the levels are shifted to the left else to the right. In the below example, we shift the levels to the left by 2 positions.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_shift_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>fct_shift(responses, 2) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "neutral"          "dislike somewhat" "dislike"          "like"            
## [5] "like somewhat"</code></pre>
<p>
To shift the levels to the right, supply a negative value to the <code>n</code> argument in <code>fct_shift()</code>. In the below example, we shift the levels to the right by 2 positions.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/fct_shift_2.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>fct_shift(responses, -2) %&gt;%
  levels()</code></pre>
<pre><code>## [1] "dislike somewhat" "dislike"          "like"             "like somewhat"   
## [5] "neutral"</code></pre>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://forcats.tidyverse.org/index.html" class="uri">https://forcats.tidyverse.org/index.html</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/factors.html" class="uri">http://r4ds.had.co.nz/factors.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>forcats</category>
  <guid>https://blog.rsquaredacademy.com/posts/working-with-categorical-data-using-forcats/</guid>
  <pubDate>Thu, 15 Nov 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/forcats.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Working with Date and Time in R</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/working-with-dates-in-r/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-11-03-working-with-dates-in-r.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In this post, we will learn to work with date/time data in R using <a href="http://lubridate.tidyverse.org/">lubridate</a>, an R package that makes it easy to work with dates and time. Let us begin by installing and loading the pacakge.
</p>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://lubridate.tidyverse.org/index.html">lubridate</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
<a href="http://magrittr.tidyverse.org/index.html">magrittr</a>
</li>
<li>
<a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The data sets can be downloaded from <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes from <a href="https://gist.github.com/aravindhebbali/7758b86c2bc13ff1e5d88d9d1c204f8c">here</a>.
</p>
<pre class="r"><code>library(lubridate)
library(dplyr)
library(magrittr)
library(readr)</code></pre>
</section>
<section id="quick-intro" class="level2">
<h2 class="anchored" data-anchor-id="quick-intro">
Quick Intro
</h2>
<section id="origin" class="level4">
<h4 class="anchored" data-anchor-id="origin">
Origin
</h4>
<p>
Let us look at the origin for the numbering system used for date and time calculations in R.
</p>
<pre class="r"><code>origin</code></pre>
<pre><code>## [1] "1970-01-01 UTC"</code></pre>
</section>
<section id="current-datetime" class="level4">
<h4 class="anchored" data-anchor-id="current-datetime">
Current Date/Time
</h4>
<p>
Next, let us check out the current date, time and whether it occurs in the am or pm. <code>now()</code> returns the date time as well as the time zone whereas <code>today()</code> will return only the current date. <code>am()</code> and <code>pm()</code> return <code>TRUE</code> or <code>FALSE</code>.
</p>
<pre class="r"><code>now()</code></pre>
<pre><code>## [1] "2020-06-10 18:52:25 IST"</code></pre>
<pre class="r"><code>today()</code></pre>
<pre><code>## [1] "2020-06-10"</code></pre>
<pre class="r"><code>am(now())  </code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>pm(now())</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>transact &lt;- read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/transact.csv')</code></pre>
<pre><code>## # A tibble: 2,466 x 3
##    Invoice    Due        Payment   
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;    
##  1 2013-01-02 2013-02-01 2013-01-15
##  2 2013-01-26 2013-02-25 2013-03-03
##  3 2013-07-03 2013-08-02 2013-07-08
##  4 2013-02-10 2013-03-12 2013-03-17
##  5 2012-10-25 2012-11-24 2012-11-28
##  6 2012-01-27 2012-02-26 2012-02-22
##  7 2013-08-13 2013-09-12 2013-09-09
##  8 2012-12-16 2013-01-15 2013-01-12
##  9 2012-05-14 2012-06-13 2012-07-01
## 10 2013-07-01 2013-07-31 2013-07-26
## # ... with 2,456 more rows</code></pre>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<p>
The data set has 3 columns. All the dates are in the format (yyyy-mm-dd).
</p>
<ul>
<li>
Invoice: invoice date
</li>
<li>
Due: due date
</li>
<li>
Payment: payment date
</li>
</ul>
<p>
We will use the functions in the lubridate package to answer a few questions we have about the transact data.
</p>
<ul>
<li>
extract date, month and year from Due
</li>
<li>
compute the number of days to settle invoice
</li>
<li>
compute days over due
</li>
<li>
check if due year is a leap year
</li>
<li>
check when due day in february is 29, whether it is a leap year
</li>
<li>
how many invoices were settled within due date
</li>
<li>
how many invoices are due in each quarter
</li>
<li>
what is the average duration between invoice date and payment date
</li>
</ul>
</section>
</section>
<section id="extract-date-month-year-from-due-date" class="level2">
<h2 class="anchored" data-anchor-id="extract-date-month-year-from-due-date">
Extract Date, Month &amp; Year from Due Date
</h2>
<p>
The first thing we will learn is to extract the date, month and year.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/day_week_month.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>this_day &lt;- as_date('2017-03-23')
day(this_day)</code></pre>
<pre><code>## [1] 23</code></pre>
<pre class="r"><code>month(this_day)</code></pre>
<pre><code>## [1] 3</code></pre>
<pre class="r"><code>year(this_day)</code></pre>
<pre><code>## [1] 2017</code></pre>
<p>
Let us now extract the date, month and year from the <code>Due</code> column.
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    due_day   = day(Due),
    due_month = month(Due),
    due_year  = year(Due)
  )</code></pre>
<pre><code>## # A tibble: 2,466 x 6
##    Invoice    Due        Payment    due_day due_month due_year
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;       &lt;int&gt;     &lt;dbl&gt;    &lt;dbl&gt;
##  1 2013-01-02 2013-02-01 2013-01-15       1         2     2013
##  2 2013-01-26 2013-02-25 2013-03-03      25         2     2013
##  3 2013-07-03 2013-08-02 2013-07-08       2         8     2013
##  4 2013-02-10 2013-03-12 2013-03-17      12         3     2013
##  5 2012-10-25 2012-11-24 2012-11-28      24        11     2012
##  6 2012-01-27 2012-02-26 2012-02-22      26         2     2012
##  7 2013-08-13 2013-09-12 2013-09-09      12         9     2013
##  8 2012-12-16 2013-01-15 2013-01-12      15         1     2013
##  9 2012-05-14 2012-06-13 2012-07-01      13         6     2012
## 10 2013-07-01 2013-07-31 2013-07-26      31         7     2013
## # ... with 2,456 more rows</code></pre>
</section>
<section id="compute-days-to-settle-invoice" class="level2">
<h2 class="anchored" data-anchor-id="compute-days-to-settle-invoice">
Compute days to settle invoice
</h2>
<p>
Time to do some arithmetic with the dates. Let us calculate the duration of a course by subtracting the course start date from the course end date.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/course_duration.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>course_start    &lt;- as_date('2017-04-12')
course_end      &lt;- as_date('2017-04-21')
course_duration &lt;- course_end - course_start
course_duration
## Time difference of 9 days</code></pre>
<p>
Let us estimate the number of days to settle the invoice by subtracting the date of invoice from the date of payment.
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    days_to_pay = Payment - Invoice
  )</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    days_to_pay
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;     &lt;drtn&gt;     
##  1 2013-01-02 2013-02-01 2013-01-15 13 days    
##  2 2013-01-26 2013-02-25 2013-03-03 36 days    
##  3 2013-07-03 2013-08-02 2013-07-08  5 days    
##  4 2013-02-10 2013-03-12 2013-03-17 35 days    
##  5 2012-10-25 2012-11-24 2012-11-28 34 days    
##  6 2012-01-27 2012-02-26 2012-02-22 26 days    
##  7 2013-08-13 2013-09-12 2013-09-09 27 days    
##  8 2012-12-16 2013-01-15 2013-01-12 27 days    
##  9 2012-05-14 2012-06-13 2012-07-01 48 days    
## 10 2013-07-01 2013-07-31 2013-07-26 25 days    
## # ... with 2,456 more rows</code></pre>
</section>
<section id="compute-days-over-due" class="level2">
<h2 class="anchored" data-anchor-id="compute-days-over-due">
Compute days over due
</h2>
<p>
How many of the invoices were settled post the due date? We can find this by:
</p>
<ul>
<li>
subtracting the due date from the payment date
</li>
<li>
counting the number of rows where delay &lt; 0
</li>
</ul>
<pre class="r"><code>transact %&gt;%
  mutate(
    delay = Due - Payment
  ) %&gt;%
  filter(delay &lt; 0) %&gt;%
  tally()</code></pre>
<pre><code>## # A tibble: 1 x 1
##       n
##   &lt;int&gt;
## 1   877</code></pre>
</section>
<section id="is-due-year-a-leap-year" class="level2">
<h2 class="anchored" data-anchor-id="is-due-year-a-leap-year">
Is due year a leap year?
</h2>
<p>
Just for fun, let us check if the due year happens to be a leap year.
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    is_leap = leap_year(Due)
  )</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    is_leap
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;     &lt;lgl&gt;  
##  1 2013-01-02 2013-02-01 2013-01-15 FALSE  
##  2 2013-01-26 2013-02-25 2013-03-03 FALSE  
##  3 2013-07-03 2013-08-02 2013-07-08 FALSE  
##  4 2013-02-10 2013-03-12 2013-03-17 FALSE  
##  5 2012-10-25 2012-11-24 2012-11-28 TRUE   
##  6 2012-01-27 2012-02-26 2012-02-22 TRUE   
##  7 2013-08-13 2013-09-12 2013-09-09 FALSE  
##  8 2012-12-16 2013-01-15 2013-01-12 FALSE  
##  9 2012-05-14 2012-06-13 2012-07-01 TRUE   
## 10 2013-07-01 2013-07-31 2013-07-26 FALSE  
## # ... with 2,456 more rows</code></pre>
</section>
<section id="if-due-day-is-february-29-is-it-a-leap-year" class="level2">
<h2 class="anchored" data-anchor-id="if-due-day-is-february-29-is-it-a-leap-year">
If due day is February 29, is it a leap year?
</h2>
<p>
Let us do some data sanitization. If the due day happens to be February 29, let us ensure that the due year is a leap year. Below are the steps to check if the due year is a leap year:
</p>
<ul>
<li>
we will extract the following from the due date:
<ul>
<li>
day
</li>
<li>
month
</li>
<li>
year
</li>
</ul>
</li>
<li>
we will then create a new column <code>is_leap</code> which will have be set to <code>TRUE</code> if the year is a leap year else it will be set to <code>FALSE</code>
</li>
<li>
filter all the payments due on 29th Feb
</li>
<li>
select the following columns:
<ul>
<li>
<code>Due</code>
</li>
<li>
<code>is_leap</code>
</li>
</ul>
</li>
</ul>
<pre class="r"><code>transact %&gt;%
  mutate(
    due_day   = day(Due),
    due_month = month(Due),
    due_year  = year(Due),
    is_leap   = leap_year(due_year)
  ) %&gt;%
  filter(due_month == 2 &amp; due_day == 29) %&gt;%
  select(Due, is_leap) </code></pre>
<pre><code>## # A tibble: 4 x 2
##   Due        is_leap
##   &lt;date&gt;     &lt;lgl&gt;  
## 1 2012-02-29 TRUE   
## 2 2012-02-29 TRUE   
## 3 2012-02-29 TRUE   
## 4 2012-02-29 TRUE</code></pre>
</section>
<section id="shift-date" class="level2">
<h2 class="anchored" data-anchor-id="shift-date">
Shift Date
</h2>
<p>
Time to shift some dates. We can shift a date by days, weeks or months. Let us shift the course start date by:
</p>
<ul>
<li>
2 days
</li>
<li>
3 weeks
</li>
<li>
1 year
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/shift_dates.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>course_start + days(2)
## [1] "2017-04-14"
course_start + weeks(3)
## [1] "2017-05-03"
course_start + years(1)
## [1] "2018-04-12"</code></pre>
</section>
<section id="interval" class="level2">
<h2 class="anchored" data-anchor-id="interval">
Interval
</h2>
<p>
Let us calculate the duration of the course using <code>interval</code>. If you observe carefully, the result is not the duration in days but an object of class <code>interval</code>. Now let us learn how we can use intervals.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/course_interval.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>interval(course_start, course_end)</code></pre>
<pre><code>## [1] 2017-04-12 UTC--2017-04-21 UTC</code></pre>
</section>
<section id="intervals-overlap" class="level2">
<h2 class="anchored" data-anchor-id="intervals-overlap">
Intervals Overlap
</h2>
<p>
Let us say you are planning a vacation and want to check if the vacation dates overlap with the course dates. You can do this by:
</p>
<ul>
<li>
creating vacation and course intervals
</li>
<li>
use <code>int_overlaps()</code> to check if two intervals overlap. It returns <code>TRUE</code> if the intervals overlap else <code>FALSE</code>.
</li>
</ul>
<p>
Let us use the vacation start and end dates to create <code>vacation_interval</code> and then check if it overlaps with <code>course_interval</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/interval_overlap.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>vacation_start    &lt;- as_date('2017-04-19')
vacation_end      &lt;- as_date('2017-04-25')
course_interval   &lt;- interval(course_start, course_end)
vacation_interval &lt;- interval(vacation_start, vacation_end)
int_overlaps(course_interval, vacation_interval)
## [1] TRUE</code></pre>
</section>
<section id="how-many-invoices-were-settled-within-due-date" class="level2">
<h2 class="anchored" data-anchor-id="how-many-invoices-were-settled-within-due-date">
How many invoices were settled within due date?
</h2>
<p>
Let us use intervals to count the number of invoices that were settled within the due date. To do this, we will:
</p>
<ul>
<li>
create an interval for the invoice and due date
</li>
<li>
create a new column <code>due_next</code> by incrementing the due date by 1 day
</li>
<li>
another interval for <code>due_next</code> and the payment date
</li>
<li>
if the intervals overlap, the payment was made within the due date
</li>
</ul>
<pre class="r"><code>transact %&gt;%
  mutate(
    inv_due_interval = interval(Invoice, Due),
    due_next         = Due + days(1),
    due_pay_interval = interval(due_next, Payment),
    overlaps         = int_overlaps(inv_due_interval, due_pay_interval)
  ) %&gt;%
  select(Invoice, Due, Payment, overlaps)</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    overlaps
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;     &lt;lgl&gt;   
##  1 2013-01-02 2013-02-01 2013-01-15 TRUE    
##  2 2013-01-26 2013-02-25 2013-03-03 FALSE   
##  3 2013-07-03 2013-08-02 2013-07-08 TRUE    
##  4 2013-02-10 2013-03-12 2013-03-17 FALSE   
##  5 2012-10-25 2012-11-24 2012-11-28 FALSE   
##  6 2012-01-27 2012-02-26 2012-02-22 TRUE    
##  7 2013-08-13 2013-09-12 2013-09-09 TRUE    
##  8 2012-12-16 2013-01-15 2013-01-12 TRUE    
##  9 2012-05-14 2012-06-13 2012-07-01 FALSE   
## 10 2013-07-01 2013-07-31 2013-07-26 TRUE    
## # ... with 2,456 more rows</code></pre>
<p>
Below we show another method to count the number of invoices paid within the due date. Instead of using <code>days</code> to change the due date, we use <code>int_shift</code> to shift it by 1 day.
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    inv_due_interval = interval(Invoice, Due),
    due_pay_interval = interval(Due, Payment),  
    due_pay_next     = int_shift(due_pay_interval, by = days(1)),
    overlaps         = int_overlaps(inv_due_interval, due_pay_next)
  ) %&gt;%
  select(Invoice, Due, Payment, overlaps)</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    overlaps
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;     &lt;lgl&gt;   
##  1 2013-01-02 2013-02-01 2013-01-15 TRUE    
##  2 2013-01-26 2013-02-25 2013-03-03 FALSE   
##  3 2013-07-03 2013-08-02 2013-07-08 TRUE    
##  4 2013-02-10 2013-03-12 2013-03-17 FALSE   
##  5 2012-10-25 2012-11-24 2012-11-28 FALSE   
##  6 2012-01-27 2012-02-26 2012-02-22 TRUE    
##  7 2013-08-13 2013-09-12 2013-09-09 TRUE    
##  8 2012-12-16 2013-01-15 2013-01-12 TRUE    
##  9 2012-05-14 2012-06-13 2012-07-01 FALSE   
## 10 2013-07-01 2013-07-31 2013-07-26 TRUE    
## # ... with 2,456 more rows</code></pre>
<p>
You might be thinking why we incremented the due date by a day before creating the interval between the due day and the payment day. If we do not increment, both the intervals will share a common date i.e.&nbsp;the due date and they will always overlap as shown below:
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    inv_due_interval = interval(Invoice, Due),
    due_pay_interval = interval(Due, Payment),
    overlaps         = int_overlaps(inv_due_interval, due_pay_interval)
  ) %&gt;%
  select(Invoice, Due, Payment, overlaps)</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    overlaps
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;     &lt;lgl&gt;   
##  1 2013-01-02 2013-02-01 2013-01-15 TRUE    
##  2 2013-01-26 2013-02-25 2013-03-03 TRUE    
##  3 2013-07-03 2013-08-02 2013-07-08 TRUE    
##  4 2013-02-10 2013-03-12 2013-03-17 TRUE    
##  5 2012-10-25 2012-11-24 2012-11-28 TRUE    
##  6 2012-01-27 2012-02-26 2012-02-22 TRUE    
##  7 2013-08-13 2013-09-12 2013-09-09 TRUE    
##  8 2012-12-16 2013-01-15 2013-01-12 TRUE    
##  9 2012-05-14 2012-06-13 2012-07-01 TRUE    
## 10 2013-07-01 2013-07-31 2013-07-26 TRUE    
## # ... with 2,456 more rows</code></pre>
</section>
<section id="shift-interval" class="level2">
<h2 class="anchored" data-anchor-id="shift-interval">
Shift Interval
</h2>
<p>
Intervals can be shifted too. In the below example, we shift the course interval by:
</p>
<ul>
<li>
1 day
</li>
<li>
3 weeks
</li>
<li>
1 year
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/shift_interval.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>course_interval &lt;- interval(course_start, course_end)
int_shift(course_interval, by = days(1))
## [1] 2017-04-13 UTC--2017-04-22 UTC
int_shift(course_interval, by = weeks(3))
## [1] 2017-05-03 UTC--2017-05-12 UTC
int_shift(course_interval, by = years(1))
## [1] 2018-04-12 UTC--2018-04-21 UTC</code></pre>
</section>
<section id="within" class="level2">
<h2 class="anchored" data-anchor-id="within">
Within
</h2>
<p>
Let us assume that we have to attend a conference in April 2017. Does it occur during the course duration? We can answer this using <code>%within%</code> which will return <code>TRUE</code> if a date falls within an interval.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/within.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>conference &lt;- as_date('2017-04-15')
conference %within% course_interval
## [1] TRUE</code></pre>
<section id="how-many-invoices-were-settled-within-due-date-1" class="level4">
<h4 class="anchored" data-anchor-id="how-many-invoices-were-settled-within-due-date-1">
How many invoices were settled within due date?
</h4>
<p>
Let us use <code>%within%</code> to count the number of invoices that were settled within the due date. We will do this by:
</p>
<ul>
<li>
creating an interval for the invoice and due date
</li>
<li>
check if the payment date falls within the above interval
</li>
</ul>
<pre class="r"><code>transact %&gt;%
  mutate(
    inv_due_interval = interval(Invoice, Due),
    overlaps         = Payment %within% inv_due_interval
  ) %&gt;%
  select(Due, Payment, overlaps)</code></pre>
<pre><code>## # A tibble: 2,466 x 3
##    Due        Payment    overlaps
##    &lt;date&gt;     &lt;date&gt;     &lt;lgl&gt;   
##  1 2013-02-01 2013-01-15 TRUE    
##  2 2013-02-25 2013-03-03 FALSE   
##  3 2013-08-02 2013-07-08 TRUE    
##  4 2013-03-12 2013-03-17 FALSE   
##  5 2012-11-24 2012-11-28 FALSE   
##  6 2012-02-26 2012-02-22 TRUE    
##  7 2013-09-12 2013-09-09 TRUE    
##  8 2013-01-15 2013-01-12 TRUE    
##  9 2012-06-13 2012-07-01 FALSE   
## 10 2013-07-31 2013-07-26 TRUE    
## # ... with 2,456 more rows</code></pre>
</section>
</section>
<section id="quarter" class="level2">
<h2 class="anchored" data-anchor-id="quarter">
Quarter
</h2>
<p>
Let us check the quarter and the semester in which the course starts.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/quarter_semester.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>course_start
## [1] "2017-04-12"
quarter(course_start)
## [1] 2
quarter(course_start, with_year = TRUE)
## [1] 2017.2
semester(course_start)  
## [1] 1</code></pre>
<p>
Let us count the invoices due for each quarter.
</p>
<pre class="r"><code>transact %&gt;%
  mutate(
    quarter_due = quarter(Due)
  ) %&gt;%
  count(quarter_due)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   quarter_due     n
## *       &lt;int&gt; &lt;int&gt;
## 1           1   521
## 2           2   661
## 3           3   618
## 4           4   666</code></pre>
<pre class="r"><code>transact %&gt;%
  mutate(
    Quarter = quarter(Due, with_year = TRUE)
  )</code></pre>
<pre><code>## # A tibble: 2,466 x 4
##    Invoice    Due        Payment    Quarter
##    &lt;date&gt;     &lt;date&gt;     &lt;date&gt;       &lt;dbl&gt;
##  1 2013-01-02 2013-02-01 2013-01-15   2013.
##  2 2013-01-26 2013-02-25 2013-03-03   2013.
##  3 2013-07-03 2013-08-02 2013-07-08   2013.
##  4 2013-02-10 2013-03-12 2013-03-17   2013.
##  5 2012-10-25 2012-11-24 2012-11-28   2012.
##  6 2012-01-27 2012-02-26 2012-02-22   2012.
##  7 2013-08-13 2013-09-12 2013-09-09   2013.
##  8 2012-12-16 2013-01-15 2013-01-12   2013.
##  9 2012-05-14 2012-06-13 2012-07-01   2012.
## 10 2013-07-01 2013-07-31 2013-07-26   2013.
## # ... with 2,456 more rows</code></pre>
<section id="case-study-1" class="level4">
<h4 class="anchored" data-anchor-id="case-study-1">
Case Study
</h4>
<p>
<img src="https://blog.rsquaredacademy.com/img/convert.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
Let us also get the course interval in different units.
</p>
<p>
<br>
</p>
<pre class="r"><code>course_interval / dseconds()
## [1] 777600
course_interval / dminutes()
## [1] 12960
course_interval / dhours()
## [1] 216
course_interval / dweeks()
## [1] 1.285714
course_interval / dyears()
## [1] 0.02464066</code></pre>
<p>
We can use <code>time_length()</code> to get the course interval in different units.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/time_length.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>time_length(course_interval, unit = "seconds")
## [1] 777600
time_length(course_interval, unit = "minutes")
## [1] 12960
time_length(course_interval, unit = "hours")
## [1] 216</code></pre>
<p>
<code>as.period()</code> is another way to get the course interval in different units.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/as_period.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>as.period(course_interval, unit = "seconds")
## [1] "777600S"
as.period(course_interval, unit = "minutes")
## [1] "12960M 0S"
as.period(course_interval, unit = "hours")
## [1] "216H 0M 0S"</code></pre>
</section>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://lubridate.tidyverse.org/" class="uri">https://lubridate.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/dates-and-times.html" class="uri">http://r4ds.had.co.nz/dates-and-times.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>lubridate</category>
  <guid>https://blog.rsquaredacademy.com/posts/working-with-dates-in-r/</guid>
  <pubDate>Sat, 03 Nov 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/lubridate.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Hacking strings with stringr</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/hacking-strings-with-stringr/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-10-22-hacking-strings-with-stringr.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In this post, we will learn to work with string data in R using <a href="http://stringr.tidyverse.org">stringr</a>. As we did in the other posts, we will use a case study to explore the various features of the stringr package. Let us begin by installing and loading stringr and a set of other pacakges we will be using.
</p>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following libraries:
</p>
<ul>
<li>
<a href="http://lubridate.tidyverse.org/index.html">stringr</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
<a href="http://magrittr.tidyverse.org/index.html">magrittr</a>
</li>
<li>
<a href="http://tibble.tidyverse.org/index.html">tibble</a>
</li>
<li>
<a href="http://purrr.tidyverse.org/index.html">purrr</a>
</li>
<li>
and <a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The data sets can be downloaded from <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes from <a href="https://gist.github.com/aravindhebbali/5d1799744e55dc76cdf1af6b1cc03c82">here</a>.
</p>
<pre class="r"><code>library(stringr)
library(tibble)
library(magrittr)
library(purrr)
library(dplyr)
library(readr)</code></pre>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<ul>
<li>
extract domain name from random email ids
</li>
<li>
extract image type from url
</li>
<li>
extract image dimension from url
</li>
<li>
extract extension from domain name
</li>
<li>
extract http protocol from url
</li>
<li>
extract file type from url
</li>
</ul>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>mockstring &lt;- read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/mock_strings.csv')
mockstring</code></pre>
<pre><code>## # A tibble: 1,000 x 12
##       id image_url domain imageurl email filename phone address url   full_name
##    &lt;dbl&gt; &lt;chr&gt;     &lt;chr&gt;  &lt;chr&gt;    &lt;chr&gt; &lt;chr&gt;    &lt;chr&gt; &lt;chr&gt;   &lt;chr&gt; &lt;chr&gt;    
##  1     1 https://~ addto~ http://~ mnew~ PedeMal~ 66-(~ 8 Anha~ http~ Mufi Ruit
##  2     2 https://~ gmpg.~ http://~ mdan~ Loborti~ 351-~ 697 Ea~ http~ Leese Fu~
##  3     3 https://~ samsu~ http://~ hgir~ CongueD~ 33-(~ 89 Dot~ http~ Blakelee~
##  4     4 https://~ spoti~ http://~ pmcm~ Eleifen~ 86-(~ 98135 ~ http~ Terencio~
##  5     5 https://~ wunde~ http://~ dris~ PurusPh~ 223-~ 7814 P~ http~ Debee Mc~
##  6     6 https://~ alexa~ http://~ cphl~ Element~ 420-~ 4897 L~ http~ Fran Pai~
##  7     7 https://~ googl~ http://~ kdod~ Mattis.~ 1-(7~ 53541 ~ http~ Frasco B~
##  8     8 https://~ ed.gov http://~ vhou~ PurusEu~ 62-(~ 4819 H~ http~ Car Pont~
##  9     9 https://~ jigsy~ http://~ rdik~ JustoEt~ 1-(6~ 68096 ~ http~ Tades Ch~
## 10    10 https://~ jugem~ http://~ tdud~ Ante.ti~ 30-(~ 9595 S~ http~ Wilton K~
## # ... with 990 more rows, and 2 more variables: currency &lt;chr&gt;, passwords &lt;chr&gt;</code></pre>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<ul>
<li>
domain: dummy website domain
</li>
<li>
imageurl: url of an image
</li>
<li>
email: dummy email id
</li>
<li>
filename: dummy file name with different extensions
</li>
<li>
phone: dummy phone number
</li>
<li>
address: dummy address with door and street names
</li>
<li>
url: randomyly generated urls
</li>
<li>
full_name: dummy first and last names
</li>
<li>
currency: different currencies
</li>
<li>
passwords: dummy passwords
</li>
</ul>
</section>
</section>
<section id="overview" class="level2">
<h2 class="anchored" data-anchor-id="overview">
Overview
</h2>
<p>
Before we start with the case study, let us take a quick tour of <strong>stringr</strong> and introduce ourselves to some of the functions we will be using later in the case study. One of the columns in the case study data is <code>email</code>. It contains random email ids. We want to ensure that the email ids adher to a particular format .i.e
</p>
<ul>
<li>
they contain <code>@</code>
</li>
<li>
they contain only one <code>@</code>
</li>
</ul>
<p>
Let us first detect if the email ids contain <code>@</code>. Since the data set has 1000 rows, we will use a smaller sample in the examples.
</p>
<pre class="r"><code>mockdata &lt;- slice(mockstring, 1:10)
mockdata</code></pre>
<pre><code>## # A tibble: 10 x 12
##       id image_url domain imageurl email filename phone address url   full_name
##    &lt;dbl&gt; &lt;chr&gt;     &lt;chr&gt;  &lt;chr&gt;    &lt;chr&gt; &lt;chr&gt;    &lt;chr&gt; &lt;chr&gt;   &lt;chr&gt; &lt;chr&gt;    
##  1     1 https://~ addto~ http://~ mnew~ PedeMal~ 66-(~ 8 Anha~ http~ Mufi Ruit
##  2     2 https://~ gmpg.~ http://~ mdan~ Loborti~ 351-~ 697 Ea~ http~ Leese Fu~
##  3     3 https://~ samsu~ http://~ hgir~ CongueD~ 33-(~ 89 Dot~ http~ Blakelee~
##  4     4 https://~ spoti~ http://~ pmcm~ Eleifen~ 86-(~ 98135 ~ http~ Terencio~
##  5     5 https://~ wunde~ http://~ dris~ PurusPh~ 223-~ 7814 P~ http~ Debee Mc~
##  6     6 https://~ alexa~ http://~ cphl~ Element~ 420-~ 4897 L~ http~ Fran Pai~
##  7     7 https://~ googl~ http://~ kdod~ Mattis.~ 1-(7~ 53541 ~ http~ Frasco B~
##  8     8 https://~ ed.gov http://~ vhou~ PurusEu~ 62-(~ 4819 H~ http~ Car Pont~
##  9     9 https://~ jigsy~ http://~ rdik~ JustoEt~ 1-(6~ 68096 ~ http~ Tades Ch~
## 10    10 https://~ jugem~ http://~ tdud~ Ante.ti~ 30-(~ 9595 S~ http~ Wilton K~
## # ... with 2 more variables: currency &lt;chr&gt;, passwords &lt;chr&gt;</code></pre>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_count.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
Use <code>str_detect()</code> to detect <code>@</code> and <code>str_count()</code> to count the number of times <code>@</code> appears in the email ids.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_detect.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code># detect @
str_detect(mockdata$email, pattern = "@")</code></pre>
<pre><code>##  [1] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE</code></pre>
<pre class="r"><code># count @
str_count(mockdata$email, pattern = "@")</code></pre>
<pre><code>##  [1] 1 1 1 1 1 1 1 1 1 1</code></pre>
<p>
We can use <code>str_c()</code> to concatenate strings. Let us add the string <code>email id:</code> before each email id in the data set.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_c.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_c("email id:", mockdata$email)</code></pre>
<pre><code>##  [1] "email id:mnewburn0@fastcompany.com"   
##  [2] "email id:mdankersley1@digg.com"       
##  [3] "email id:hgirhard2@altervista.org"    
##  [4] "email id:pmcmenamy3@sciencedirect.com"
##  [5] "email id:drisbrough4@bandcamp.com"    
##  [6] "email id:cphlippi5@surveymonkey.com"  
##  [7] "email id:kdodswell6@un.org"           
##  [8] "email id:vhourihane7@ovh.net"         
##  [9] "email id:rdike8@timesonline.co.uk"    
## [10] "email id:tdudbridge9@clickbank.net"</code></pre>
<p>
If we want to split a string into two parts using a particular pattern, we use <code>str_split()</code>. Let us split the domain name and extension from the domain column in the data. The domain name and extension are separated by <code>.</code> and we will use it to split the domain column. Since <code>.</code> is a special character, we will use two slashes to escape the special character.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_split.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_split(mockdata$domain, pattern = "\\.")</code></pre>
<pre><code>## [[1]]
## [1] "addtoany" "com"     
## 
## [[2]]
## [1] "gmpg" "org" 
## 
## [[3]]
## [1] "samsung" "com"    
## 
## [[4]]
## [1] "spotify" "com"    
## 
## [[5]]
## [1] "wunderground" "com"         
## 
## [[6]]
## [1] "alexa" "com"  
## 
## [[7]]
## [1] "google" "it"    
## 
## [[8]]
## [1] "ed"  "gov"
## 
## [[9]]
## [1] "jigsy" "com"  
## 
## [[10]]
## [1] "jugem" "jp"</code></pre>
<p>
We can truncate a string using <code>str_trunc()</code>. The default truncation happens at the beggining of the string but we can truncate the central part or the end of the string as well.
</p>
<pre class="r"><code>str_trunc(mockdata$email, width = 10)</code></pre>
<pre><code>##  [1] "mnewbur..." "mdanker..." "hgirhar..." "pmcmena..." "drisbro..."
##  [6] "cphlipp..." "kdodswe..." "vhourih..." "rdike8@..." "tdudbri..."</code></pre>
<pre class="r"><code>str_trunc(mockdata$email, width = 10, side = "left")</code></pre>
<pre><code>##  [1] "...any.com" "...igg.com" "...sta.org" "...ect.com" "...amp.com"
##  [6] "...key.com" "...@un.org" "...ovh.net" "...e.co.uk" "...ank.net"</code></pre>
<pre class="r"><code>str_trunc(mockdata$email, width = 10, side = "center")</code></pre>
<pre><code>##  [1] "mnew...com" "mdan...com" "hgir...org" "pmcm...com" "dris...com"
##  [6] "cphl...com" "kdod...org" "vhou...net" "rdik....uk" "tdud...net"</code></pre>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_sort_descending.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
Strings can be sorted using <code>str_sort()</code>. Let us quickly sort the emails in both ascending and descending orders.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_sort.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_sort(mockdata$email)</code></pre>
<pre><code>##  [1] "cphlippi5@surveymonkey.com"   "drisbrough4@bandcamp.com"    
##  [3] "hgirhard2@altervista.org"     "kdodswell6@un.org"           
##  [5] "mdankersley1@digg.com"        "mnewburn0@fastcompany.com"   
##  [7] "pmcmenamy3@sciencedirect.com" "rdike8@timesonline.co.uk"    
##  [9] "tdudbridge9@clickbank.net"    "vhourihane7@ovh.net"</code></pre>
<pre class="r"><code>str_sort(mockdata$email, decreasing = TRUE)</code></pre>
<pre><code>##  [1] "vhourihane7@ovh.net"          "tdudbridge9@clickbank.net"   
##  [3] "rdike8@timesonline.co.uk"     "pmcmenamy3@sciencedirect.com"
##  [5] "mnewburn0@fastcompany.com"    "mdankersley1@digg.com"       
##  [7] "kdodswell6@un.org"            "hgirhard2@altervista.org"    
##  [9] "drisbrough4@bandcamp.com"     "cphlippi5@surveymonkey.com"</code></pre>
<p>
The case of a string can be changed to upper, lower or title case as shown below.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_to_upper.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_to_upper(mockdata$full_name)</code></pre>
<pre><code>##  [1] "MUFI RUIT"          "LEESE FURMAGIER"    "BLAKELEE WILSHIRE" 
##  [4] "TERENCIO MCILLRICK" "DEBEE MCERLAINE"    "FRAN PAINTEN"      
##  [7] "FRASCO BOWICH"      "CAR PONTEN"         "TADES CHECCUCCI"   
## [10] "WILTON KEMMEY"</code></pre>
<pre class="r"><code>str_to_lower(mockdata$full_name)</code></pre>
<pre><code>##  [1] "mufi ruit"          "leese furmagier"    "blakelee wilshire" 
##  [4] "terencio mcillrick" "debee mcerlaine"    "fran painten"      
##  [7] "frasco bowich"      "car ponten"         "tades checcucci"   
## [10] "wilton kemmey"</code></pre>
<p>
Parts of a string can be replaced using <code>str_replace()</code>. In the <code>address</code> column of the data set, let us replace:
</p>
<ul>
<li>
Street with ST
</li>
<li>
Road with RD
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_replace.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_replace(mockdata$address, "Street", "ST")</code></pre>
<pre><code>##  [1] "8 Anhalt Crossing"          "697 East Avenue"           
##  [3] "89 Dottie Circle"           "98135 Blue Bill Park Drive"
##  [5] "7814 Pennsylvania ST"       "4897 Little Fleur Drive"   
##  [7] "53541 Morrow Center"        "4819 Hermina Parkway"      
##  [9] "68096 Monument Park"        "9595 Spaight Avenue"</code></pre>
<pre class="r"><code>str_replace(mockdata$address, "Road", "RD")</code></pre>
<pre><code>##  [1] "8 Anhalt Crossing"          "697 East Avenue"           
##  [3] "89 Dottie Circle"           "98135 Blue Bill Park Drive"
##  [5] "7814 Pennsylvania Street"   "4897 Little Fleur Drive"   
##  [7] "53541 Morrow Center"        "4819 Hermina Parkway"      
##  [9] "68096 Monument Park"        "9595 Spaight Avenue"</code></pre>
<p>
We can extract parts of the string that match a particular pattern using <code>str_extract()</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_extract.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_extract(mockdata$email, pattern = "org")</code></pre>
<pre><code>##  [1] NA    NA    "org" NA    NA    NA    "org" NA    NA    NA</code></pre>
<p>
Before we extract, we need to know whether the string contains text that match our pattern. Use <code>str_match()</code> to see if the pattern is present in the string.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_match.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_match(mockdata$email, pattern = "org")</code></pre>
<pre><code>##       [,1] 
##  [1,] NA   
##  [2,] NA   
##  [3,] "org"
##  [4,] NA   
##  [5,] NA   
##  [6,] NA   
##  [7,] "org"
##  [8,] NA   
##  [9,] NA   
## [10,] NA</code></pre>
<p>
If we are dealing with a character vector and know that the pattern we are looking at is present in the vector, we might want to know the index of the strings in which it is present. Use <code>str_which()</code> to identify the index of the strings that match our pattern.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_which.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_which(mockdata$email, pattern = "org")</code></pre>
<pre><code>## [1] 3 7</code></pre>
<p>
Another objective might be to locate the position of the pattern we are looking for in the string. For example, if we want to know the position of <code>@</code> in the email ids, we can use <code>str_locate()</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_locate.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_locate(mockdata$email, pattern = "@")</code></pre>
<pre><code>##       start end
##  [1,]    10  10
##  [2,]    13  13
##  [3,]    10  10
##  [4,]    11  11
##  [5,]    12  12
##  [6,]    10  10
##  [7,]    11  11
##  [8,]    12  12
##  [9,]     7   7
## [10,]    12  12</code></pre>
<p>
The length of the string can be computed using <code>str_length()</code>. Let us ensure that the length of the strings in the <code>password</code> column is 16.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_length.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_length(mockdata$passwords)</code></pre>
<pre><code>##  [1] 16 16 16 16 16 16 16 16 16 16</code></pre>
<p>
We can extract parts of a string by specifying the starting and ending position using <code>str_sub()</code>. Let us extract the currency type from the <code>currency</code> column.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/str_sub.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>str_sub(mockdata$currency, start = 1, end = 1)</code></pre>
<pre><code>##  [1] "¥" "$" "\200" "\200" "\200" "¥" "$" "¥" "\200" "\200"</code></pre>
<p>
One final function that we will look at before the case study is <code>word()</code>. It extracts word(s) from sentences. We do not have any sentences in the data set, but let us use it to extract the first and last name from the <code>full_name</code> column.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/word.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>word(mockdata$full_name, 1)</code></pre>
<pre><code>##  [1] "Mufi"     "Leese"    "Blakelee" "Terencio" "Debee"    "Fran"    
##  [7] "Frasco"   "Car"      "Tades"    "Wilton"</code></pre>
<pre class="r"><code>word(mockdata$full_name, 2)</code></pre>
<pre><code>##  [1] "Ruit"      "Furmagier" "Wilshire"  "McIllrick" "McErlaine" "Painten"  
##  [7] "Bowich"    "Ponten"    "Checcucci" "Kemmey"</code></pre>
<p>
Alright, now let us apply what we have learned so far to our case study.
</p>
</section>
<section id="extract-domain-name-from-email-ids" class="level2">
<h2 class="anchored" data-anchor-id="extract-domain-name-from-email-ids">
Extract domain name from email ids
</h2>
<section id="steps" class="level3">
<h3 class="anchored" data-anchor-id="steps">
Steps
</h3>
<ul>
<li>
split email using pattern <code>@</code>
</li>
<li>
extract the second element from the resulting list
</li>
<li>
split the above using pattern <code>\.</code>
</li>
<li>
extract the first element from the resulting list
</li>
</ul>
<p>
Let us take a look at the emails before we extract the domain names.
</p>
<pre class="r"><code>emails &lt;- 
  mockstring %&gt;%
  pull(email) %&gt;%
  head()

emails</code></pre>
<pre><code>## [1] "mnewburn0@fastcompany.com"    "mdankersley1@digg.com"       
## [3] "hgirhard2@altervista.org"     "pmcmenamy3@sciencedirect.com"
## [5] "drisbrough4@bandcamp.com"     "cphlippi5@surveymonkey.com"</code></pre>
<section id="step-1-split-email-using-pattern-." class="level4">
<h4 class="anchored" data-anchor-id="step-1-split-email-using-pattern-.">
Step 1: Split email using pattern <code>@</code>.
</h4>
<p>
We will split the email using <code>str_split</code>. It will split a string using the pattern supplied. In our case the pattern is <code>@</code>.
</p>
<pre class="r"><code> str_split(emails, pattern = '@')</code></pre>
<pre><code>## [[1]]
## [1] "mnewburn0"       "fastcompany.com"
## 
## [[2]]
## [1] "mdankersley1" "digg.com"    
## 
## [[3]]
## [1] "hgirhard2"      "altervista.org"
## 
## [[4]]
## [1] "pmcmenamy3"        "sciencedirect.com"
## 
## [[5]]
## [1] "drisbrough4"  "bandcamp.com"
## 
## [[6]]
## [1] "cphlippi5"        "surveymonkey.com"</code></pre>
</section>
<section id="step-2-extract-the-second-element-from-the-resulting-list." class="level4">
<h4 class="anchored" data-anchor-id="step-2-extract-the-second-element-from-the-resulting-list.">
Step 2: Extract the second element from the resulting list.
</h4>
<p>
Step 1 returned a list. Each element of the list has two values. The first one is the username and the second is the domain name. Since we are extracting the domain name, we want the second value from each element of the list.
</p>
<p>
We will use <code>map_chr()</code> from purrr to extract the domain names. It will return the second value from each element in the list. Since the domain name is a string, <code>map_chr()</code> will return a character vector.
</p>
<pre class="r"><code>emails %&gt;%
  str_split(pattern = '@') %&gt;%
  map_chr(2)</code></pre>
<pre><code>## [1] "fastcompany.com"   "digg.com"          "altervista.org"   
## [4] "sciencedirect.com" "bandcamp.com"      "surveymonkey.com"</code></pre>
</section>
<section id="step-3-split-the-above-using-pattern-.." class="level4">
<h4 class="anchored" data-anchor-id="step-3-split-the-above-using-pattern-..">
Step 3: Split the above using pattern <code>\.</code>.
</h4>
<p>
We want the domain name and not the extension. Step 2 returned a character vector and we need to split the domain name and the domain extension. They are separated by <code>.</code>. Since <code>.</code> is a special character, we will use <code>\</code> before <code>.</code> to escape it. Let us split the domain name and domain extension using <code>str_split</code> and <code>\.</code> as the pattern.
</p>
<pre class="r"><code>emails %&gt;%
  str_split(pattern = '@') %&gt;%
  map_chr(2) %&gt;%
  str_split(pattern = '\\.') </code></pre>
<pre><code>## [[1]]
## [1] "fastcompany" "com"        
## 
## [[2]]
## [1] "digg" "com" 
## 
## [[3]]
## [1] "altervista" "org"       
## 
## [[4]]
## [1] "sciencedirect" "com"          
## 
## [[5]]
## [1] "bandcamp" "com"     
## 
## [[6]]
## [1] "surveymonkey" "com"</code></pre>
</section>
<section id="step-4-extract-the-first-element-from-the-resulting-list." class="level4">
<h4 class="anchored" data-anchor-id="step-4-extract-the-first-element-from-the-resulting-list.">
Step 4: Extract the first element from the resulting list.
</h4>
<p>
Now that we have separated the domain name from its extension, let us extract the first value from each element in the list returned in step 3. We will again use <code>map_chr</code> to achieve this.
</p>
<pre class="r"><code>emails %&gt;%
  str_split(pattern = '@') %&gt;%
  map_chr(2) %&gt;%
  str_split(pattern = '\\.') %&gt;%
  map_chr(extract(1))</code></pre>
<pre><code>## [1] "fastcompany"   "digg"          "altervista"    "sciencedirect"
## [5] "bandcamp"      "surveymonkey"</code></pre>
</section>
</section>
</section>
<section id="extract-domain-extension" class="level2">
<h2 class="anchored" data-anchor-id="extract-domain-extension">
Extract Domain Extension
</h2>
<p>
The below code extracts the domain extension instead of the domain name.
</p>
<pre class="r"><code>emails %&gt;%
  str_split(pattern = '@') %&gt;%
  map_chr(2) %&gt;%
  str_split(pattern = '\\.', simplify = TRUE) %&gt;%
  extract(, 2)</code></pre>
<pre><code>## [1] "com" "com" "org" "com" "com" "com"</code></pre>
</section>
<section id="extract-image-type-from-url" class="level2">
<h2 class="anchored" data-anchor-id="extract-image-type-from-url">
Extract image type from URL
</h2>
<section id="steps-1" class="level3">
<h3 class="anchored" data-anchor-id="steps-1">
Steps
</h3>
<ul>
<li>
split imageurl using pattern <code>\.</code>
</li>
<li>
extract the third value from each element of the resulting list
</li>
<li>
subset the string using the index position
</li>
</ul>
<p>
Let us take a look at the URL of the image.
</p>
<pre class="r"><code>img &lt;- 
  mockstring %&gt;%
  pull(imageurl) %&gt;%
  head()

img</code></pre>
<pre><code>## [1] "http://dummyimage.com/130x183.jpg/dddddd/000000"
## [2] "http://dummyimage.com/106x217.bmp/dddddd/000000"
## [3] "http://dummyimage.com/146x127.bmp/cc0000/ffffff"
## [4] "http://dummyimage.com/181x194.png/5fa2dd/ffffff"
## [5] "http://dummyimage.com/220x123.jpg/ff4444/ffffff"
## [6] "http://dummyimage.com/118x176.bmp/dddddd/000000"</code></pre>
<section id="step-1-split-imageurl-using-pattern-." class="level4">
<h4 class="anchored" data-anchor-id="step-1-split-imageurl-using-pattern-.">
Step 1: Split imageurl using pattern <code>\.</code>
</h4>
<p>
Let us split imageurl using <code>str_split</code> and the pattern <code>\.</code>.
</p>
<pre class="r"><code>str_split(img, pattern = '\\.')</code></pre>
<pre><code>## [[1]]
## [1] "http://dummyimage" "com/130x183"       "jpg/dddddd/000000"
## 
## [[2]]
## [1] "http://dummyimage" "com/106x217"       "bmp/dddddd/000000"
## 
## [[3]]
## [1] "http://dummyimage" "com/146x127"       "bmp/cc0000/ffffff"
## 
## [[4]]
## [1] "http://dummyimage" "com/181x194"       "png/5fa2dd/ffffff"
## 
## [[5]]
## [1] "http://dummyimage" "com/220x123"       "jpg/ff4444/ffffff"
## 
## [[6]]
## [1] "http://dummyimage" "com/118x176"       "bmp/dddddd/000000"</code></pre>
</section>
<section id="step-2-extract-the-third-value-from-each-element-of-the-resulting-list" class="level4">
<h4 class="anchored" data-anchor-id="step-2-extract-the-third-value-from-each-element-of-the-resulting-list">
Step 2: Extract the third value from each element of the resulting list
</h4>
<p>
Step 1 returned a list the elements of which have 3 values each. If you observe the list, the image type is in the 3rd value. We will now extract the third value from each element of the list using <code>map_chr</code>.
</p>
<pre class="r"><code>img %&gt;%
  str_split(pattern = '\\.') %&gt;%
  map_chr(extract(3))</code></pre>
<pre><code>## [1] "jpg/dddddd/000000" "bmp/dddddd/000000" "bmp/cc0000/ffffff"
## [4] "png/5fa2dd/ffffff" "jpg/ff4444/ffffff" "bmp/dddddd/000000"</code></pre>
</section>
<section id="step-3-subset-the-string-using-the-index-position" class="level4">
<h4 class="anchored" data-anchor-id="step-3-subset-the-string-using-the-index-position">
Step 3: Subset the string using the index position
</h4>
<p>
We can now extract the image type in two ways:
</p>
<ul>
<li>
subset the first 3 characters of the string
</li>
<li>
split the string using pattern <code>/</code> and extract the first value from the elements of the resulting list
</li>
</ul>
<p>
Below is the first method. We know that the image type is 3 characters. So we use <code>str_sub</code> to subset the first 3 characters. The index positions are mentioned using <code>start</code> and <code>stop</code>.
</p>
<pre class="r"><code>img %&gt;%
  str_split(pattern = '\\.') %&gt;%
  map_chr(extract(3)) %&gt;%
  str_sub(start = 1, end = 3)</code></pre>
<pre><code>## [1] "jpg" "bmp" "bmp" "png" "jpg" "bmp"</code></pre>
<p>
In case you are not sure about the length of the image type. In such cases, we will split the string using pattern <code>/</code> and then use <code>map_chr</code> to extract the first value of each element of the resulting list.
</p>
<pre class="r"><code>img %&gt;%
  str_split(pattern = '\\.') %&gt;%
  map_chr(extract(3)) %&gt;%
  str_split(pattern = '/') %&gt;%
  map_chr(extract(1))</code></pre>
<pre><code>## [1] "jpg" "bmp" "bmp" "png" "jpg" "bmp"</code></pre>
</section>
</section>
</section>
<section id="extract-image-dimesion-from-url" class="level2">
<h2 class="anchored" data-anchor-id="extract-image-dimesion-from-url">
Extract Image Dimesion from URL
</h2>
<section id="steps-2" class="level3">
<h3 class="anchored" data-anchor-id="steps-2">
Steps
</h3>
<ul>
<li>
locate numbers between 0 and 9
</li>
<li>
extract part of url starting with image dimension
</li>
<li>
split the string using the pattern <code>\.</code>
</li>
<li>
extract the first element
</li>
</ul>
<section id="step-1-locate-numbers-between-0-and-9." class="level4">
<h4 class="anchored" data-anchor-id="step-1-locate-numbers-between-0-and-9.">
Step 1: Locate numbers between 0 and 9.
</h4>
<p>
Let us inspect the image url. The dimension of the image appears after the domain extension and there are no numbers in the url before. We will locate the position or index of the first number in the url using <code>str_locate()</code> and using the pattern <code>[0-9]</code> which instructs to look for any number between and including 0 and 9.
</p>
<pre class="r"><code>str_locate(img, pattern = "[0-9]") </code></pre>
<pre><code>##      start end
## [1,]    23  23
## [2,]    23  23
## [3,]    23  23
## [4,]    23  23
## [5,]    23  23
## [6,]    23  23</code></pre>
</section>
<section id="step-2-extract-url" class="level4">
<h4 class="anchored" data-anchor-id="step-2-extract-url">
Step 2: Extract url
</h4>
<p>
We know where the dimension is located in the url. Let us extract the part of the url that contains the image dimension using <code>str_sub()</code>.
</p>
<pre class="r"><code>str_sub(img, start = 23) </code></pre>
<pre><code>## [1] "130x183.jpg/dddddd/000000" "106x217.bmp/dddddd/000000"
## [3] "146x127.bmp/cc0000/ffffff" "181x194.png/5fa2dd/ffffff"
## [5] "220x123.jpg/ff4444/ffffff" "118x176.bmp/dddddd/000000"</code></pre>
</section>
<section id="step-3-split-the-string-using-the-pattern-.." class="level4">
<h4 class="anchored" data-anchor-id="step-3-split-the-string-using-the-pattern-..">
Step 3: Split the string using the pattern <code>\.</code>.
</h4>
<p>
From the previous step, we have the part of the url that contains the image dimension. To extract the dimension, we will split it from the rest of the url using <code>str_split()</code> and using the pattern <code>\.</code> as it separates the dimension and the image extension.
</p>
<pre class="r"><code>img %&gt;%
  str_sub(start = 23) %&gt;%
  str_split(pattern = '\\.') </code></pre>
<pre><code>## [[1]]
## [1] "130x183"           "jpg/dddddd/000000"
## 
## [[2]]
## [1] "106x217"           "bmp/dddddd/000000"
## 
## [[3]]
## [1] "146x127"           "bmp/cc0000/ffffff"
## 
## [[4]]
## [1] "181x194"           "png/5fa2dd/ffffff"
## 
## [[5]]
## [1] "220x123"           "jpg/ff4444/ffffff"
## 
## [[6]]
## [1] "118x176"           "bmp/dddddd/000000"</code></pre>
</section>
<section id="step-4-extract-the-first-element." class="level4">
<h4 class="anchored" data-anchor-id="step-4-extract-the-first-element.">
Step 4: Extract the first element.
</h4>
<p>
The above step resulted in a list which contains the image dimension and the rest of the url. Each element of the list is a character vector. We want to extract the first value in the character vector. Let us use <code>map_chr()</code> to extract the first value from each element of the list.
</p>
<pre class="r"><code>img %&gt;%
  str_sub(start = 23) %&gt;%
  str_split(pattern = '\\.') %&gt;%
  map_chr(extract(1))</code></pre>
<pre><code>## [1] "130x183" "106x217" "146x127" "181x194" "220x123" "118x176"</code></pre>
</section>
</section>
</section>
<section id="extract-http-protocol-from-url" class="level2">
<h2 class="anchored" data-anchor-id="extract-http-protocol-from-url">
Extract HTTP Protocol from URL
</h2>
<pre class="r"><code>url1 &lt;- 
  mockstring %&gt;%
  pull(url) %&gt;%
  first()

url1</code></pre>
<pre><code>## [1] "https://engadget.com/nascetur/ridiculus/mus/vivamus/vestibulum.jsp?eu=est&amp;tincidunt=risus&amp;in=auctor&amp;leo=sed&amp;maecenas=tristique&amp;pulvinar=in&amp;lobortis=tempus&amp;est=sit&amp;phasellus=amet&amp;sit=sem&amp;amet=fusce&amp;erat=consequat&amp;nulla=nulla&amp;tempus=nisl&amp;vivamus=nunc&amp;in=nisl&amp;felis=duis&amp;eu=bibendum&amp;sapien=felis&amp;cursus=sed&amp;vestibulum=interdum&amp;proin=venenatis&amp;eu=turpis&amp;mi=enim&amp;nulla=blandit&amp;ac=mi&amp;enim=in&amp;in=porttitor&amp;tempor=pede&amp;turpis=justo&amp;nec=eu&amp;euismod=massa&amp;scelerisque=donec&amp;quam=dapibus&amp;turpis=duis&amp;adipiscing=at&amp;lorem=velit&amp;vitae=eu&amp;mattis=est&amp;nibh=congue&amp;ligula=elementum&amp;nec=in&amp;sem=hac&amp;duis=habitasse&amp;aliquam=platea&amp;convallis=dictumst&amp;nunc=morbi&amp;proin=vestibulum&amp;at=velit&amp;turpis=id&amp;a=pretium&amp;pede=iaculis&amp;posuere=diam&amp;nonummy=erat&amp;integer=fermentum&amp;non=justo&amp;velit=nec&amp;donec=condimentum&amp;diam=neque&amp;neque=sapien&amp;vestibulum=placerat&amp;eget=ante&amp;vulputate=nulla&amp;ut=justo&amp;ultrices=aliquam&amp;vel=quis&amp;augue=turpis&amp;vestibulum=eget&amp;ante=elit&amp;ipsum=sodales&amp;primis=scelerisque&amp;in=mauris&amp;faucibus=sit&amp;orci=amet&amp;luctus=eros&amp;et=suspendisse&amp;ultrices=accumsan&amp;posuere=tortor&amp;cubilia=quis&amp;curae=turpis&amp;donec=sed&amp;pharetra=ante&amp;magna=vivamus&amp;vestibulum=tortor&amp;aliquet=duis&amp;ultrices=mattis&amp;erat=egestas&amp;tortor=metus&amp;sollicitudin=aenean&amp;mi=fermentum&amp;sit=donec"</code></pre>
<section id="steps-3" class="level3">
<h3 class="anchored" data-anchor-id="steps-3">
Steps
</h3>
<ul>
<li>
split the url using the pattern <code>://</code>
</li>
<li>
extract the first element
</li>
</ul>
<section id="step-1-split-the-url-using-the-pattern-." class="level4">
<h4 class="anchored" data-anchor-id="step-1-split-the-url-using-the-pattern-.">
Step 1: Split the url using the pattern <code>://</code>.
</h4>
<p>
The HTTP protocol is the first part of the url and is separated from the rest of the url by <code>:</code>. Let us split the url using <code>str_split()</code> and using the pattern <code>:</code>. Since <code>:</code> is a special character, we will escape it using <code>\</code>.
</p>
<pre class="r"><code>str_split(url1, pattern = '://') </code></pre>
<pre><code>## [[1]]
## [1] "https"                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               
## [2] "engadget.com/nascetur/ridiculus/mus/vivamus/vestibulum.jsp?eu=est&amp;tincidunt=risus&amp;in=auctor&amp;leo=sed&amp;maecenas=tristique&amp;pulvinar=in&amp;lobortis=tempus&amp;est=sit&amp;phasellus=amet&amp;sit=sem&amp;amet=fusce&amp;erat=consequat&amp;nulla=nulla&amp;tempus=nisl&amp;vivamus=nunc&amp;in=nisl&amp;felis=duis&amp;eu=bibendum&amp;sapien=felis&amp;cursus=sed&amp;vestibulum=interdum&amp;proin=venenatis&amp;eu=turpis&amp;mi=enim&amp;nulla=blandit&amp;ac=mi&amp;enim=in&amp;in=porttitor&amp;tempor=pede&amp;turpis=justo&amp;nec=eu&amp;euismod=massa&amp;scelerisque=donec&amp;quam=dapibus&amp;turpis=duis&amp;adipiscing=at&amp;lorem=velit&amp;vitae=eu&amp;mattis=est&amp;nibh=congue&amp;ligula=elementum&amp;nec=in&amp;sem=hac&amp;duis=habitasse&amp;aliquam=platea&amp;convallis=dictumst&amp;nunc=morbi&amp;proin=vestibulum&amp;at=velit&amp;turpis=id&amp;a=pretium&amp;pede=iaculis&amp;posuere=diam&amp;nonummy=erat&amp;integer=fermentum&amp;non=justo&amp;velit=nec&amp;donec=condimentum&amp;diam=neque&amp;neque=sapien&amp;vestibulum=placerat&amp;eget=ante&amp;vulputate=nulla&amp;ut=justo&amp;ultrices=aliquam&amp;vel=quis&amp;augue=turpis&amp;vestibulum=eget&amp;ante=elit&amp;ipsum=sodales&amp;primis=scelerisque&amp;in=mauris&amp;faucibus=sit&amp;orci=amet&amp;luctus=eros&amp;et=suspendisse&amp;ultrices=accumsan&amp;posuere=tortor&amp;cubilia=quis&amp;curae=turpis&amp;donec=sed&amp;pharetra=ante&amp;magna=vivamus&amp;vestibulum=tortor&amp;aliquet=duis&amp;ultrices=mattis&amp;erat=egestas&amp;tortor=metus&amp;sollicitudin=aenean&amp;mi=fermentum&amp;sit=donec"</code></pre>
</section>
<section id="step-2-extract-the-first-element." class="level4">
<h4 class="anchored" data-anchor-id="step-2-extract-the-first-element.">
Step 2: Extract the first element.
</h4>
<p>
The HTTP protocol is the first value in each element of the list. As we did in the previous example, we will extact it using <code>map_chr()</code> and <code>extract()</code>.
</p>
<pre class="r"><code>url1 %&gt;%
  str_split(pattern = '://') %&gt;%
  map_chr(extract(1))</code></pre>
<pre><code>## [1] "https"</code></pre>
</section>
</section>
</section>
<section id="extract-file-type" class="level2">
<h2 class="anchored" data-anchor-id="extract-file-type">
Extract file type
</h2>
<pre class="r"><code>urls &lt;-
  mockstring %&gt;%
  use_series(url) %&gt;%
  extract(1:3)</code></pre>
<section id="steps-4" class="level3">
<h3 class="anchored" data-anchor-id="steps-4">
Steps
</h3>
<ul>
<li>
check if there are only 2 dots in the URL
</li>
<li>
check if there is only 1 question mark in the URL
</li>
<li>
detect the staritng position of file type
</li>
<li>
tetect the ending position of file type
</li>
<li>
use the locations to specify the index position for extracting file type
</li>
</ul>
<section id="step-1-check-if-there-are-only-2-dots-in-the-url" class="level4">
<h4 class="anchored" data-anchor-id="step-1-check-if-there-are-only-2-dots-in-the-url">
Step 1: Check if there are only 2 dots in the URL
</h4>
<p>
Let us locate all the dots in the url using <code>str_locate_all()</code> and see if any of them contain more than 2 dots.
</p>
<pre class="r"><code>urls %&gt;%
  str_locate_all(pattern = '\\.') %&gt;%
  map_int(nrow) %&gt;%
  is_greater_than(2) %&gt;%
  sum()</code></pre>
<pre><code>## [1] 0</code></pre>
</section>
<section id="step-2-check-if-there-is-only-1-question-mark-in-the-url" class="level4">
<h4 class="anchored" data-anchor-id="step-2-check-if-there-is-only-1-question-mark-in-the-url">
Step 2: Check if there is only 1 question mark in the URL
</h4>
<p>
The next step is to check if there is only one <code>?</code> (question mark) in the url.
</p>
<pre class="r"><code>urls %&gt;%
  str_locate_all(pattern = "[?]") %&gt;%
  map_int(nrow) %&gt;%
  is_greater_than(1) %&gt;%
  sum()</code></pre>
<pre><code>## [1] 0</code></pre>
</section>
<section id="step-3-detect-the-staritng-position-of-file-type" class="level4">
<h4 class="anchored" data-anchor-id="step-3-detect-the-staritng-position-of-file-type">
Step 3: Detect the staritng position of file type
</h4>
<p>
Since the file type is located between the second dot and the first quesiton mark in the url, let us extract the location of the second dot and add 1 as the file type starts after the dot.
</p>
<pre class="r"><code>d &lt;- 
  urls %&gt;%
  str_locate_all(pattern = '\\.') %&gt;%
  map_int(extract(2)) %&gt;%
  add(1)

d  </code></pre>
<pre><code>## [1] 64 47 48</code></pre>
</section>
<section id="step-4-detect-the-ending-position-of-file-type" class="level4">
<h4 class="anchored" data-anchor-id="step-4-detect-the-ending-position-of-file-type">
Step 4: Detect the ending position of file type
</h4>
<p>
In step 2, we confirmed that the url has only one question mark. Let us locate the question mark in the url and subtract 1 (as the file type ends before the question mark) so that we get the ending postion of the file type. .
</p>
<pre class="r"><code>q &lt;-  
  urls %&gt;%
  str_locate_all(pattern = "[?]") %&gt;%
  map_int(extract(1)) %&gt;%
  subtract(1)

q</code></pre>
<pre><code>## [1] 66 50 51</code></pre>
</section>
<section id="step-5-specify-the-index-position-for-extracting-file-type" class="level4">
<h4 class="anchored" data-anchor-id="step-5-specify-the-index-position-for-extracting-file-type">
Step 5: Specify the index position for extracting file type
</h4>
<p>
From steps 3 and 4, we have the location of the second dot and the first question mark in the url. Let us use them with <code>str_sub()</code> to extract the file type.
</p>
<pre class="r"><code>str_sub(urls, start = d, end = q)</code></pre>
<pre><code>## [1] "jsp"  "json" "json"</code></pre>
</section>
</section>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://stringr.tidyverse.org/" class="uri">https://stringr.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/strings.html" class="uri">http://r4ds.had.co.nz/strings.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>stringr</category>
  <guid>https://blog.rsquaredacademy.com/posts/hacking-strings-with-stringr/</guid>
  <pubDate>Mon, 22 Oct 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/stringr.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Introduction to tibbles</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/introduction-to-tibbles/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-09-28-introduction-to-tibbles.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<blockquote class="blockquote">
<p>
A <strong>tibble</strong>, or <code>tbl_df</code>, is a modern reimagining of the data.frame, keeping what time has proven to be effective, and throwing out what is not. Tibbles are data.frames that are lazy and surly: they do less (i.e.&nbsp;they don’t change variable names or types, and don’t do partial matching) and complain more (e.g.&nbsp;when a variable does not exist). This forces you to confront problems earlier, typically leading to cleaner, more expressive code. Tibbles also have an enhanced <code>print method()</code> which makes them easier to use with large datasets containing complex objects.
</p>
</blockquote>
<blockquote class="blockquote">
<p>
Source: <a href="https://tibble.tidyverse.org/" class="uri">https://tibble.tidyverse.org/</a>
</p>
</blockquote>
<p>
In this post, we will explore tibbles. To be more precise, we will learn:
</p>
<ul>
<li>
how tibbles are different from data frames?
</li>
<li>
how to create tibbles?
</li>
<li>
how to manipulate tibbles?
</li>
</ul>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://tibble.tidyverse.org/">tibble</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/">dplyr</a>
</li>
</ul>
<p>
The code can be found <a href="https://gist.github.com/aravindhebbali/9a3814b9b4bb5c271d030b15ce4ecdf1">here</a>.
</p>
<pre class="r"><code>library(tibble)
library(dplyr)</code></pre>
</section>
<section id="creating-tibbles" class="level2">
<h2 class="anchored" data-anchor-id="creating-tibbles">
Creating tibbles
</h2>
<p>
tibble can be created using any of the following:
</p>
<ul>
<li>
<code>tibble()</code>
</li>
<li>
<code>as_tibble()</code>
</li>
<li>
<code>tribble()</code>
</li>
</ul>
<p>
Let us start with <code>tibble()</code>.
</p>
<pre class="r"><code>tibble(x = letters,
       y = 1:26,
       z = sample(100, 26))</code></pre>
<pre><code>## # A tibble: 26 x 3
##    x         y     z
##    &lt;chr&gt; &lt;int&gt; &lt;int&gt;
##  1 a         1    71
##  2 b         2    56
##  3 c         3    59
##  4 d         4    14
##  5 e         5    58
##  6 f         6    60
##  7 g         7    16
##  8 h         8     9
##  9 i         9    66
## 10 j        10    48
## # ... with 16 more rows</code></pre>
<p>
We mentioned the column names followed by the data. If you do not specify the column names, <code>tibble()</code> will supply them. Ensure that the length of each column is same.
</p>
</section>
<section id="tibble-features" class="level2">
<h2 class="anchored" data-anchor-id="tibble-features">
tibble features
</h2>
<ul>
<li>
never changes input’s types
</li>
</ul>
<p>
<code>tibble()</code> will never alter the input’s type. For example, if you supply a character vector it will not be converted to factor unlike data.frame where you need to set <code>stringsAsFactors</code> to <code>FALSE</code>.
</p>
<pre class="r"><code>tibble(x = letters,
       y = 1:26,
       z = sample(100, 26))</code></pre>
<pre><code>## # A tibble: 26 x 3
##    x         y     z
##    &lt;chr&gt; &lt;int&gt; &lt;int&gt;
##  1 a         1    62
##  2 b         2    13
##  3 c         3    75
##  4 d         4    17
##  5 e         5    82
##  6 f         6    83
##  7 g         7     9
##  8 h         8    76
##  9 i         9    19
## 10 j        10    97
## # ... with 16 more rows</code></pre>
<ul>
<li>
never adjusts variable names
</li>
</ul>
<p>
<code>tibble()</code> will never modify the column names. In the below example, you can observe that while <code>data.frame</code> adds a <code>.</code>, <code>tibble()</code> retains the column names as is.
</p>
<pre class="r"><code>names(data.frame(`order value` = 10))</code></pre>
<pre><code>## [1] "order.value"</code></pre>
<pre class="r"><code>names(tibble(`order value` = 10))</code></pre>
<pre><code>## [1] "order value"</code></pre>
<ul>
<li>
never prints all rows
</li>
</ul>
<p>
<code>tibble()</code> will never print all the rows and clutter your console. It will only print the first 10 rows and only as many columns that fit the width of the console.
</p>
<pre class="r"><code>x &lt;- 1:100
y &lt;- letters[1]
z &lt;- sample(c(TRUE, FALSE), 100, replace = TRUE)
tibble(x, y, z)</code></pre>
<pre><code>## # A tibble: 100 x 3
##        x y     z    
##    &lt;int&gt; &lt;chr&gt; &lt;lgl&gt;
##  1     1 a     TRUE 
##  2     2 a     FALSE
##  3     3 a     TRUE 
##  4     4 a     TRUE 
##  5     5 a     FALSE
##  6     6 a     TRUE 
##  7     7 a     TRUE 
##  8     8 a     FALSE
##  9     9 a     TRUE 
## 10    10 a     FALSE
## # ... with 90 more rows</code></pre>
<ul>
<li>
never recycles vector of length greater than 1
</li>
</ul>
<p>
Recycling vectors of length greater than 1 often leads to errors and as such <code>tibble()</code> will only recycle vectors of length 1.
</p>
<pre class="r"><code>x &lt;- 1:100
y &lt;- letters
z &lt;- sample(c(TRUE, FALSE), 100, replace = TRUE)
tibble(x, y, z)
Error in overscope_eval_next(overscope, expr) : object 'y' not found</code></pre>
</section>
<section id="membership-testing" class="level2">
<h2 class="anchored" data-anchor-id="membership-testing">
Membership Testing
</h2>
<p>
We can test if an object is a tibble using <code>is_tibble()</code>.
</p>
<pre class="r"><code>is_tibble(mtcars)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>is_tibble(as_tibble(mtcars))</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="tribble" class="level2">
<h2 class="anchored" data-anchor-id="tribble">
Tribble
</h2>
<p>
Another way to create tibbles is using <code>tribble()</code>:
</p>
<ul>
<li>
it is short for transposed tibbles
</li>
<li>
it is customized for data entry in code
</li>
<li>
column names start with <code>~</code>
</li>
<li>
and values are separated by commas
</li>
</ul>
<pre class="r"><code>tribble(
  ~x, ~y, ~z,
  #--|--|----
  1, TRUE, 'a',
  2, FALSE, 'b'
)</code></pre>
<pre><code>## # A tibble: 2 x 3
##       x y     z    
##   &lt;dbl&gt; &lt;lgl&gt; &lt;chr&gt;
## 1     1 TRUE  a    
## 2     2 FALSE b</code></pre>
</section>
<section id="column-names" class="level2">
<h2 class="anchored" data-anchor-id="column-names">
Column Names
</h2>
<p>
Names of the columns in tibbles need not be valid R variable names. They can contain unusual characters like a space or a smiley but must be enclosed in ticks.
</p>
<pre class="r"><code>tibble(
  ` ` = 'space',
  `2` = 'integer',
  `:)` = 'smiley'
)</code></pre>
<pre><code>## # A tibble: 1 x 3
##   ` `   `2`     `:)`  
##   &lt;chr&gt; &lt;chr&gt;   &lt;chr&gt; 
## 1 space integer smiley</code></pre>
</section>
<section id="add-rows" class="level2">
<h2 class="anchored" data-anchor-id="add-rows">
Add Rows
</h2>
<p>
Let us add data related to <strong>Safari</strong> browser to the web traffic data using <code>add_row()</code>.
</p>
<pre class="r"><code>browsers &lt;- enframe(c(chrome = 40, firefox = 20, edge = 30))
browsers</code></pre>
<pre><code>## # A tibble: 3 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 firefox    20
## 3 edge       30</code></pre>
<pre class="r"><code>add_row(browsers, name = 'safari', value = 10)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 firefox    20
## 3 edge       30
## 4 safari     10</code></pre>
<p>
If we want to add the data at a particular row, we can specify the row number using the <code>.before</code> argument. Let us add the data related to <strong>Safari</strong> browser in the second row instead of the last row.
</p>
<pre class="r"><code>add_row(browsers, name = 'safari', value = 10, .before = 2)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 safari     10
## 3 firefox    20
## 4 edge       30</code></pre>
</section>
<section id="add-columns" class="level2">
<h2 class="anchored" data-anchor-id="add-columns">
Add Columns
</h2>
<p>
<code>add_column()</code> adds a new column to tibbles.
</p>
<pre class="r"><code>browsers &lt;- enframe(c(chrome = 40, firefox = 20, edge = 30, safari = 10))
add_column(browsers, visits = c(4000, 2000, 3000, 1000))</code></pre>
<pre><code>## # A tibble: 4 x 3
##   name    value visits
##   &lt;chr&gt;   &lt;dbl&gt;  &lt;dbl&gt;
## 1 chrome     40   4000
## 2 firefox    20   2000
## 3 edge       30   3000
## 4 safari     10   1000</code></pre>
</section>
<section id="rownames" class="level2">
<h2 class="anchored" data-anchor-id="rownames">
Rownames
</h2>
<p>
The <a href="tibble.tidyverse.org">tibble</a> package provides a set of functions to deal with rownames. Remember, <code>tibble</code> does not have <code>rownames</code> unlike <code>data.frame</code>. To check whether a data set has rownames, use <code>has_rownames()</code>.
</p>
<pre class="r"><code>has_rownames(mtcars)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<section id="remove-rownames" class="level4">
<h4 class="anchored" data-anchor-id="remove-rownames">
Remove Rownames
</h4>
<pre class="r"><code>remove_rownames(mtcars)</code></pre>
<pre><code>##     mpg cyl  disp  hp drat    wt  qsec vs am gear carb
## 1  21.0   6 160.0 110 3.90 2.620 16.46  0  1    4    4
## 2  21.0   6 160.0 110 3.90 2.875 17.02  0  1    4    4
## 3  22.8   4 108.0  93 3.85 2.320 18.61  1  1    4    1
## 4  21.4   6 258.0 110 3.08 3.215 19.44  1  0    3    1
## 5  18.7   8 360.0 175 3.15 3.440 17.02  0  0    3    2
## 6  18.1   6 225.0 105 2.76 3.460 20.22  1  0    3    1
## 7  14.3   8 360.0 245 3.21 3.570 15.84  0  0    3    4
## 8  24.4   4 146.7  62 3.69 3.190 20.00  1  0    4    2
## 9  22.8   4 140.8  95 3.92 3.150 22.90  1  0    4    2
## 10 19.2   6 167.6 123 3.92 3.440 18.30  1  0    4    4
## 11 17.8   6 167.6 123 3.92 3.440 18.90  1  0    4    4
## 12 16.4   8 275.8 180 3.07 4.070 17.40  0  0    3    3
## 13 17.3   8 275.8 180 3.07 3.730 17.60  0  0    3    3
## 14 15.2   8 275.8 180 3.07 3.780 18.00  0  0    3    3
## 15 10.4   8 472.0 205 2.93 5.250 17.98  0  0    3    4
## 16 10.4   8 460.0 215 3.00 5.424 17.82  0  0    3    4
## 17 14.7   8 440.0 230 3.23 5.345 17.42  0  0    3    4
## 18 32.4   4  78.7  66 4.08 2.200 19.47  1  1    4    1
## 19 30.4   4  75.7  52 4.93 1.615 18.52  1  1    4    2
## 20 33.9   4  71.1  65 4.22 1.835 19.90  1  1    4    1
## 21 21.5   4 120.1  97 3.70 2.465 20.01  1  0    3    1
## 22 15.5   8 318.0 150 2.76 3.520 16.87  0  0    3    2
## 23 15.2   8 304.0 150 3.15 3.435 17.30  0  0    3    2
## 24 13.3   8 350.0 245 3.73 3.840 15.41  0  0    3    4
## 25 19.2   8 400.0 175 3.08 3.845 17.05  0  0    3    2
## 26 27.3   4  79.0  66 4.08 1.935 18.90  1  1    4    1
## 27 26.0   4 120.3  91 4.43 2.140 16.70  0  1    5    2
## 28 30.4   4  95.1 113 3.77 1.513 16.90  1  1    5    2
## 29 15.8   8 351.0 264 4.22 3.170 14.50  0  1    5    4
## 30 19.7   6 145.0 175 3.62 2.770 15.50  0  1    5    6
## 31 15.0   8 301.0 335 3.54 3.570 14.60  0  1    5    8
## 32 21.4   4 121.0 109 4.11 2.780 18.60  1  1    4    2</code></pre>
</section>
<section id="rownames-to-column" class="level4">
<h4 class="anchored" data-anchor-id="rownames-to-column">
Rownames to Column
</h4>
<pre class="r"><code>head(rownames_to_column(mtcars))</code></pre>
<pre><code>##             rowname  mpg cyl disp  hp drat    wt  qsec vs am gear carb
## 1         Mazda RX4 21.0   6  160 110 3.90 2.620 16.46  0  1    4    4
## 2     Mazda RX4 Wag 21.0   6  160 110 3.90 2.875 17.02  0  1    4    4
## 3        Datsun 710 22.8   4  108  93 3.85 2.320 18.61  1  1    4    1
## 4    Hornet 4 Drive 21.4   6  258 110 3.08 3.215 19.44  1  0    3    1
## 5 Hornet Sportabout 18.7   8  360 175 3.15 3.440 17.02  0  0    3    2
## 6           Valiant 18.1   6  225 105 2.76 3.460 20.22  1  0    3    1</code></pre>
</section>
<section id="column-to-rownames" class="level4">
<h4 class="anchored" data-anchor-id="column-to-rownames">
Column to Rownames
</h4>
<p>
To convert the first column in the data set to rownames, use <code>column_to_rownames()</code>:
</p>
<pre class="r"><code>mtcars_tbl &lt;- rownames_to_column(mtcars)
column_to_rownames(mtcars_tbl)</code></pre>
<pre><code>##                      mpg cyl  disp  hp drat    wt  qsec vs am gear carb
## Mazda RX4           21.0   6 160.0 110 3.90 2.620 16.46  0  1    4    4
## Mazda RX4 Wag       21.0   6 160.0 110 3.90 2.875 17.02  0  1    4    4
## Datsun 710          22.8   4 108.0  93 3.85 2.320 18.61  1  1    4    1
## Hornet 4 Drive      21.4   6 258.0 110 3.08 3.215 19.44  1  0    3    1
## Hornet Sportabout   18.7   8 360.0 175 3.15 3.440 17.02  0  0    3    2
## Valiant             18.1   6 225.0 105 2.76 3.460 20.22  1  0    3    1
## Duster 360          14.3   8 360.0 245 3.21 3.570 15.84  0  0    3    4
## Merc 240D           24.4   4 146.7  62 3.69 3.190 20.00  1  0    4    2
## Merc 230            22.8   4 140.8  95 3.92 3.150 22.90  1  0    4    2
## Merc 280            19.2   6 167.6 123 3.92 3.440 18.30  1  0    4    4
## Merc 280C           17.8   6 167.6 123 3.92 3.440 18.90  1  0    4    4
## Merc 450SE          16.4   8 275.8 180 3.07 4.070 17.40  0  0    3    3
## Merc 450SL          17.3   8 275.8 180 3.07 3.730 17.60  0  0    3    3
## Merc 450SLC         15.2   8 275.8 180 3.07 3.780 18.00  0  0    3    3
## Cadillac Fleetwood  10.4   8 472.0 205 2.93 5.250 17.98  0  0    3    4
## Lincoln Continental 10.4   8 460.0 215 3.00 5.424 17.82  0  0    3    4
## Chrysler Imperial   14.7   8 440.0 230 3.23 5.345 17.42  0  0    3    4
## Fiat 128            32.4   4  78.7  66 4.08 2.200 19.47  1  1    4    1
## Honda Civic         30.4   4  75.7  52 4.93 1.615 18.52  1  1    4    2
## Toyota Corolla      33.9   4  71.1  65 4.22 1.835 19.90  1  1    4    1
## Toyota Corona       21.5   4 120.1  97 3.70 2.465 20.01  1  0    3    1
## Dodge Challenger    15.5   8 318.0 150 2.76 3.520 16.87  0  0    3    2
## AMC Javelin         15.2   8 304.0 150 3.15 3.435 17.30  0  0    3    2
## Camaro Z28          13.3   8 350.0 245 3.73 3.840 15.41  0  0    3    4
## Pontiac Firebird    19.2   8 400.0 175 3.08 3.845 17.05  0  0    3    2
## Fiat X1-9           27.3   4  79.0  66 4.08 1.935 18.90  1  1    4    1
## Porsche 914-2       26.0   4 120.3  91 4.43 2.140 16.70  0  1    5    2
## Lotus Europa        30.4   4  95.1 113 3.77 1.513 16.90  1  1    5    2
## Ford Pantera L      15.8   8 351.0 264 4.22 3.170 14.50  0  1    5    4
## Ferrari Dino        19.7   6 145.0 175 3.62 2.770 15.50  0  1    5    6
## Maserati Bora       15.0   8 301.0 335 3.54 3.570 14.60  0  1    5    8
## Volvo 142E          21.4   4 121.0 109 4.11 2.780 18.60  1  1    4    2</code></pre>
</section>
</section>
<section id="glimpse" class="level2">
<h2 class="anchored" data-anchor-id="glimpse">
Glimpse
</h2>
<p>
Use <code>glimpse()</code> to get an overview of the data.
</p>
<pre class="r"><code>glimpse(mtcars)</code></pre>
<pre><code>## Rows: 32
## Columns: 11
## $ mpg  &lt;dbl&gt; 21.0, 21.0, 22.8, 21.4, 18.7, 18.1, 14.3, 24.4, 22.8, 19.2, 17...
## $ cyl  &lt;dbl&gt; 6, 6, 4, 6, 8, 6, 8, 4, 4, 6, 6, 8, 8, 8, 8, 8, 8, 4, 4, 4, 4,...
## $ disp &lt;dbl&gt; 160.0, 160.0, 108.0, 258.0, 360.0, 225.0, 360.0, 146.7, 140.8,...
## $ hp   &lt;dbl&gt; 110, 110, 93, 110, 175, 105, 245, 62, 95, 123, 123, 180, 180, ...
## $ drat &lt;dbl&gt; 3.90, 3.90, 3.85, 3.08, 3.15, 2.76, 3.21, 3.69, 3.92, 3.92, 3....
## $ wt   &lt;dbl&gt; 2.620, 2.875, 2.320, 3.215, 3.440, 3.460, 3.570, 3.190, 3.150,...
## $ qsec &lt;dbl&gt; 16.46, 17.02, 18.61, 19.44, 17.02, 20.22, 15.84, 20.00, 22.90,...
## $ vs   &lt;dbl&gt; 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1,...
## $ am   &lt;dbl&gt; 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0,...
## $ gear &lt;dbl&gt; 4, 4, 4, 3, 3, 3, 3, 4, 4, 4, 4, 3, 3, 3, 3, 3, 3, 4, 4, 4, 3,...
## $ carb &lt;dbl&gt; 4, 4, 1, 1, 2, 1, 4, 2, 2, 4, 4, 3, 3, 3, 4, 4, 4, 1, 2, 1, 1,...</code></pre>
</section>
<section id="check-column" class="level2">
<h2 class="anchored" data-anchor-id="check-column">
Check Column
</h2>
<p>
<code>has_name()</code> can be used to check if a tibble has a specific column.
</p>
<pre class="r"><code>has_name(mtcars, 'cyl')</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>has_name(mtcars, 'gears')</code></pre>
<pre><code>## [1] FALSE</code></pre>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">
Summary
</h2>
<section id="creating-tibbles-1" class="level4">
<h4 class="anchored" data-anchor-id="creating-tibbles-1">
Creating tibbles
</h4>
<ul>
<li>
use <code>tibble()</code> to create tibbles
</li>
<li>
use <code>as_tibble()</code> to coerce other objects to tibble
</li>
<li>
use <code>enframe()</code> to coerce vector to tibble
</li>
<li>
use <code>tribble()</code> to create tibble using data entry
</li>
</ul>
</section>
<section id="modifying-tibbles" class="level4">
<h4 class="anchored" data-anchor-id="modifying-tibbles">
Modifying tibbles
</h4>
<ul>
<li>
use <code>add_row()</code> to add a new row
</li>
<li>
use <code>add_column()</code> to add a new column
</li>
<li>
use <code>remove_rownames()</code> to remove rownames from data
</li>
<li>
use <code>rownames_to_colum()</code> to coerce rowname to first column
</li>
<li>
use <code>column_to_rownames()</code> to coerce first column to rownames
</li>
</ul>
</section>
<section id="testing-tibbles" class="level4">
<h4 class="anchored" data-anchor-id="testing-tibbles">
Testing tibbles
</h4>
<ul>
<li>
use <code>is_tibble()</code> to test if an object is a tibble
</li>
<li>
use <code>has_rownames()</code> to check whether a data set has rownames
</li>
<li>
use <code>has_name()</code> to check if tibble has a specific column
</li>
<li>
use <code>glimpse()</code> to get an overview of data
</li>
</ul>
</section>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://tibble.tidyverse.org/" class="uri">https://tibble.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/tibbles.html" class="uri">http://r4ds.had.co.nz/tibbles.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>tibbles</category>
  <guid>https://blog.rsquaredacademy.com/posts/introduction-to-tibbles/</guid>
  <pubDate>Fri, 28 Sep 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/tibble_intro.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Data Wrangling with dplyr - Part 3</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-3/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-09-16-data-wrangling-with-dplyr-part-3.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous <a href="https://blog.aravindhebbali.com/2017/12/25/data-wrangling-with-dplyr-part-2/">post</a>, we learnt to combine tables using dplyr. In this post, we will explore a set of helper functions in order to:
</p>
<ul>
<li>
extract unique rows
</li>
<li>
rename columns
</li>
<li>
sample data
</li>
<li>
extract columns
</li>
<li>
slice rows
</li>
<li>
arrange rows
</li>
<li>
compare tables
</li>
<li>
extract/mutate data using predicate functions
</li>
<li>
count observations for different levels of a variable
</li>
</ul>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
<a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The data sets can be downloaded from <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes from <a href="https://gist.github.com/aravindhebbali/55c4f40476028c09949b73af97bb1619">here</a>.
</p>
<pre class="r"><code>library(dplyr)
library(readr)</code></pre>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<p>
Let us look at a case study (e-commerce data) and see how we can use dplyr helper functions to answer questions we have about and to modify/transform the underlying data set.
</p>
<section id="data" class="level3">
<h3 class="anchored" data-anchor-id="data">
Data
</h3>
<pre class="r"><code>ecom &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web.csv',
    col_types = cols_only(device = col_factor(levels = c("laptop", "tablet", "mobile")),
      referrer = col_factor(levels = c("bing", "direct", "social", "yahoo", "google")),
      purchase = col_logical(), bouncers = col_logical(), duration = col_double(),
      n_visit = col_double(), n_pages = col_double()
    )
  )

ecom</code></pre>
<pre><code>## # A tibble: 1,000 x 7
##    referrer device bouncers n_visit n_pages duration purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 google   laptop TRUE          10       1      693 FALSE   
##  2 yahoo    tablet TRUE           9       1      459 FALSE   
##  3 direct   laptop TRUE           0       1      996 FALSE   
##  4 bing     tablet FALSE          3      18      468 TRUE    
##  5 yahoo    mobile TRUE           9       1      955 FALSE   
##  6 yahoo    laptop FALSE          5       5      135 FALSE   
##  7 yahoo    mobile TRUE          10       1       75 FALSE   
##  8 direct   mobile TRUE          10       1      908 FALSE   
##  9 bing     mobile FALSE          3      19      209 FALSE   
## 10 google   mobile TRUE           6       1      208 FALSE   
## # ... with 990 more rows</code></pre>
</section>
<section id="data-dictionary" class="level3">
<h3 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h3>
<ul>
<li>
referrer: referrer website/search engine
</li>
<li>
device: device used to visit the website
</li>
<li>
bouncers: whether a visit bounced (exited from landing page)
</li>
<li>
duration: time spent on the website (in seconds)
</li>
<li>
purchase: whether visitor purchased
</li>
<li>
n_visit: number of visits
</li>
<li>
n_pages: number of pages visited/browsed
</li>
</ul>
</section>
</section>
<section id="data-sanitization" class="level2">
<h2 class="anchored" data-anchor-id="data-sanitization">
Data Sanitization
</h2>
<p>
Let us ensure that the data is sanitized by checking the sources of traffic and devices used to visit the site. We will use <code>distinct</code> to examine the values in the <code>referrer</code> column
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/distinct_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>distinct(ecom, referrer)</code></pre>
<pre><code>## # A tibble: 5 x 1
##   referrer
##   &lt;fct&gt;   
## 1 google  
## 2 yahoo   
## 3 direct  
## 4 bing    
## 5 social</code></pre>
<p>
and the <code>device</code> column as well.
</p>
<pre class="r"><code>distinct(ecom, device)</code></pre>
<pre><code>## # A tibble: 3 x 1
##   device
##   &lt;fct&gt; 
## 1 laptop
## 2 tablet
## 3 mobile</code></pre>
</section>
<section id="rename-columns" class="level2">
<h2 class="anchored" data-anchor-id="rename-columns">
Rename Columns
</h2>
<p>
Columns can be renamed using <code>rename()</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/rename_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>rename(ecom, time_on_site = duration)</code></pre>
<pre><code>## # A tibble: 1,000 x 7
##    referrer device bouncers n_visit n_pages time_on_site purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;        &lt;dbl&gt; &lt;lgl&gt;   
##  1 google   laptop TRUE          10       1          693 FALSE   
##  2 yahoo    tablet TRUE           9       1          459 FALSE   
##  3 direct   laptop TRUE           0       1          996 FALSE   
##  4 bing     tablet FALSE          3      18          468 TRUE    
##  5 yahoo    mobile TRUE           9       1          955 FALSE   
##  6 yahoo    laptop FALSE          5       5          135 FALSE   
##  7 yahoo    mobile TRUE          10       1           75 FALSE   
##  8 direct   mobile TRUE          10       1          908 FALSE   
##  9 bing     mobile FALSE          3      19          209 FALSE   
## 10 google   mobile TRUE           6       1          208 FALSE   
## # ... with 990 more rows</code></pre>
</section>
<section id="data-tabulation" class="level2">
<h2 class="anchored" data-anchor-id="data-tabulation">
Data Tabulation
</h2>
<p>
Let us now look at the proportion or share of visits driven by different sources of traffic.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/tally_count.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>ecom %&gt;%
  group_by(referrer) %&gt;%
  tally()</code></pre>
<pre><code>## # A tibble: 5 x 2
##   referrer     n
## * &lt;fct&gt;    &lt;int&gt;
## 1 bing       194
## 2 direct     191
## 3 social     200
## 4 yahoo      207
## 5 google     208</code></pre>
<p>
We would also like to know the number of bouncers driven by the different sources of traffic.
</p>
<pre class="r"><code>ecom %&gt;%
  group_by(referrer, bouncers) %&gt;%
  tally()</code></pre>
<pre><code>## # A tibble: 10 x 3
## # Groups:   referrer [5]
##    referrer bouncers     n
##    &lt;fct&gt;    &lt;lgl&gt;    &lt;int&gt;
##  1 bing     FALSE      104
##  2 bing     TRUE        90
##  3 direct   FALSE       98
##  4 direct   TRUE        93
##  5 social   FALSE       93
##  6 social   TRUE       107
##  7 yahoo    FALSE      110
##  8 yahoo    TRUE        97
##  9 google   FALSE      101
## 10 google   TRUE       107</code></pre>
<p>
Let us look at how many conversions happen across different devices.
</p>
<pre class="r"><code>ecom %&gt;%
  group_by(device, purchase) %&gt;%
  tally() %&gt;%
  filter(purchase)</code></pre>
<pre><code>## # A tibble: 3 x 3
## # Groups:   device [3]
##   device purchase     n
##   &lt;fct&gt;  &lt;lgl&gt;    &lt;int&gt;
## 1 laptop TRUE        31
## 2 tablet TRUE        36
## 3 mobile TRUE        36</code></pre>
<p>
Another way to extract the above information is by using <code>count</code>
</p>
<pre class="r"><code>ecom %&gt;%
  count(referrer, purchase) %&gt;%
  filter(purchase)</code></pre>
<pre><code>## # A tibble: 5 x 3
##   referrer purchase     n
##   &lt;fct&gt;    &lt;lgl&gt;    &lt;int&gt;
## 1 bing     TRUE        17
## 2 direct   TRUE        25
## 3 social   TRUE        20
## 4 yahoo    TRUE        22
## 5 google   TRUE        19</code></pre>
</section>
<section id="sampling-data" class="level2">
<h2 class="anchored" data-anchor-id="sampling-data">
Sampling Data
</h2>
<p>
dplyr offers sampling functions which allow us to specify either the number or percentage of observations. <code>sample_n()</code> allows sampling a specific number of observations.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/sample_frac_n.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>sample_n(ecom, 700)</code></pre>
<pre><code>## # A tibble: 700 x 7
##    referrer device bouncers n_visit n_pages duration purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 yahoo    laptop FALSE          2       9      162 FALSE   
##  2 yahoo    tablet FALSE          2      14      364 TRUE    
##  3 social   laptop TRUE           1       1      111 FALSE   
##  4 direct   laptop TRUE           4       1      896 FALSE   
##  5 yahoo    mobile FALSE          5       8       80 FALSE   
##  6 social   laptop TRUE           0       1      720 FALSE   
##  7 bing     mobile TRUE           5       1      190 FALSE   
##  8 direct   mobile TRUE           2       1      501 FALSE   
##  9 yahoo    tablet TRUE           1       1      605 FALSE   
## 10 bing     laptop TRUE           1       1      169 FALSE   
## # ... with 690 more rows</code></pre>
<p>
We can combine the sampling functions with other dplyr functions as shown below where we sample observation after grouping them according to the source of traffic.
</p>
<pre class="r"><code>ecom %&gt;%
  group_by(referrer) %&gt;%
  sample_n(100)</code></pre>
<pre><code>## # A tibble: 500 x 7
## # Groups:   referrer [5]
##    referrer device bouncers n_visit n_pages duration purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 bing     mobile FALSE          3       8      120 FALSE   
##  2 bing     laptop FALSE          9      13      299 FALSE   
##  3 bing     tablet FALSE          2      17      510 FALSE   
##  4 bing     laptop TRUE           9       1      709 FALSE   
##  5 bing     tablet TRUE           0       1      845 FALSE   
##  6 bing     tablet TRUE           1       1      721 FALSE   
##  7 bing     tablet TRUE           0       1      425 FALSE   
##  8 bing     mobile FALSE          0       7      196 FALSE   
##  9 bing     tablet TRUE           4       1      493 FALSE   
## 10 bing     mobile TRUE           6       1      604 FALSE   
## # ... with 490 more rows</code></pre>
<p>
<code>sample_frac()</code> allows a specific percentage of observations.
</p>
<pre class="r"><code>sample_frac(ecom, size = 0.7)</code></pre>
<pre><code>## # A tibble: 700 x 7
##    referrer device bouncers n_visit n_pages duration purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 google   mobile TRUE           0       1      132 FALSE   
##  2 yahoo    tablet TRUE           7       1      889 FALSE   
##  3 bing     mobile FALSE          7       1       22 FALSE   
##  4 google   laptop TRUE           5       1      376 FALSE   
##  5 social   tablet FALSE         10      11      330 TRUE    
##  6 social   mobile FALSE          7       8      216 FALSE   
##  7 bing     mobile FALSE          1      12      168 FALSE   
##  8 bing     tablet TRUE           7       1      489 FALSE   
##  9 social   laptop TRUE           0       1      581 FALSE   
## 10 direct   laptop FALSE          6       6       96 FALSE   
## # ... with 690 more rows</code></pre>
</section>
<section id="data-extraction" class="level2">
<h2 class="anchored" data-anchor-id="data-extraction">
Data Extraction
</h2>
<p>
In the first <a href="https://blog.aravindhebbali.com/2017/12/25/data-wrangling-with-dplyr-part-1/">post</a>, we had observed that dplyr verbs always returned a tibble. What if you want to extract a specific column or a bunch of rows but not as a tibble?
</p>
<p>
Use <code>pull</code> to extract columns either by name or position. It will return a vector. In the below example, we extract the <code>device</code> column as a vector. I am using <code>head</code> in addition to limit the output printed.
</p>
<section id="sample-data" class="level3">
<h3 class="anchored" data-anchor-id="sample-data">
Sample Data
</h3>
<pre class="r"><code>ecom_mini &lt;- sample_n(ecom, size = 10)</code></pre>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/pull_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>pull(ecom_mini, device)</code></pre>
<pre><code>##  [1] laptop mobile mobile laptop tablet laptop mobile laptop tablet tablet
## Levels: laptop tablet mobile</code></pre>
<p>
Let us extract the first column from <code>ecom</code> using column position instead of name.
</p>
<pre class="r"><code>pull(ecom_mini, 1) </code></pre>
<pre><code>##  [1] bing   google yahoo  direct bing   direct bing   bing   bing   bing  
## Levels: bing direct social yahoo google</code></pre>
<p>
You can use <code>-</code> before the column position to indicate the position in reverse. The below example extracts data from the last column.
</p>
<pre class="r"><code>pull(ecom_mini, -1) </code></pre>
<pre><code>##  [1] FALSE FALSE FALSE  TRUE FALSE FALSE FALSE FALSE FALSE FALSE</code></pre>
<p>
Let us now look at extracting rows using <code>slice()</code>. In the below example, we extract data starting from the 5th row and upto the 15th row.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/slice_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>slice(ecom, 5:15)</code></pre>
<pre><code>## # A tibble: 11 x 7
##    referrer device bouncers n_visit n_pages duration purchase
##    &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 yahoo    mobile TRUE           9       1      955 FALSE   
##  2 yahoo    laptop FALSE          5       5      135 FALSE   
##  3 yahoo    mobile TRUE          10       1       75 FALSE   
##  4 direct   mobile TRUE          10       1      908 FALSE   
##  5 bing     mobile FALSE          3      19      209 FALSE   
##  6 google   mobile TRUE           6       1      208 FALSE   
##  7 direct   laptop TRUE           9       1      738 FALSE   
##  8 direct   tablet FALSE          6      12      132 FALSE   
##  9 direct   mobile FALSE          9      14      406 TRUE    
## 10 yahoo    tablet FALSE          5       8       80 FALSE   
## 11 yahoo    mobile FALSE          7       1       19 FALSE</code></pre>
<p>
Use <code>n()</code> inside <code>slice()</code> to extract the last row.
</p>
<pre class="r"><code>slice(ecom, n())</code></pre>
<pre><code>## # A tibble: 1 x 7
##   referrer device bouncers n_visit n_pages duration purchase
##   &lt;fct&gt;    &lt;fct&gt;  &lt;lgl&gt;      &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
## 1 google   mobile TRUE           9       1      269 FALSE</code></pre>
</section>
</section>
<section id="between" class="level2">
<h2 class="anchored" data-anchor-id="between">
Between
</h2>
<p>
<code>between()</code> allows us to test if the values in a column lie between two specific values. In the below example, we check how many visits browsed pages between 5 and 15.
</p>
<pre class="r"><code>ecom_sample &lt;- sample_n(ecom, 30)
  
ecom_sample %&gt;%
  pull(n_pages) %&gt;%
  between(5, 15) </code></pre>
<pre><code>##  [1] FALSE FALSE FALSE  TRUE  TRUE FALSE FALSE FALSE FALSE  TRUE  TRUE FALSE
## [13]  TRUE  TRUE FALSE  TRUE FALSE FALSE FALSE FALSE FALSE FALSE FALSE  TRUE
## [25] FALSE FALSE FALSE FALSE FALSE FALSE</code></pre>
</section>
<section id="case-when" class="level2">
<h2 class="anchored" data-anchor-id="case-when">
Case When
</h2>
<p>
<code>case_when()</code> is an alternative to <code>if else</code>. It allows us to lay down the conditions clearly and makes the code more readable. In the below example, we create a new column <code>repeat_visit</code> from <code>n_visit</code> (the number of previous visits).
</p>
<pre class="r"><code>ecom %&gt;%
  mutate(
    repeat_visit = case_when(
      n_visit &gt; 0 ~ TRUE,
      TRUE ~ FALSE
    )
  ) %&gt;%
  select(n_visit, repeat_visit) </code></pre>
<pre><code>## # A tibble: 1,000 x 2
##    n_visit repeat_visit
##      &lt;dbl&gt; &lt;lgl&gt;       
##  1      10 TRUE        
##  2       9 TRUE        
##  3       0 FALSE       
##  4       3 TRUE        
##  5       9 TRUE        
##  6       5 TRUE        
##  7      10 TRUE        
##  8      10 TRUE        
##  9       3 TRUE        
## 10       6 TRUE        
## # ... with 990 more rows</code></pre>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://dplyr.tidyverse.org/" class="uri">https://dplyr.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/transform.html" class="uri">http://r4ds.had.co.nz/transform.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>dplyr</category>
  <guid>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-3/</guid>
  <pubDate>Sun, 16 Sep 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/dplyr.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Data Wrangling with dplyr - Part 2</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-2/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-09-04-data-wrangling-with-dplyr-part-2.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post we learnt about dplyr verbs and used them to compute average order value for an online retail company data. In this post, we will learn to combine tables using different <code>*_join</code> functions provided in dplyr.
</p>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
<a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The data sets can be downloaded from <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes from <a href="https://gist.github.com/aravindhebbali/3e31f13a4194a8f1003034aa7d70d5ee">here</a>.
</p>
<pre class="r"><code>library(dplyr)
library(readr)
options(tibble.width = Inf)</code></pre>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<p>
For our case study, we will use two data sets. The first one, <code>order</code>, contains details of orders placed by different customers. The second data set, <code>customer</code> contains details of each customer. The below table displays the details stored in each data set.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/join_data.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
Let us import both the data sets using <code>read_csv</code>.
</p>
<section id="data-orders" class="level3">
<h3 class="anchored" data-anchor-id="data-orders">
Data: Orders
</h3>
<pre class="r"><code>order &lt;- read_delim('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/order.csv', delim = ';')
order</code></pre>
<pre><code>## # A tibble: 300 x 3
##       id order_date amount
##    &lt;dbl&gt; &lt;chr&gt;       &lt;dbl&gt;
##  1   368 7/2/2016     365.
##  2   286 11/2/2016   2064.
##  3    28 2/22/2017    432.
##  4   309 3/5/2017     480.
##  5     2 12/28/2016   235.
##  6    31 12/30/2016  2745.
##  7   179 12/21/2016  2358.
##  8   484 11/24/2016  1031.
##  9   115 9/9/2016    1218.
## 10   340 5/6/2017    1184.
## # ... with 290 more rows</code></pre>
</section>
<section id="data-customers" class="level3">
<h3 class="anchored" data-anchor-id="data-customers">
Data: Customers
</h3>
<pre class="r"><code>customer &lt;- read_delim('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/customer.csv', delim = ';')
customer</code></pre>
<pre><code>## # A tibble: 91 x 3
##       id first_name city      
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;     
##  1     1 Elbertine  California
##  2     2 Marcella   Colorado  
##  3     3 Daria      Florida   
##  4     4 Sherilyn   Distric...
##  5     5 Ketty      Texas     
##  6     6 Jethro     California
##  7     7 Jeremiah   California
##  8     8 Constancia Texas     
##  9     9 Muire      Idaho     
## 10    10 Abigail    Texas     
## # ... with 81 more rows</code></pre>
<p>
We will explore the following in the case study:
</p>
<ul>
<li>
details of customers who have placed orders and their order details
</li>
<li>
details of customers and their orders irrespective of whether a customer has placed orders or not
</li>
<li>
customer details for each order
</li>
<li>
details of customers who have placed orders
</li>
<li>
details of customers who have not placed orders
</li>
<li>
details of all customers and all orders
</li>
</ul>
</section>
</section>
<section id="example-data" class="level2">
<h2 class="anchored" data-anchor-id="example-data">
Example Data
</h2>
<p>
We will use another data set to illustrate how the different joins work. You can view the example data sets below.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/join.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
</section>
<section id="inner-join" class="level2">
<h2 class="anchored" data-anchor-id="inner-join">
Inner Join
</h2>
<p>
<br>
</p>
<p>
Inner join return all rows from <code>Age</code> where there are matching values in <code>Height</code>, and all columns from <code>Age</code> and <code>Height</code>. If there are multiple matches between <code>Age</code> and <code>Height</code>, all combination of the matches are returned.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_inner.png" width="100%" style="display: block; margin: auto;">
</p>
<section id="case-study-details-of-customers-who-have-placed-orders-and-their-order-details" class="level4">
<h4 class="anchored" data-anchor-id="case-study-details-of-customers-who-have-placed-orders-and-their-order-details">
Case Study: Details of customers who have placed orders and their order details
</h4>
<p>
To get data for all those customers who have placed orders in the past let us join the <code>order</code> data with the <code>customer</code> data using <code>inner_join</code>.
</p>
<pre class="r"><code>inner_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 55 x 5
##       id first_name city       order_date amount
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;      &lt;chr&gt;       &lt;dbl&gt;
##  1     2 Marcella   Colorado   12/28/2016   235.
##  2     2 Marcella   Colorado   8/31/2016   1150.
##  3     5 Ketty      Texas      1/17/2017    346.
##  4     6 Jethro     California 1/27/2017   2317.
##  5     7 Jeremiah   California 6/21/2016    136.
##  6     7 Jeremiah   California 2/13/2017   1407.
##  7     7 Jeremiah   California 7/8/2016    1914.
##  8     8 Constancia Texas      11/5/2016   2461.
##  9     8 Constancia Texas      5/19/2017   2714.
## 10     9 Muire      Idaho      12/28/2016   187.
## # ... with 45 more rows</code></pre>
</section>
</section>
<section id="left-join" class="level2">
<h2 class="anchored" data-anchor-id="left-join">
Left Join
</h2>
<p>
<br>
</p>
<p>
Left join return all rows from <code>Age</code>, and all columns from <code>Age</code> and <code>Height</code>. Rows in <code>Age</code> with no match in <code>Height</code> will have NA values in the new columns. If there are multiple matches between <code>Age</code> and <code>Height</code>, all combinations of the matches are returned.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_left.png" width="100%" style="display: block; margin: auto;">
</p>
<section id="case-study-details-of-customers-and-their-orders-irrespective-of-whether-a-customer-has" class="level4">
<h4 class="anchored" data-anchor-id="case-study-details-of-customers-and-their-orders-irrespective-of-whether-a-customer-has">
Case Study: Details of customers and their orders irrespective of whether a customer has
</h4>
<p>
placed orders or not.
</p>
<p>
To get data for all those customers and their orders irrespective of whether a customer has placed orders or not let us join the <code>order</code> data with the <code>customer</code> data using <code>left_join</code>.
</p>
<pre class="r"><code>left_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 104 x 5
##       id first_name city       order_date amount
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;      &lt;chr&gt;       &lt;dbl&gt;
##  1     1 Elbertine  California &lt;NA&gt;          NA 
##  2     2 Marcella   Colorado   12/28/2016   235.
##  3     2 Marcella   Colorado   8/31/2016   1150.
##  4     3 Daria      Florida    &lt;NA&gt;          NA 
##  5     4 Sherilyn   Distric... &lt;NA&gt;          NA 
##  6     5 Ketty      Texas      1/17/2017    346.
##  7     6 Jethro     California 1/27/2017   2317.
##  8     7 Jeremiah   California 6/21/2016    136.
##  9     7 Jeremiah   California 2/13/2017   1407.
## 10     7 Jeremiah   California 7/8/2016    1914.
## # ... with 94 more rows</code></pre>
</section>
</section>
<section id="right-join" class="level2">
<h2 class="anchored" data-anchor-id="right-join">
Right Join
</h2>
<p>
<br>
</p>
<p>
Right join return all rows from <code>Height</code>, and all columns from <code>Age</code> and <code>Height</code>. Rows in <code>Height</code> with no match in <code>Age</code> will have NA values in the new columns. If there are multiple matches between <code>Age</code> and <code>Height</code>, all combinations of the matches are returned.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_right.png" width="100%" style="display: block; margin: auto;">
</p>
<section id="case-study-customer-details-for-each-order" class="level4">
<h4 class="anchored" data-anchor-id="case-study-customer-details-for-each-order">
Case Study: Customer details for each order
</h4>
<p>
To get customer data for all orders, let us join the <code>order</code> data with the <code>customer</code> data using <code>right_join</code>.
</p>
<pre class="r"><code>right_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 300 x 5
##       id first_name city       order_date amount
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;      &lt;chr&gt;       &lt;dbl&gt;
##  1     2 Marcella   Colorado   12/28/2016   235.
##  2     2 Marcella   Colorado   8/31/2016   1150.
##  3     5 Ketty      Texas      1/17/2017    346.
##  4     6 Jethro     California 1/27/2017   2317.
##  5     7 Jeremiah   California 6/21/2016    136.
##  6     7 Jeremiah   California 2/13/2017   1407.
##  7     7 Jeremiah   California 7/8/2016    1914.
##  8     8 Constancia Texas      11/5/2016   2461.
##  9     8 Constancia Texas      5/19/2017   2714.
## 10     9 Muire      Idaho      12/28/2016   187.
## # ... with 290 more rows</code></pre>
</section>
</section>
<section id="semi-join" class="level2">
<h2 class="anchored" data-anchor-id="semi-join">
Semi Join
</h2>
<p>
<br>
</p>
<p>
Semi join return all rows from <code>Age</code> where there are matching values in <code>Height</code>, keeping just columns from <code>Age</code>. A semi join differs from an inner join because an inner join will return one row of <code>Age</code> for each matching row of <code>Height</code>, where a semi join will never duplicate rows of <code>Age</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_semi.png" width="100%" style="display: block; margin: auto;">
</p>
<section id="case-study-details-of-customers-who-have-placed-orders" class="level4">
<h4 class="anchored" data-anchor-id="case-study-details-of-customers-who-have-placed-orders">
Case Study: Details of customers who have placed orders
</h4>
<p>
To get customer data for all orders where customer data exists, let us join the <code>order</code> data with the <code>customer</code> data using <code>semi_join</code>. You can observe that data is returned only for those cases where customer data is present.
</p>
<pre class="r"><code>semi_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 42 x 3
##       id first_name city      
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;     
##  1     2 Marcella   Colorado  
##  2     5 Ketty      Texas     
##  3     6 Jethro     California
##  4     7 Jeremiah   California
##  5     8 Constancia Texas     
##  6     9 Muire      Idaho     
##  7    15 Valentijn  California
##  8    16 Monique    Missouri  
##  9    20 Colette    Texas     
## 10    28 Avrit      Texas     
## # ... with 32 more rows</code></pre>
</section>
</section>
<section id="anti-join" class="level2">
<h2 class="anchored" data-anchor-id="anti-join">
Anti Join
</h2>
<p>
<br>
</p>
<p>
Anti join return all rows from <code>Age</code> where there are not matching values in <code>Height</code>, keeping just columns from <code>Age</code>.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_anti.png" width="100%" style="display: block; margin: auto;">
</p>
<section id="case-study-details-of-customers-who-have-not-placed-orders" class="level4">
<h4 class="anchored" data-anchor-id="case-study-details-of-customers-who-have-not-placed-orders">
Case Study: Details of customers who have not placed orders
</h4>
<p>
To get details of customers who have not placed orders, let us join the <code>order</code> data with the <code>customer</code> data using <code>anti_join</code>.
</p>
<pre class="r"><code>anti_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 49 x 3
##       id first_name city      
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;     
##  1     1 Elbertine  California
##  2     3 Daria      Florida   
##  3     4 Sherilyn   Distric...
##  4    10 Abigail    Texas     
##  5    11 Wynne      Georgia   
##  6    12 Pietra     Minnesota 
##  7    13 Bram       Iowa      
##  8    14 Rees       New York  
##  9    17 Orazio     Louisiana 
## 10    18 Mason      Texas     
## # ... with 39 more rows</code></pre>
</section>
</section>
<section id="full-join" class="level2">
<h2 class="anchored" data-anchor-id="full-join">
Full Join
</h2>
<p>
<br>
</p>
<p>
Full join return all rows and all columns from both <code>Age</code> and <code>Height</code>. Where there are not matching values, returns NA for the one missing.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/draw_full.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<section id="case-study-details-of-all-customers-and-all-orders" class="level4">
<h4 class="anchored" data-anchor-id="case-study-details-of-all-customers-and-all-orders">
Case Study: Details of all customers and all orders
</h4>
<p>
To get details of all customers and all orders, let us join the <code>order</code> data with the <code>customer</code> data using <code>full_join</code>.
</p>
<pre class="r"><code>full_join(customer, order, by = "id")</code></pre>
<pre><code>## # A tibble: 349 x 5
##       id first_name city       order_date amount
##    &lt;dbl&gt; &lt;chr&gt;      &lt;chr&gt;      &lt;chr&gt;       &lt;dbl&gt;
##  1     1 Elbertine  California &lt;NA&gt;          NA 
##  2     2 Marcella   Colorado   12/28/2016   235.
##  3     2 Marcella   Colorado   8/31/2016   1150.
##  4     3 Daria      Florida    &lt;NA&gt;          NA 
##  5     4 Sherilyn   Distric... &lt;NA&gt;          NA 
##  6     5 Ketty      Texas      1/17/2017    346.
##  7     6 Jethro     California 1/27/2017   2317.
##  8     7 Jeremiah   California 6/21/2016    136.
##  9     7 Jeremiah   California 2/13/2017   1407.
## 10     7 Jeremiah   California 7/8/2016    1914.
## # ... with 339 more rows</code></pre>
</section>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://dplyr.tidyverse.org/" class="uri">https://dplyr.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/relational-data.html" class="uri">http://r4ds.had.co.nz/relational-data.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>dplyr</category>
  <guid>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-2/</guid>
  <pubDate>Tue, 04 Sep 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/dplyr.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Data Wrangling with dplyr - Part 1</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-1/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-08-23-data-wrangling-with-dplyr-part-1.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
According to a <a href="http://visit.crowdflower.com/rs/416-ZBE-142/images/CrowdFlower_DataScienceReport_2016.pdf">survey</a> by <a href="https://www.crowdflower.com/">CrowdFlower</a>, data scientists spend most of their time cleaning and manipulating data rather than mining or modeling them for insights. As such, it becomes important to have tools that make data manipulation faster and easier. In today’s post, we introduce you to <a href="http://dplyr.tidyverse.org/">dplyr</a>, a grammar of data manipulation.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/crowd_flower.png" width="70%" style="display: block; margin: auto;">
</p>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following libraries:
</p>
<ul>
<li>
<a href="http://dplyr.tidyverse.org/index.html">dplyr</a>
</li>
<li>
and <a href="http://readr.tidyverse.org/index.html">readr</a>
</li>
</ul>
<p>
The data sets can be downloaded from <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes from <a href="https://gist.github.com/aravindhebbali/7758b86c2bc13ff1e5d88d9d1c204f8c">here</a>.
</p>
<pre class="r"><code>library(dplyr)
library(readr)</code></pre>
</section>
<section id="dplyr-verbs" class="level2">
<h2 class="anchored" data-anchor-id="dplyr-verbs">
dplyr Verbs
</h2>
<p>
dplyr provides a set of verbs that help us solve the most common data manipulation challenges while working with tabular data (dataframes, tibbles):
</p>
<ul>
<li>
<code>select</code>
</li>
<li>
<code>filter</code>
</li>
<li>
<code>arrange</code>
</li>
<li>
<code>mutate</code>
</li>
<li>
<code>summarise</code>
</li>
</ul>
</section>
<section id="data" class="level2">
<h2 class="anchored" data-anchor-id="data">
Data
</h2>
<pre class="r"><code>ecom &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web.csv',
    col_types = cols_only(device = col_factor(levels = c("laptop", "tablet", "mobile")),
      referrer = col_factor(levels = c("bing", "direct", "social", "yahoo", "google")),
      purchase = col_logical(), n_pages = col_double(), n_visit = col_double(), 
      duration = col_double(), order_value = col_double(), order_items = col_double()
    )
  )

ecom</code></pre>
<pre><code>## # A tibble: 1,000 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 google   laptop      10       1      693 FALSE              0           0
##  2 yahoo    tablet       9       1      459 FALSE              0           0
##  3 direct   laptop       0       1      996 FALSE              0           0
##  4 bing     tablet       3      18      468 TRUE               6         434
##  5 yahoo    mobile       9       1      955 FALSE              0           0
##  6 yahoo    laptop       5       5      135 FALSE              0           0
##  7 yahoo    mobile      10       1       75 FALSE              0           0
##  8 direct   mobile      10       1      908 FALSE              0           0
##  9 bing     mobile       3      19      209 FALSE              0           0
## 10 google   mobile       6       1      208 FALSE              0           0
## # ... with 990 more rows</code></pre>
<section id="data-dictionary" class="level6">
<h6 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h6>
<p>
Below is the description of the data set:
</p>
<ul>
<li>
referrer: referrer website/search engine
</li>
<li>
device: device used to visit the website
</li>
<li>
n_pages: number of pages visited
</li>
<li>
duration: time spent on the website (in seconds)
</li>
<li>
purchase: whether visitor purchased
</li>
<li>
order_value: order value of visitor (in dollars)
</li>
<li>
n_visit: number of visits
</li>
</ul>
</section>
</section>
<section id="case-study" class="level2">
<h2 class="anchored" data-anchor-id="case-study">
Case Study
</h2>
<p>
We will use dplyr to answer the following:
</p>
<ul>
<li>
what is the average order value by device types?
</li>
<li>
what is the average number of pages visited by purchasers and non-purchasers?
</li>
<li>
what is the average time on site for purchasers vs non-purchasers?
</li>
<li>
what is the average number of pages visited by purchasers and non-purchasers using mobile?
</li>
</ul>
</section>
<section id="average-order-value" class="level2">
<h2 class="anchored" data-anchor-id="average-order-value">
Average Order Value
</h2>
<p>
<img src="https://blog.rsquaredacademy.com/img/image.jpg" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="aov-by-devices" class="level2">
<h2 class="anchored" data-anchor-id="aov-by-devices">
AOV by Devices
</h2>
<pre class="r"><code>ecom %&gt;%
  filter(purchase) %&gt;%
  select(device, order_value) %&gt;%
  group_by(device) %&gt;%
  summarise_all(funs(revenue = sum, orders = n())) %&gt;%
  mutate(
    aov = revenue / orders
  ) %&gt;%
  select(device, aov) %&gt;%
  arrange(aov)</code></pre>
<pre><code>## Warning: `funs()` is deprecated as of dplyr 0.8.0.
## Please use a list of either functions or lambdas: 
## 
##   # Simple named list: 
##   list(mean = mean, median = median)
## 
##   # Auto named with `tibble::lst()`: 
##   tibble::lst(mean, median)
## 
##   # Using lambdas
##   list(~ mean(., trim = .2), ~ median(., na.rm = TRUE))
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_warnings()` to see where this warning was generated.</code></pre>
<pre><code>## # A tibble: 3 x 2
##   device   aov
##   &lt;fct&gt;  &lt;dbl&gt;
## 1 tablet 1426.
## 2 mobile 1431.
## 3 laptop 1824.</code></pre>
</section>
<section id="syntax" class="level2">
<h2 class="anchored" data-anchor-id="syntax">
Syntax
</h2>
<p>
Before we start exploring the dplyr verbs, let us look at their syntax:
</p>
<ul>
<li>
the first argument is always a <code>data.frame</code> or <code>tibble</code>
</li>
<li>
the subsequent arguments provide the information required for the verbs to take action
</li>
<li>
the name of columns in the data need not be surrounded by quotes
</li>
</ul>
</section>
<section id="filter-rows" class="level2">
<h2 class="anchored" data-anchor-id="filter-rows">
Filter Rows
</h2>
<p>
In order to compute the AOV, we must first separate the purchasers from non-purchasers. We will do this by filtering the data related to purchasers using the <code>filter()</code> function. It allows us to filter rows that meet a specific criteria/condition. The first argument is the name of the data frame and the rest of the arguments are expressions for filtering the data. Let us look at a few examples:
</p>
<p>
The first example we will look at filters all visits from device <strong>mobile</strong>. As we had learnt in the previous section, the first argument is our data set <code>ecom</code> and the next argument is the condition for filtering rows.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/filter_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>filter(ecom, device == "mobile")</code></pre>
<pre><code>## # A tibble: 344 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 yahoo    mobile       9       1      955 FALSE              0           0
##  2 yahoo    mobile      10       1       75 FALSE              0           0
##  3 direct   mobile      10       1      908 FALSE              0           0
##  4 bing     mobile       3      19      209 FALSE              0           0
##  5 google   mobile       6       1      208 FALSE              0           0
##  6 direct   mobile       9      14      406 TRUE               3         651
##  7 yahoo    mobile       7       1       19 FALSE              7        2423
##  8 google   mobile       5       1      147 FALSE              0           0
##  9 bing     mobile       0       7      196 FALSE              4         237
## 10 google   mobile      10       1      338 FALSE              0           0
## # ... with 334 more rows</code></pre>
<p>
We can specify multiple filtering conditions as well. In the below example, we specify two filter conditions:
</p>
<ul>
<li>
visit from device <strong>mobile</strong>
</li>
<li>
resulted in a purchase or conversion
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/filter_2.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>filter(ecom, device == "mobile", purchase)</code></pre>
<pre><code>## # A tibble: 36 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 direct   mobile       9      14      406 TRUE               3         651
##  2 bing     mobile       4      20      440 TRUE               3         184
##  3 bing     mobile       3      18      288 TRUE               6         764
##  4 social   mobile      10      11      242 TRUE               4         287
##  5 yahoo    mobile       6      14      322 TRUE               3        1443
##  6 google   mobile       1      18      252 TRUE               3        2449
##  7 social   mobile       7      16      352 TRUE              10        2824
##  8 direct   mobile       4      18      324 TRUE               3        1670
##  9 social   mobile       1      20      520 TRUE               5        1021
## 10 yahoo    mobile       0      13      351 TRUE              10         288
## # ... with 26 more rows</code></pre>
<p>
Here is another example where we specify multiple conditions:
</p>
<ul>
<li>
visit from device <strong>tablet</strong>
</li>
<li>
made a purchase
</li>
<li>
browsed less than 15 pages
</li>
</ul>
<pre class="r"><code>filter(ecom, device == "tablet", purchase, n_pages &lt; 15)</code></pre>
<pre><code>## # A tibble: 12 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 social   tablet       7      10      290 TRUE               9        1304
##  2 yahoo    tablet       2      14      364 TRUE               6        1667
##  3 google   tablet       7      12      324 TRUE               2        1358
##  4 direct   tablet       3      12      324 TRUE              10        1257
##  5 yahoo    tablet       0      13      390 TRUE               5        1748
##  6 social   tablet       2      12      300 TRUE               2        2754
##  7 direct   tablet       6      13      338 TRUE               5         683
##  8 yahoo    tablet       2      10      280 TRUE               4         293
##  9 social   tablet      10      10      290 TRUE               9          37
## 10 direct   tablet       3      10      260 TRUE               7         980
## 11 google   tablet       9      14      308 TRUE               7        2436
## 12 social   tablet      10      11      330 TRUE               1        2171</code></pre>
<section id="case-study-1" class="level5">
<h5 class="anchored" data-anchor-id="case-study-1">
Case Study
</h5>
<p>
Let us apply what we have learnt to the case study. We want to filter all visits that resulted in a purchase.
</p>
<pre class="r"><code>filter(ecom, purchase)</code></pre>
<pre><code>## # A tibble: 103 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 bing     tablet       3      18      468 TRUE               6         434
##  2 direct   mobile       9      14      406 TRUE               3         651
##  3 bing     tablet       5      16      368 TRUE               6        1049
##  4 social   tablet       7      10      290 TRUE               9        1304
##  5 direct   tablet       2      19      342 TRUE               5         622
##  6 social   tablet       9      20      420 TRUE               7        1613
##  7 bing     mobile       4      20      440 TRUE               3         184
##  8 yahoo    tablet       2      16      480 TRUE               9         286
##  9 bing     mobile       3      18      288 TRUE               6         764
## 10 yahoo    tablet       2      14      364 TRUE               6        1667
## # ... with 93 more rows</code></pre>
</section>
</section>
<section id="select-columns" class="level2">
<h2 class="anchored" data-anchor-id="select-columns">
Select Columns
</h2>
<p>
After filtering the data, we need to select relevent variables to compute the AOV. Remember, we do not need all the columns in the data to compute a required metric (in our case, AOV). The <code>select()</code> function allows us to select a subset of columns. The first argument is the name of the data frame and the subsequent arguments specify the columns by name or position.
</p>
<p>
To select the <code>device</code> and <code>duration</code> columns, we specify the data set i.e.&nbsp; <code>ecom</code> followed by the name of the columns.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/select_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>select(ecom, device, duration)</code></pre>
<pre><code>## # A tibble: 1,000 x 2
##    device duration
##    &lt;fct&gt;     &lt;dbl&gt;
##  1 laptop      693
##  2 tablet      459
##  3 laptop      996
##  4 tablet      468
##  5 mobile      955
##  6 laptop      135
##  7 mobile       75
##  8 mobile      908
##  9 mobile      209
## 10 mobile      208
## # ... with 990 more rows</code></pre>
<p>
We can select a set of columns using <code>:</code>. In the below example, we select all the columns starting from <code>referrer</code> up to <code>order_items</code>. Remember that we can use <code>:</code> only when the columns are adjacent to each other in the data set.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/select_2.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>select(ecom, referrer:order_items)</code></pre>
<pre><code>## # A tibble: 1,000 x 7
##    referrer device n_visit n_pages duration purchase order_items
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;
##  1 google   laptop      10       1      693 FALSE              0
##  2 yahoo    tablet       9       1      459 FALSE              0
##  3 direct   laptop       0       1      996 FALSE              0
##  4 bing     tablet       3      18      468 TRUE               6
##  5 yahoo    mobile       9       1      955 FALSE              0
##  6 yahoo    laptop       5       5      135 FALSE              0
##  7 yahoo    mobile      10       1       75 FALSE              0
##  8 direct   mobile      10       1      908 FALSE              0
##  9 bing     mobile       3      19      209 FALSE              0
## 10 google   mobile       6       1      208 FALSE              0
## # ... with 990 more rows</code></pre>
<p>
What if you want to select all columns except a few? Typing the name of many columns can be cumbersome and may also result in spelling errors. We may use <code>:</code> only if the columns are adjacent to each other but that may not always be the case. dplyr allows us to specify columns that need not be selected using <code>-</code>. In the below example, we select all columns except <code>n_pages</code> and <code>duration</code>. Notice the <code>-</code> before both of them.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/select_3.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>select(ecom, -n_pages, -duration)</code></pre>
<pre><code>## # A tibble: 1,000 x 6
##    referrer device n_visit purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 google   laptop      10 FALSE              0           0
##  2 yahoo    tablet       9 FALSE              0           0
##  3 direct   laptop       0 FALSE              0           0
##  4 bing     tablet       3 TRUE               6         434
##  5 yahoo    mobile       9 FALSE              0           0
##  6 yahoo    laptop       5 FALSE              0           0
##  7 yahoo    mobile      10 FALSE              0           0
##  8 direct   mobile      10 FALSE              0           0
##  9 bing     mobile       3 FALSE              0           0
## 10 google   mobile       6 FALSE              0           0
## # ... with 990 more rows</code></pre>
<section id="case-study-2" class="level5">
<h5 class="anchored" data-anchor-id="case-study-2">
Case Study
</h5>
<p>
For our case study, we need to select the column <code>order_value</code> to calculate the AOV. We also need to select the <code>device</code> column as we are computing the AOV for each device type.
</p>
<pre class="r"><code>select(ecom, device, order_value)</code></pre>
<pre><code>## # A tibble: 1,000 x 2
##    device order_value
##    &lt;fct&gt;        &lt;dbl&gt;
##  1 laptop           0
##  2 tablet           0
##  3 laptop           0
##  4 tablet         434
##  5 mobile           0
##  6 laptop           0
##  7 mobile           0
##  8 mobile           0
##  9 mobile           0
## 10 mobile           0
## # ... with 990 more rows</code></pre>
<p>
But we want the above data only for purchasers. Let us combine <code>filter()</code> and <code>select()</code> functions to extract <code>order_value</code> and <code>order_items</code> only for those visis that resulted in a purchase.
</p>
<pre class="r"><code># filter all visits that resulted in a purchase
ecom1 &lt;- filter(ecom, purchase)

# select the relevant columns
ecom2 &lt;- select(ecom1, device, order_value)

# view data
ecom2</code></pre>
<pre><code>## # A tibble: 103 x 2
##    device order_value
##    &lt;fct&gt;        &lt;dbl&gt;
##  1 tablet         434
##  2 mobile         651
##  3 tablet        1049
##  4 tablet        1304
##  5 tablet         622
##  6 tablet        1613
##  7 mobile         184
##  8 tablet         286
##  9 mobile         764
## 10 tablet        1667
## # ... with 93 more rows</code></pre>
</section>
</section>
<section id="grouping-data" class="level2">
<h2 class="anchored" data-anchor-id="grouping-data">
Grouping Data
</h2>
<p>
We need to compute the total order value and total order items for each device in order to compute their AOV. To achieve this, we need to group the selected <code>order_value</code> and <code>order_items</code> by device type. <code>group_by()</code> allows us to group or split data based on particular (discrete) variable. The first argument is the name of the data set and the second argument is the name of the column based on which the data will be split.
</p>
<p>
To split the data by referrer type, we use <code>group_by</code> and specify the data set i.e.&nbsp;<code>ecom</code> and the column based on which to split the data i.e.&nbsp;<code>referrer</code>.
</p>
<pre class="r"><code>group_by(ecom, referrer)</code></pre>
<pre><code>## # A tibble: 1,000 x 8
## # Groups:   referrer [5]
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 google   laptop      10       1      693 FALSE              0           0
##  2 yahoo    tablet       9       1      459 FALSE              0           0
##  3 direct   laptop       0       1      996 FALSE              0           0
##  4 bing     tablet       3      18      468 TRUE               6         434
##  5 yahoo    mobile       9       1      955 FALSE              0           0
##  6 yahoo    laptop       5       5      135 FALSE              0           0
##  7 yahoo    mobile      10       1       75 FALSE              0           0
##  8 direct   mobile      10       1      908 FALSE              0           0
##  9 bing     mobile       3      19      209 FALSE              0           0
## 10 google   mobile       6       1      208 FALSE              0           0
## # ... with 990 more rows</code></pre>
<section id="case-study-3" class="level5">
<h5 class="anchored" data-anchor-id="case-study-3">
Case Study
</h5>
<p>
In the second line in the previous output, you can observe <code>Groups: referrer [5]</code> . The data is split into 5 groups as the referrer variable has 5 distinct values. For our case study, we need to group the data by <code>device</code> type.
</p>
<pre class="r"><code># split ecom2 by device type
ecom3 &lt;- group_by(ecom2, device)
ecom3</code></pre>
<pre><code>## # A tibble: 103 x 2
## # Groups:   device [3]
##    device order_value
##    &lt;fct&gt;        &lt;dbl&gt;
##  1 tablet         434
##  2 mobile         651
##  3 tablet        1049
##  4 tablet        1304
##  5 tablet         622
##  6 tablet        1613
##  7 mobile         184
##  8 tablet         286
##  9 mobile         764
## 10 tablet        1667
## # ... with 93 more rows</code></pre>
</section>
</section>
<section id="summarise-data" class="level2">
<h2 class="anchored" data-anchor-id="summarise-data">
Summarise Data
</h2>
<p>
The next step is to compute the total order value and total order items for each device. i.e.&nbsp;we need to reduce the order value and order items data to a single summary. We can achieve this using <code>summarise()</code>. As usual, the first argument is the name of a data set and the subsequent arguments are functions that can summarise data. For example, we can use <code>min</code>, <code>max</code>, <code>sum</code>, <code>mean</code> etc.
</p>
<p>
Let us compute the average number of pages browsed by referrer type:
</p>
<ul>
<li>
split data by <code>referrer</code> type
</li>
<li>
compute the average number of pages using <code>mean</code>
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/groupby_summarise.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code># split data by referrer type
step_1 &lt;- group_by(ecom, referrer)

# compute average number of pages
step_2 &lt;- summarise(step_1, mean(n_pages))</code></pre>
<pre><code>## `summarise()` ungrouping (override with `.groups` argument)</code></pre>
<pre class="r"><code>step_2</code></pre>
<pre><code>## # A tibble: 5 x 2
##   referrer `mean(n_pages)`
## * &lt;fct&gt;              &lt;dbl&gt;
## 1 bing                6.13
## 2 direct              6.38
## 3 social              5.42
## 4 yahoo               5.99
## 5 google              5.73</code></pre>
<p>
Now let us compute both the <code>mean</code> and the <code>median</code>.
</p>
<pre class="r"><code># split data by referrer type
step_1 &lt;- group_by(ecom, referrer)

# compute average number of pages
step_2 &lt;- summarise(step_1, mean(n_pages), median(n_pages))</code></pre>
<pre><code>## `summarise()` ungrouping (override with `.groups` argument)</code></pre>
<pre class="r"><code>step_2</code></pre>
<pre><code>## # A tibble: 5 x 3
##   referrer `mean(n_pages)` `median(n_pages)`
## * &lt;fct&gt;              &lt;dbl&gt;             &lt;dbl&gt;
## 1 bing                6.13                 1
## 2 direct              6.38                 1
## 3 social              5.42                 1
## 4 yahoo               5.99                 2
## 5 google              5.73                 1</code></pre>
<p>
Another way to achieve the above result is to use the <code>summarise_all()</code> function. How does that work? It generates the specified summary for all the columns in the data set except for the column based on which the data has been grouped or split. So we need to ensure that the data does not have any irrelevant columns.
</p>
<ul>
<li>
split data by <code>referrer</code> type
</li>
<li>
select <code>order_value</code> and <code>order_items</code>
</li>
<li>
compute the average number of pages by applying the <code>mean</code> function to all the columns
</li>
</ul>
<pre class="r"><code># select relevant columns
step_1 &lt;- select(ecom, referrer, order_value)

# split data by referrer type
step_2 &lt;- group_by(step_1, referrer)

# compute average number of pages
step_3 &lt;- summarise_all(step_2, funs(mean))
step_3</code></pre>
<pre><code>## # A tibble: 5 x 2
##   referrer order_value
## * &lt;fct&gt;          &lt;dbl&gt;
## 1 bing            316.
## 2 direct          441.
## 3 social          380.
## 4 yahoo           470.
## 5 google          328.</code></pre>
<p>
Let us compute <code>mean</code> and <code>median</code> number of pages for each referre type using <code>summarise_all</code>.
</p>
<pre class="r"><code># select relevant columns
step_1 &lt;- select(ecom, referrer, order_value)

# split data by referrer type
step_2 &lt;- group_by(step_1, referrer)

# compute mean and median number of pages
step_3 &lt;- summarise_all(step_2, funs(mean, median))
step_3</code></pre>
<pre><code>## # A tibble: 5 x 3
##   referrer  mean median
## * &lt;fct&gt;    &lt;dbl&gt;  &lt;dbl&gt;
## 1 bing      316.      0
## 2 direct    441.      0
## 3 social    380.      0
## 4 yahoo     470.      0
## 5 google    328.      0</code></pre>
<section id="case-study-4" class="level5">
<h5 class="anchored" data-anchor-id="case-study-4">
Case Study
</h5>
<p>
So far, we have split the data based on the <code>device</code> type and we have selected 2 columns, <code>order_value</code> and <code>order_items</code>. We need the sum of order value and order items. What function can we use to obtain them? The <code>sum()</code> function will generate the sum of the values and hence we will use it inside the <code>summarise()</code> function. Remember, we need to provide a name to the summary being generated.
</p>
<pre class="r"><code>ecom4 &lt;- summarise(ecom3, revenue = sum(order_value),
          orders = n())</code></pre>
<pre><code>## `summarise()` ungrouping (override with `.groups` argument)</code></pre>
<pre class="r"><code>ecom4</code></pre>
<pre><code>## # A tibble: 3 x 3
##   device revenue orders
## * &lt;fct&gt;    &lt;dbl&gt;  &lt;int&gt;
## 1 laptop   56531     31
## 2 tablet   51321     36
## 3 mobile   51504     36</code></pre>
<p>
There you go, we have the total order value and total order items for each device type. If we use <code>summarise_all()</code>, it will generate the summary for the selected columns based on the function specified. To specify the functions, we need to use another argument <code>funs</code> and it can take any number of valid functions.
</p>
<pre class="r"><code>ecom4 &lt;- summarise_all(ecom3, funs(revenue = sum, orders = n()))
ecom4</code></pre>
<pre><code>## # A tibble: 3 x 3
##   device revenue orders
## * &lt;fct&gt;    &lt;dbl&gt;  &lt;int&gt;
## 1 laptop   56531     31
## 2 tablet   51321     36
## 3 mobile   51504     36</code></pre>
</section>
</section>
<section id="create-columns" class="level2">
<h2 class="anchored" data-anchor-id="create-columns">
Create Columns
</h2>
<p>
To create a new column, we will use <code>mutate()</code>. The first argument is the name of the data set and the subsequent arguments are expressions for creating new columns based out of existing columns.
</p>
<p>
Let us add a new column <code>avg_page_time</code> i.e.&nbsp;time on site divided by number of pages visited.
</p>
<pre class="r"><code># select duration and n_pages from ecom
mutate_1 &lt;- select(ecom, n_pages, duration)
mutate(mutate_1, avg_page_time = duration / n_pages)</code></pre>
<pre><code>## # A tibble: 1,000 x 3
##    n_pages duration avg_page_time
##      &lt;dbl&gt;    &lt;dbl&gt;         &lt;dbl&gt;
##  1       1      693           693
##  2       1      459           459
##  3       1      996           996
##  4      18      468            26
##  5       1      955           955
##  6       5      135            27
##  7       1       75            75
##  8       1      908           908
##  9      19      209            11
## 10       1      208           208
## # ... with 990 more rows</code></pre>
<p>
We can create new columns based on other columns created using <code>mutate</code>. Let us create another column <code>sqrt_avg_page_time</code> i.e.&nbsp;square root of the average time on page using <code>avg_page_time</code>.
</p>
<pre class="r"><code>mutate(mutate_1,
       avg_page_time = duration / n_pages,
       sqrt_avg_page_time = sqrt(avg_page_time))</code></pre>
<pre><code>## # A tibble: 1,000 x 4
##    n_pages duration avg_page_time sqrt_avg_page_time
##      &lt;dbl&gt;    &lt;dbl&gt;         &lt;dbl&gt;              &lt;dbl&gt;
##  1       1      693           693              26.3 
##  2       1      459           459              21.4 
##  3       1      996           996              31.6 
##  4      18      468            26               5.10
##  5       1      955           955              30.9 
##  6       5      135            27               5.20
##  7       1       75            75               8.66
##  8       1      908           908              30.1 
##  9      19      209            11               3.32
## 10       1      208           208              14.4 
## # ... with 990 more rows</code></pre>
<section id="case-study-5" class="level5">
<h5 class="anchored" data-anchor-id="case-study-5">
Case Study
</h5>
<p>
Back to our case study, from the last step we have the total order value and total order items for each device category and can compute the AOV. We will create a new column to store AOV.
</p>
<p>
<br>
</p>
<pre class="r"><code>ecom5 &lt;- mutate(ecom4, aov = revenue / orders)
ecom5</code></pre>
<pre><code>## # A tibble: 3 x 4
##   device revenue orders   aov
## * &lt;fct&gt;    &lt;dbl&gt;  &lt;int&gt; &lt;dbl&gt;
## 1 laptop   56531     31 1824.
## 2 tablet   51321     36 1426.
## 3 mobile   51504     36 1431.</code></pre>
</section>
</section>
<section id="select-columns-1" class="level2">
<h2 class="anchored" data-anchor-id="select-columns-1">
Select Columns
</h2>
<p>
The last step is to select the relevant columns. We will select the device type and the corresponding aov while getting rid of other columns. Use <code>select()</code> to extract the relevant columns.
</p>
<pre class="r"><code>ecom6 &lt;- select(ecom5, device, aov)
ecom6</code></pre>
<pre><code>## # A tibble: 3 x 2
##   device   aov
##   &lt;fct&gt;  &lt;dbl&gt;
## 1 laptop 1824.
## 2 tablet 1426.
## 3 mobile 1431.</code></pre>
</section>
<section id="arrange-data" class="level2">
<h2 class="anchored" data-anchor-id="arrange-data">
Arrange Data
</h2>
<p>
Arranging data in ascending or descending order is one of the most common tasks in data manipulation. We can use <code>arrange</code> to arrange data by different columns. Let us say we want to arrange data by the number of pages browsed.
</p>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/arrange_1.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>arrange(ecom, n_pages)</code></pre>
<pre><code>## # A tibble: 1,000 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 google   laptop      10       1      693 FALSE              0           0
##  2 yahoo    tablet       9       1      459 FALSE              0           0
##  3 direct   laptop       0       1      996 FALSE              0           0
##  4 yahoo    mobile       9       1      955 FALSE              0           0
##  5 yahoo    mobile      10       1       75 FALSE              0           0
##  6 direct   mobile      10       1      908 FALSE              0           0
##  7 google   mobile       6       1      208 FALSE              0           0
##  8 direct   laptop       9       1      738 FALSE              0           0
##  9 yahoo    mobile       7       1       19 FALSE              7        2423
## 10 bing     laptop       1       1      995 FALSE              0           0
## # ... with 990 more rows</code></pre>
<p>
If we want to arrange the data in descending order, we can use <code>desc()</code>. Let us arrange the data in descending order.
</p>
<pre class="r"><code>arrange(ecom , desc(n_pages))</code></pre>
<pre><code>## # A tibble: 1,000 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 social   tablet       9      20      420 TRUE               7        1613
##  2 bing     mobile       4      20      440 TRUE               3         184
##  3 yahoo    tablet       0      20      200 FALSE              0           0
##  4 direct   tablet       6      20      580 TRUE               5        1155
##  5 social   mobile       1      20      520 TRUE               5        1021
##  6 google   mobile       8      20      300 TRUE               7        2091
##  7 social   laptop       4      20      200 FALSE              0           0
##  8 yahoo    mobile       3      20      480 FALSE              0           0
##  9 social   laptop      10      20      280 TRUE               1        2011
## 10 yahoo    mobile       2      20      240 FALSE              0           0
## # ... with 990 more rows</code></pre>
<p>
Data can be arranged by multiple variables as well. Let us arrange data first by number of visits and then by number of pages in a descending order.
</p>
<pre class="r"><code>arrange(ecom, n_visit, desc(n_pages))</code></pre>
<pre><code>## # A tibble: 1,000 x 8
##    referrer device n_visit n_pages duration purchase order_items order_value
##    &lt;fct&gt;    &lt;fct&gt;    &lt;dbl&gt;   &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;          &lt;dbl&gt;       &lt;dbl&gt;
##  1 yahoo    tablet       0      20      200 FALSE              0           0
##  2 google   laptop       0      19      418 TRUE               2         996
##  3 bing     laptop       0      18      180 FALSE              0           0
##  4 yahoo    laptop       0      18      522 TRUE               8        1523
##  5 direct   tablet       0      18      252 FALSE              0           0
##  6 social   laptop       0      17      204 FALSE              0           0
##  7 bing     laptop       0      17      272 TRUE               9        1384
##  8 bing     mobile       0      16      272 FALSE              0           0
##  9 yahoo    mobile       0      15      255 FALSE              0           0
## 10 direct   laptop       0      15      255 FALSE              0           0
## # ... with 990 more rows</code></pre>
<section id="case-study-6" class="level5">
<h5 class="anchored" data-anchor-id="case-study-6">
Case Study
</h5>
<p>
If you observe <code>ecom6</code>, the <code>aov</code> column is arranged in descending order.
</p>
<pre class="r"><code>arrange(ecom6, aov)</code></pre>
<pre><code>## # A tibble: 3 x 2
##   device   aov
##   &lt;fct&gt;  &lt;dbl&gt;
## 1 tablet 1426.
## 2 mobile 1431.
## 3 laptop 1824.</code></pre>
</section>
</section>
<section id="aov-by-devices-1" class="level2">
<h2 class="anchored" data-anchor-id="aov-by-devices-1">
AOV by Devices
</h2>
<p>
Let us combine all the code from the above steps:
</p>
<pre class="r"><code>ecom1 &lt;- filter(ecom, purchase)
ecom2 &lt;- select(ecom1, device, order_value)
ecom3 &lt;- group_by(ecom2, device)
ecom4 &lt;- summarise_all(ecom3, funs(revenue = sum, orders = n()))
ecom5 &lt;- mutate(ecom4, aov = revenue / orders)
ecom6 &lt;- select(ecom5, device, aov)
ecom7 &lt;- arrange(ecom6, aov)
ecom7</code></pre>
<pre><code>## # A tibble: 3 x 2
##   device   aov
##   &lt;fct&gt;  &lt;dbl&gt;
## 1 tablet 1426.
## 2 mobile 1431.
## 3 laptop 1824.</code></pre>
<p>
If you observe, at each step we create a new variable(data frame) and then use it as an input in the next step i.e.&nbsp;the output from one step becomes the input for the next. Can we achieve the final outcome i.e.&nbsp;<code>ecom7</code> without creating the intermediate data (ecom1 - ecom6)? Yes, we can. We will use the <code>%&gt;%</code> operator to chain the steps and get rid of the intermediate data.
</p>
<pre class="r"><code>ecom %&gt;%
  filter(purchase) %&gt;%
  select(device, order_value) %&gt;%
  group_by(device) %&gt;%
  summarise_all(funs(revenue = sum, orders = n())) %&gt;%
  mutate(
    aov = revenue / orders
  ) %&gt;%
  select(device, aov) %&gt;%
  arrange(aov)</code></pre>
<pre><code>## # A tibble: 3 x 2
##   device   aov
##   &lt;fct&gt;  &lt;dbl&gt;
## 1 tablet 1426.
## 2 mobile 1431.
## 3 laptop 1824.</code></pre>
</section>
<section id="your-turn" class="level2">
<h2 class="anchored" data-anchor-id="your-turn">
Your Turn
</h2>
<ul>
<li>
what is the average number of pages visited by purchasers and non-purchasers?
</li>
<li>
what is the average time on site for purchasers vs non-purchasers?
</li>
<li>
what is the average number of pages visited by purchasers and non-purchasers using mobile?
</li>
</ul>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://dplyr.tidyverse.org/" class="uri">https://dplyr.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/transform.html" class="uri">http://r4ds.had.co.nz/transform.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>dplyr</category>
  <guid>https://blog.rsquaredacademy.com/posts/data-wrangling-with-dplyr-part-1/</guid>
  <pubDate>Thu, 23 Aug 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/dplyr.png" medium="image" type="image/png" height="89" width="144"/>
</item>
</channel>
</rss>
