<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Rsquared Academy Blog</title>
<link>https://blog.rsquaredacademy.com/r-programming/</link>
<atom:link href="https://blog.rsquaredacademy.com/r-programming/index.xml" rel="self" type="application/rss+xml"/>
<description>Core R: syntax, data structures, dates, strings, and getting help.</description>
<generator>quarto-1.10.18</generator>
<lastBuildDate>Tue, 05 Mar 2019 00:00:00 GMT</lastBuildDate>
<item>
  <title>Getting Help in R</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/getting-help-in-r-updated/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2019-03-05-getting-help-in-r.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<p>
<img src="https://blog.rsquaredacademy.com/img/help_banner.png" width="80%" style="display: block; margin: auto;">
</p>
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In this post, we will learn about the different methods of getting help in R. Often, we get stuck while doing some analysis as either we do not know the correct function to use or its syntax. It is important for anyone who is new to R to know the right place to look for help. There are two ways to look for help in R:
</p>
<ul>
<li>
built in help system
</li>
<li>
online
</li>
</ul>
<p>
In the first section, we will look at various online resources that can supplement the built in help system. In the second section, we will look at various ways to access the built in help system of R. Let us get started!
</p>
</section>
<section id="online-resources" class="level2">
<h2 class="anchored" data-anchor-id="online-resources">
Online Resources
</h2>
{{% youtube “ZgU6ICh2YzI” %}}
<section id="r-bloggers" class="level3">
<h3 class="anchored" data-anchor-id="r-bloggers">
R Bloggers
</h3>
<p>
<a href="https://www.r-bloggers.com/">R Bloggers</a> aggregates blogs written in English from across the globe. This is the first place you want to visit if you want help with R, data analysis, visualization and machine learning. There are blogs on a wide range of topics and the latest content is delivered to your inbox if you subscribe. If you are a blogger yourself, share it with th R community by adding your blog to R Bloggers.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/r-bloggers.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="stack-overflow" class="level3">
<h3 class="anchored" data-anchor-id="stack-overflow">
Stack Overflow
</h3>
<p>
<a href="https://stackoverflow.com/questions/tagged/r">Stack Overflow</a> is a great place to visit if you are having trouble with R code or packages. Chances are high that someone has already encountered the same or similar problem and you can use the answers given by R experts. In case you have encountered a new problem or issue, you can ask for help by providing a reproducible example of your analysis along with the R code. Use the <a href="http://reprex.tidyverse.org/">reprex</a> package to create reproducible examples.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/stack-overflow.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="twitter" class="level3">
<h3 class="anchored" data-anchor-id="twitter">
Twitter
</h3>
<p>
The R community is very active on Twitter and there are lot of R experts who are willing to help those who are new to R. Use the hashtag <a href="https://twitter.com/search?q=%23rstats">#rstats</a> if you are asking for help or guidance on Twitter.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/twitter.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="rstudio-community" class="level3">
<h3 class="anchored" data-anchor-id="rstudio-community">
RStudio Community
</h3>
<p>
<a href="https://community.rstudio.com/">RStudio Community</a> is similar to Stack Overflow. You can ask questions related to RStudio, Shiny, tidyverse and other RStudio products.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/rstudio.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="r4ds-community" class="level3">
<h3 class="anchored" data-anchor-id="r4ds-community">
r4ds Community
</h3>
<p>
An online learning community that brings learners and mentors on a single platform. You can learn more about the community <a href="http://www.rfordatasci.com/">here</a>.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/r4ds.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="rstudio-resources" class="level3">
<h3 class="anchored" data-anchor-id="rstudio-resources">
RStudio Resources
</h3>
<p>
<a href="https://resources.rstudio.com/">RStudio</a> has very good resources including cheatsheets, webinars and blogs.
</p>
</section>
<section id="reddit" class="level3">
<h3 class="anchored" data-anchor-id="reddit">
Reddit
</h3>
<p>
<a href="https://www.reddit.com/r/rstats/">Reddit</a> is another place where you can look for help. The discussions are moderated by R experts. There are subreddits for <a href="https://www.reddit.com/r/rstats/">Rstats</a>, <a href="https://www.reddit.com/r/Rlanguage/">Rlanguage</a>, <a href="https://www.reddit.com/r/rshiny/">Rstudio</a> and <a href="https://www.reddit.com/r/RStudio/">Rshiny</a>.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/reddit.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="r-weekly" class="level3">
<h3 class="anchored" data-anchor-id="r-weekly">
R Weekly
</h3>
<p>
Visit <a href="https://rweekly.org/">RWeekly</a> to get regular updates about the R community. You can find information about new packages, blogs, conferences, workshops, tutorials and R jobs.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/r-weekly.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="r-user-groups" class="level3">
<h3 class="anchored" data-anchor-id="r-user-groups">
R User Groups
</h3>
<p>
There are several R User Groups active across the globe. You can find the list <a href="https://jumpingrivers.github.io/meetingsR/r-user-groups.html">here</a>. Join the local user group to meet, discuss and learn from other R enthusiasts and experts.
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/user-group.png" width="80%" style="display: block; margin: auto;">
</p>
</section>
<section id="data-helpers" class="level3">
<h3 class="anchored" data-anchor-id="data-helpers">
Data Helpers
</h3>
<p>
<a href="http://www.datahelpers.org/">Data Helpers</a> is a list of data analysts, scientists and engineers willing to offer guidance put together by <a href="https://twitter.com/AngeBassa/">Angela Bassa</a>. Visit the website to learn more about how you can approach for help and guidance.
</p>
</section>
</section>
<section id="internal" class="level2">
<h2 class="anchored" data-anchor-id="internal">
Internal
</h2>
{{% youtube “gwpFCPDa8Tw” %}}
<p>
In this section, we will look at the following functions:
</p>
<ul>
<li>
<code>help.start()</code>
</li>
<li>
<code>help()</code>
</li>
<li>
<code>?</code>
</li>
<li>
<code>??</code>
</li>
<li>
<code>help.search()</code>
</li>
<li>
<code>demo()</code>
</li>
<li>
<code>example()</code>
</li>
<li>
<code>library()</code>
</li>
<li>
<code>vignette()</code>
</li>
<li>
<code>browseVignettes()</code>
</li>
</ul>
<section id="help.start" class="level3">
<h3 class="anchored" data-anchor-id="help.start">
help.start
</h3>
<p>
The <code>help.start()</code> function opens the documetation page in your browser. Here you can find manuals, reference and other materials.
</p>
<pre class="r"><code>help.start()</code></pre>
<pre><code>## starting httpd help server ... done</code></pre>
<pre><code>## If nothing happens, you should open
## 'http://127.0.0.1:19951/doc/html/index.html' yourself</code></pre>
</section>
<section id="help" class="level3">
<h3 class="anchored" data-anchor-id="help">
help
</h3>
<p>
Use <code>help()</code> to access the documentation of functions and data sets. <code>?</code> is a shortcut for <code>help()</code> and returns the same information.
</p>
<pre class="r"><code>help(plot)
?plot</code></pre>
</section>
<section id="help.search" class="level3">
<h3 class="anchored" data-anchor-id="help.search">
help.search
</h3>
<p>
<code>help.search()</code> will search all sources of documentation and return those that match the search string. <code>??</code> is a shortcut for <code>help.search()</code> and returns the same information.
</p>
<pre class="r"><code>help.search('regression')
??regression</code></pre>
<p>{{% youtube “Y9lcHOT7tJc” %}}</p>
</section>
<section id="demo" class="level3">
<h3 class="anchored" data-anchor-id="demo">
demo
</h3>
<p>
demo displays an interactive demonstration of certain topics provided in a R package. Typing <code>demo()</code> in the console will list the demos available in all the R packages installed.
</p>
<pre class="r"><code>demo()
demo(scoping)</code></pre>
<pre><code>## 
## 
##  demo(scoping)
##  ---- ~~~~~~~
## 
## &gt; ## Here is a little example which shows a fundamental difference between
## &gt; ## R and S.  It is a little example from Abelson and Sussman which models
## &gt; ## the way in which bank accounts work.    It shows how R functions can
## &gt; ## encapsulate state information.
## &gt; ##
## &gt; ## When invoked, "open.account" defines and returns three functions
## &gt; ## in a list.  Because the variable "total" exists in the environment
## &gt; ## where these functions are defined they have access to its value.
## &gt; ## This is even true when "open.account" has returned.  The only way
## &gt; ## to access the value of "total" is through the accessor functions
## &gt; ## withdraw, deposit and balance.  Separate accounts maintain their
## &gt; ## own balances.
## &gt; ##
## &gt; ## This is a very nifty way of creating "closures" and a little thought
## &gt; ## will show you that there are many ways of using this in statistics.
## &gt; 
## &gt; #  Copyright (C) 1997-8 The R Core Team
## &gt; 
## &gt; open.account &lt;- function(total) {
## + 
## +     list(
## +     deposit = function(amount) {
## +         if(amount &lt;= 0)
## +         stop("Deposits must be positive!\n")
## +         total &lt;&lt;- total + amount
## +         cat(amount,"deposited. Your balance is", total, "\n\n")
## +     },
## +     withdraw = function(amount) {
## +         if(amount &gt; total)
## +         stop("You don't have that much money!\n")
## +         total &lt;&lt;- total - amount
## +         cat(amount,"withdrawn.  Your balance is", total, "\n\n")
## +     },
## +     balance = function() {
## +         cat("Your balance is", total, "\n\n")
## +     }
## +     )
## + }
## 
## &gt; ross &lt;- open.account(100)
## 
## &gt; robert &lt;- open.account(200)
## 
## &gt; ross$withdraw(30)
## 30 withdrawn.  Your balance is 70 
## 
## 
## &gt; ross$balance()
## Your balance is 70 
## 
## 
## &gt; robert$balance()
## Your balance is 200 
## 
## 
## &gt; ross$deposit(50)
## 50 deposited. Your balance is 120 
## 
## 
## &gt; ross$balance()
## Your balance is 120 
## 
## 
## &gt; try(ross$withdraw(500)) # no way..
## Error in ross$withdraw(500) : You don't have that much money!</code></pre>
</section>
<section id="example" class="level3">
<h3 class="anchored" data-anchor-id="example">
example
</h3>
<p>
<code>example()</code> displays examples of the specified topic if available.
</p>
<pre class="r"><code>example('mean')</code></pre>
<pre><code>## 
## mean&gt; x &lt;- c(0:10, 50)
## 
## mean&gt; xm &lt;- mean(x)
## 
## mean&gt; c(xm, mean(x, trim = 0.10))
## [1] 8.75 5.50</code></pre>
</section>
</section>
<section id="package-documentation" class="level2">
<h2 class="anchored" data-anchor-id="package-documentation">
Package Documentation
</h2>
<section id="library" class="level3">
<h3 class="anchored" data-anchor-id="library">
library
</h3>
<p>
Access the documentation of a package using <code>help()</code> inside <code>library()</code>. The package need not be installed for accessing the documentation.
</p>
<pre class="r"><code>library(help = 'ggplot2')</code></pre>
</section>
<section id="vignette" class="level3">
<h3 class="anchored" data-anchor-id="vignette">
vignette
</h3>
<p>
A vignette is a long form guide to a R package. You can access the vignettes available using <code>vignette()</code>. It will display alist of vignettes available in installed packages.
</p>
<pre class="r"><code>vignette()</code></pre>
<p>
To access a specific vignette from a package, specify the name of the vignette and the package.
</p>
<pre class="r"><code>vignette('dplyr', package = 'dplyr')</code></pre>
<pre><code>## Warning: vignette 'dplyr' not found</code></pre>
</section>
<section id="browsevignettes" class="level3">
<h3 class="anchored" data-anchor-id="browsevignettes">
browseVignettes
</h3>
<p>
<code>browseVignettes()</code> is another way to access the vignettes in installed packages. It will list the vignettes in each package along with links to the web page and R code.
</p>
<pre class="r"><code>browseVignettes()</code></pre>
</section>
<section id="rsitesearch" class="level3">
<h3 class="anchored" data-anchor-id="rsitesearch">
RsiteSearch
</h3>
<p>
<code>RsiteSearch()</code> will search for a specified topics in help pages, vignettes and task views using the search engine at this <a href="http://search.r-project.org/">link</a> and return the result in a browser.
</p>
<pre class="r"><code>RSiteSearch('glm')</code></pre>
<pre><code>## A search query has been submitted to http://search.r-project.org
## The results page should open in your browser shortly</code></pre>
<p>{{% youtube “FsgJbDxsuT0” %}}</p>
</section>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">
Summary
</h2>
<p>
To sum it up, the R community is very beginner friendly and we hope you will find all the above resources, both internal help system and online resources useful.
</p>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/getting-help-in-r-updated/</guid>
  <pubDate>Tue, 05 Mar 2019 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/help_banner.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Readable Code with Pipes</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/readable-code-with-pipes/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-10-10-readable-code-with-pipes.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
R code contain a lot of parentheses in case of a sequence of multiple operations. When you are dealing with complex code, it results in nested function calls which are hard to read and maintain. The <a href="https://CRAN.R-project.org/package=magrittr">magrittr</a> package by <a href="http://stefanbache.dk/">Stefan Milton Bache</a> provides pipes enabling us to write R code that is readable.
</p>
<p>
Pipes allow us to clearly express a sequence of multiple operations by:
</p>
<ul>
<li>
structuring operations from left to right
</li>
<li>
avoiding
<ul>
<li>
nested function calls
</li>
<li>
intermediate steps
</li>
<li>
overwriting of original data
</li>
</ul>
</li>
<li>
minimizing creation of local variables
</li>
</ul>
</section>
<section id="pipes" class="level2">
<h2 class="anchored" data-anchor-id="pipes">
Pipes
</h2>
<p>
If you are using <a href="https://www.tidyverse.org/">tidyverse</a>, magrittr will be automatically loaded. We will look at 3 different types of pipes:
</p>
<ul>
<li>
<code>%&gt;%</code> : pipe a value forward into an expression or function call
</li>
<li>
<code>%&lt;&gt;%</code>: result assigned to left hand side object instead of returning it
</li>
<li>
<code>%$%</code> : expose names within left hand side objects to right hand side expressions
</li>
</ul>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages in this post:
</p>
<ul>
<li>
<a href="http://magrittr.tidyverse.org/">magrittr</a>
</li>
<li>
<a href="http://readr.tidyverse.org/">readr</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/">dplyr</a>
</li>
<li>
<a href="http://stringr.tidyverse.org/">stringr</a>
</li>
<li>
and <a href="http://readr.tidyverse.org/">purrr</a>
</li>
</ul>
<p>
You can find the data sets <a href="https://github.com/rsquaredacademy/datasets">here</a> and the codes <a href="https://gist.github.com/aravindhebbali/26d85ab4a4dadd2fe7c1f58d854cc950">here</a>.
</p>
<pre class="r"><code>library(magrittr)
library(readr)
library(dplyr)
library(stringr)
library(purrr)</code></pre>
</section>
<section id="data" class="level2">
<h2 class="anchored" data-anchor-id="data">
Data
</h2>
<pre class="r"><code>ecom &lt;- 
  read_csv('https://raw.githubusercontent.com/rsquaredacademy/datasets/master/web.csv',
    col_types = cols_only(
      referrer = col_factor(levels = c("bing", "direct", "social", "yahoo", "google")),
      n_pages = col_double(), duration = col_double(), purchase = col_logical()
    )
  )

ecom</code></pre>
<pre><code>## # A tibble: 1,000 x 4
##    referrer n_pages duration purchase
##    &lt;fct&gt;      &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 google         1      693 FALSE   
##  2 yahoo          1      459 FALSE   
##  3 direct         1      996 FALSE   
##  4 bing          18      468 TRUE    
##  5 yahoo          1      955 FALSE   
##  6 yahoo          5      135 FALSE   
##  7 yahoo          1       75 FALSE   
##  8 direct         1      908 FALSE   
##  9 bing          19      209 FALSE   
## 10 google         1      208 FALSE   
## # ... with 990 more rows</code></pre>
<p>
We will create a smaller data set from the above data to be used in some examples:
</p>
<pre class="r"><code>ecom_mini &lt;- sample_n(ecom, size = 10)
ecom_mini</code></pre>
<pre><code>## # A tibble: 10 x 4
##    referrer n_pages duration purchase
##    &lt;fct&gt;      &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 direct         1      136 FALSE   
##  2 direct         1      314 FALSE   
##  3 direct        18      324 TRUE    
##  4 social        10      290 TRUE    
##  5 yahoo          7      140 FALSE   
##  6 direct         1      658 FALSE   
##  7 bing          17      493 FALSE   
##  8 bing           1      406 FALSE   
##  9 social        15      405 FALSE   
## 10 google        15      210 FALSE</code></pre>
<section id="data-dictionary" class="level4">
<h4 class="anchored" data-anchor-id="data-dictionary">
Data Dictionary
</h4>
<ul>
<li>
referrer: referrer website/search engine
</li>
<li>
n_pages: number of pages visited
</li>
<li>
duration: time spent on the website (in seconds)
</li>
<li>
purchase: whether visitor purchased
</li>
</ul>
</section>
</section>
<section id="first-example" class="level2">
<h2 class="anchored" data-anchor-id="first-example">
First Example
</h2>
<p>
Let us start with a simple example. You must be aware of <code>head()</code>. If not, do not worry. It returns the first few observations/rows of data. We can specify the number of observations it should return as well. Let us use it to view the first 10 rows of our data set.
</p>
<pre class="r"><code>head(ecom, 10)</code></pre>
<pre><code>## # A tibble: 10 x 4
##    referrer n_pages duration purchase
##    &lt;fct&gt;      &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 google         1      693 FALSE   
##  2 yahoo          1      459 FALSE   
##  3 direct         1      996 FALSE   
##  4 bing          18      468 TRUE    
##  5 yahoo          1      955 FALSE   
##  6 yahoo          5      135 FALSE   
##  7 yahoo          1       75 FALSE   
##  8 direct         1      908 FALSE   
##  9 bing          19      209 FALSE   
## 10 google         1      208 FALSE</code></pre>
<section id="using-pipe" class="level4">
<h4 class="anchored" data-anchor-id="using-pipe">
Using Pipe
</h4>
<p>
Now let us do the same but with <code>%&gt;%</code>.
</p>
<pre class="r"><code>ecom %&gt;% head(10)</code></pre>
<pre><code>## # A tibble: 10 x 4
##    referrer n_pages duration purchase
##    &lt;fct&gt;      &lt;dbl&gt;    &lt;dbl&gt; &lt;lgl&gt;   
##  1 google         1      693 FALSE   
##  2 yahoo          1      459 FALSE   
##  3 direct         1      996 FALSE   
##  4 bing          18      468 TRUE    
##  5 yahoo          1      955 FALSE   
##  6 yahoo          5      135 FALSE   
##  7 yahoo          1       75 FALSE   
##  8 direct         1      908 FALSE   
##  9 bing          19      209 FALSE   
## 10 google         1      208 FALSE</code></pre>
</section>
</section>
<section id="square-root" class="level2">
<h2 class="anchored" data-anchor-id="square-root">
Square Root
</h2>
<p>
Time to try a slightly more challenging example. We want the square root of <code>n_pages</code> column from the data set.
</p>
<pre class="r"><code>y &lt;- sqrt(ecom_mini$n_pages)</code></pre>
<p>
Let us break down the above computation into small steps:
</p>
<ul>
<li>
select/expose the <code>n_pages</code> column from <code>ecom</code> data
</li>
<li>
compute the square root
</li>
<li>
assign the first few observations to <code>y</code>
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/pipes_square_root.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<p>
Let us reproduce <code>y</code> using pipes.
</p>
<pre class="r"><code># select n_pages variable and assign it to y
y &lt;-
    ecom_mini %$%
    n_pages

# compute square root of y and assign it to y 
y %&lt;&gt;% sqrt</code></pre>
<p>
Another way to compute the square root of y is shown below.
</p>
<pre class="r"><code>y &lt;-
  ecom_mini %$% 
  n_pages %&gt;% 
  sqrt()</code></pre>
</section>
<section id="visualization" class="level2">
<h2 class="anchored" data-anchor-id="visualization">
Visualization
</h2>
<p>
Let us look at a data visualization example. We will create a bar plot to visualize the frequency of different referrer types that drove purchasers to the website. Let us look at the steps involved in creating the bar plot:
</p>
<ul>
<li>
extract rows where purchase is TRUE
</li>
<li>
select/expose <code>referrer</code> column
</li>
<li>
tabulate referrer data using <code>table()</code>
</li>
<li>
use the tabulated data to create bar plot using <code>barplot()</code>
</li>
</ul>
<pre class="r"><code>barplot(table(subset(ecom, purchase)$referrer))</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/readable-code-with-pipes/2018-10-10-readable-code-with-pipes_files/figure-html/mag21-1.png" width="576" style="display: block; margin: auto;">
</p>
<section id="using-pipe-1" class="level4">
<h4 class="anchored" data-anchor-id="using-pipe-1">
Using pipe
</h4>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/pipes_data_visualization.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>ecom %&gt;%
  subset(purchase) %&gt;%
  extract('referrer') %&gt;%
  table() %&gt;%
  barplot()</code></pre>
<p>
<img src="https://blog.rsquaredacademy.com/posts/readable-code-with-pipes/2018-10-10-readable-code-with-pipes_files/figure-html/mag7-1.png" width="576" style="display: block; margin: auto;">
</p>
</section>
</section>
<section id="correlation" class="level2">
<h2 class="anchored" data-anchor-id="correlation">
Correlation
</h2>
<p>
Correlation is a statistical measure that indicates the extent to which two or more variables fluctuate together. In R, correlation is computed using <code>cor()</code>. Let us look at the correlation between the number of pages browsed and time spent on the site for visitors who purchased some product. Below are the steps for computing correlation:
</p>
<ul>
<li>
extract rows where purchase is TRUE
</li>
<li>
select/expose <code>n_pages</code> and <code>duration</code> columns
</li>
<li>
use <code>cor()</code> to compute the correlation
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/pipes_correlation.png" width="100%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code># without pipe
ecom1 &lt;- subset(ecom, purchase)
cor(ecom1$n_pages, ecom1$duration)</code></pre>
<pre><code>## [1] 0.4290905</code></pre>
<pre class="r"><code># with pipe
ecom %&gt;%
  subset(purchase) %$% 
  cor(n_pages, duration)</code></pre>
<pre><code>## [1] 0.4290905</code></pre>
<pre class="r"><code># with pipe
ecom %&gt;%
  filter(purchase) %$% 
  cor(n_pages, duration)</code></pre>
<pre><code>## [1] 0.4290905</code></pre>
</section>
<section id="regression" class="level2">
<h2 class="anchored" data-anchor-id="regression">
Regression
</h2>
<p>
Let us look at a regression example. We regress time spent on the site on number of pages visited. Below are the steps involved in running the regression:
</p>
<ul>
<li>
use <code>duration</code> and <code>n_pages</code> columns from ecom data
</li>
<li>
pass the above data to <code>lm()</code>
</li>
<li>
pass the output from <code>lm()</code> to <code>summary()</code>
</li>
</ul>
<pre class="r"><code>summary(lm(duration ~ n_pages, data = ecom))</code></pre>
<pre><code>## 
## Call:
## lm(formula = duration ~ n_pages, data = ecom)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -386.45 -213.03  -38.93  179.31  602.55 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(&gt;|t|)    
## (Intercept)  404.803     11.323  35.750  &lt; 2e-16 ***
## n_pages       -8.355      1.296  -6.449 1.76e-10 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 263.3 on 998 degrees of freedom
## Multiple R-squared:   0.04,  Adjusted R-squared:  0.03904 
## F-statistic: 41.58 on 1 and 998 DF,  p-value: 1.756e-10</code></pre>
<section id="using-pipe-2" class="level4">
<h4 class="anchored" data-anchor-id="using-pipe-2">
Using pipe
</h4>
<pre class="r"><code>ecom %$%
  lm(duration ~ n_pages) %&gt;%
  summary()</code></pre>
<pre><code>## 
## Call:
## lm(formula = duration ~ n_pages)
## 
## Residuals:
##     Min      1Q  Median      3Q     Max 
## -386.45 -213.03  -38.93  179.31  602.55 
## 
## Coefficients:
##             Estimate Std. Error t value Pr(&gt;|t|)    
## (Intercept)  404.803     11.323  35.750  &lt; 2e-16 ***
## n_pages       -8.355      1.296  -6.449 1.76e-10 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 263.3 on 998 degrees of freedom
## Multiple R-squared:   0.04,  Adjusted R-squared:  0.03904 
## F-statistic: 41.58 on 1 and 998 DF,  p-value: 1.756e-10</code></pre>
</section>
</section>
<section id="string-manipulation" class="level2">
<h2 class="anchored" data-anchor-id="string-manipulation">
String Manipulation
</h2>
<p>
We want to extract the first name (jovial) from the below email id and convert it to upper case. Below are the steps to achieve this:
</p>
<ul>
<li>
split the email id using the pattern <code>@</code> using <code>str_split()</code>
</li>
<li>
extract the first element from the resulting list using <code>extract2()</code>
</li>
<li>
extract the first element from the character vector using <code>extract()</code>
</li>
<li>
extract the first six characters using <code>str_sub()</code>
</li>
<li>
convert to upper case using <code>str_to_upper()</code>
</li>
</ul>
<p>
<br>
</p>
<p>
<img src="https://blog.rsquaredacademy.com/img/pipes_string.png" width="70%" style="display: block; margin: auto;">
</p>
<p>
<br>
</p>
<pre class="r"><code>email &lt;- 'jovialcann@anymail.com'

# without pipe
str_to_upper(str_sub(str_split(email, '@')[[1]][1], start = 1, end = 6))</code></pre>
<pre><code>## [1] "JOVIAL"</code></pre>
<pre class="r"><code># with pipe
email %&gt;%
  str_split(pattern = '@') %&gt;%
  extract2(1) %&gt;%
  extract(1) %&gt;%
  str_sub(start = 1, end = 6) %&gt;%
  str_to_upper()</code></pre>
<pre><code>## [1] "JOVIAL"</code></pre>
<p>
Another method that uses <code>map_chr()</code> from the <a href="https://readr.tidyverse.org/">purrr</a> package.
</p>
<pre class="r"><code>email %&gt;%
  str_split(pattern = '@') %&gt;%
  map_chr(1) %&gt;%
  str_sub(start = 1, end = 6) %&gt;%
  str_to_upper()</code></pre>
<pre><code>## [1] "JOVIAL"</code></pre>
</section>
<section id="data-extraction" class="level2">
<h2 class="anchored" data-anchor-id="data-extraction">
Data Extraction
</h2>
<p>
Let us turn our attention towards data extraction. magrittr provides alternatives to <code><img src="https://latex.codecogs.com/png.latex?%3C/code%3E,%20%3Ccode%3E%5B%3C/code%3E%20and%20%3Ccode%3E%5B%5B%3C/code%3E.%3C/p%3E%0A%3Cul%3E%0A%3Cli%3E%3Ccode%3Eextract()%3C/code%3E%3C/li%3E%0A%3Cli%3E%3Ccode%3Eextract2()%3C/code%3E%3C/li%3E%0A%3Cli%3E%3Ccode%3Euse_series()%3C/code%3E%3C/li%3E%0A%3C/ul%3E%0A%3Cdiv%20id=%22extract-column-by-name%22%20class=%22section%20level4%22%3E%0A%3Ch4%3EExtract%20Column%20By%20Name%3C/h4%3E%0A%3Cp%3ETo%20extract%20a%20specific%20column%20using%20the%20column%20name,%20we%20mention%20the%20name%0Aof%20the%20column%20in%20single/double%20quotes%20within%20%3Ccode%3E%5B%3C/code%3E%20or%20%3Ccode%3E%5B%5B%3C/code%3E.%20In%20case%20of%20%3Ccode%3E"></code>, we do not use quotes.
</p>
<pre class="r"><code># base 
ecom_mini['n_pages']</code></pre>
<pre><code>## # A tibble: 10 x 1
##    n_pages
##      &lt;dbl&gt;
##  1       1
##  2       1
##  3      18
##  4      10
##  5       7
##  6       1
##  7      17
##  8       1
##  9      15
## 10      15</code></pre>
<pre class="r"><code># magrittr
extract(ecom_mini, 'n_pages') </code></pre>
<pre><code>## # A tibble: 10 x 1
##    n_pages
##      &lt;dbl&gt;
##  1       1
##  2       1
##  3      18
##  4      10
##  5       7
##  6       1
##  7      17
##  8       1
##  9      15
## 10      15</code></pre>
</section>
<section id="extract-column-by-position" class="level4">
<h4 class="anchored" data-anchor-id="extract-column-by-position">
Extract Column By Position
</h4>
<p>
We can extract columns using their index position. Keep in mind that index position starts from <strong>1</strong> in R. In the below example, we show how to extract <code>n_pages</code> column but instead of using the column name, we use the column position.
</p>
<pre class="r"><code># base 
ecom_mini[2]</code></pre>
<pre><code>## # A tibble: 10 x 1
##    n_pages
##      &lt;dbl&gt;
##  1       1
##  2       1
##  3      18
##  4      10
##  5       7
##  6       1
##  7      17
##  8       1
##  9      15
## 10      15</code></pre>
<pre class="r"><code># magrittr
extract(ecom_mini, 2) </code></pre>
<pre><code>## # A tibble: 10 x 1
##    n_pages
##      &lt;dbl&gt;
##  1       1
##  2       1
##  3      18
##  4      10
##  5       7
##  6       1
##  7      17
##  8       1
##  9      15
## 10      15</code></pre>
</section>
<section id="extract-column-as-vector" class="level4">
<h4 class="anchored" data-anchor-id="extract-column-as-vector">
Extract Column (as vector)
</h4>
<p>
One important differentiator between <code>[</code> and <code>[[</code> is that <code>[[</code> will return a atomic vector and not a <code>data.frame</code>. <code><img src="https://latex.codecogs.com/png.latex?%3C/code%3E%20will%20also%20return%0Aa%20atomic%20vector.%20In%20magrittr,%20we%20can%20use%20%3Ccode%3Euse_series()%3C/code%3E%20in%20place%20of%0A%3Ccode%3E"></code>.
</p>
<pre class="r"><code># base 
ecom_mini$n_pages</code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
<pre class="r"><code># magrittr
use_series(ecom_mini, 'n_pages') </code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
</section>
<section id="extract-list-element-by-name" class="level4">
<h4 class="anchored" data-anchor-id="extract-list-element-by-name">
Extract List Element By Name
</h4>
<p>
Let us convert <code>ecom_mini</code> into a list using as.list() as shown below:
</p>
<pre class="r"><code>ecom_list &lt;- as.list(ecom_mini)</code></pre>
<p>
To extract elements of a list, we can use <code>extract2()</code>. It is an alternative for <code>[[</code>.
</p>
<pre class="r"><code># base 
ecom_list[['n_pages']]</code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
<pre class="r"><code># magrittr
extract2(ecom_list, 'n_pages')</code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
</section>
<section id="extract-list-element-by-position" class="level4">
<h4 class="anchored" data-anchor-id="extract-list-element-by-position">
Extract List Element By Position
</h4>
<pre class="r"><code># base 
ecom_list[[1]]</code></pre>
<pre><code>##  [1] direct direct direct social yahoo  direct bing   bing   social google
## Levels: bing direct social yahoo google</code></pre>
<pre class="r"><code># magrittr
extract2(ecom_list, 1)</code></pre>
<pre><code>##  [1] direct direct direct social yahoo  direct bing   bing   social google
## Levels: bing direct social yahoo google</code></pre>
</section>
<section id="extract-list-element" class="level4">
<h4 class="anchored" data-anchor-id="extract-list-element">
Extract List Element
</h4>
<p>
We can extract the elements of a list using <code>use_series()</code> as well.
</p>
<pre class="r"><code># base 
ecom_list$n_pages</code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
<pre class="r"><code># magrittr
use_series(ecom_list, n_pages)</code></pre>
<pre><code>##  [1]  1  1 18 10  7  1 17  1 15 15</code></pre>
</section>
 ]]></description>
  <category>pipes</category>
  <guid>https://blog.rsquaredacademy.com/posts/readable-code-with-pipes/</guid>
  <pubDate>Wed, 10 Oct 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/pipes.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Introduction to tibbles</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/introduction-to-tibbles/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2018-09-28-introduction-to-tibbles.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<blockquote class="blockquote">
<p>
A <strong>tibble</strong>, or <code>tbl_df</code>, is a modern reimagining of the data.frame, keeping what time has proven to be effective, and throwing out what is not. Tibbles are data.frames that are lazy and surly: they do less (i.e.&nbsp;they don’t change variable names or types, and don’t do partial matching) and complain more (e.g.&nbsp;when a variable does not exist). This forces you to confront problems earlier, typically leading to cleaner, more expressive code. Tibbles also have an enhanced <code>print method()</code> which makes them easier to use with large datasets containing complex objects.
</p>
</blockquote>
<blockquote class="blockquote">
<p>
Source: <a href="https://tibble.tidyverse.org/" class="uri">https://tibble.tidyverse.org/</a>
</p>
</blockquote>
<p>
In this post, we will explore tibbles. To be more precise, we will learn:
</p>
<ul>
<li>
how tibbles are different from data frames?
</li>
<li>
how to create tibbles?
</li>
<li>
how to manipulate tibbles?
</li>
</ul>
</section>
<section id="libraries-code-data" class="level2">
<h2 class="anchored" data-anchor-id="libraries-code-data">
Libraries, Code &amp; Data
</h2>
<p>
We will use the following packages:
</p>
<ul>
<li>
<a href="http://tibble.tidyverse.org/">tibble</a>
</li>
<li>
<a href="http://dplyr.tidyverse.org/">dplyr</a>
</li>
</ul>
<p>
The code can be found <a href="https://gist.github.com/aravindhebbali/9a3814b9b4bb5c271d030b15ce4ecdf1">here</a>.
</p>
<pre class="r"><code>library(tibble)
library(dplyr)</code></pre>
</section>
<section id="creating-tibbles" class="level2">
<h2 class="anchored" data-anchor-id="creating-tibbles">
Creating tibbles
</h2>
<p>
tibble can be created using any of the following:
</p>
<ul>
<li>
<code>tibble()</code>
</li>
<li>
<code>as_tibble()</code>
</li>
<li>
<code>tribble()</code>
</li>
</ul>
<p>
Let us start with <code>tibble()</code>.
</p>
<pre class="r"><code>tibble(x = letters,
       y = 1:26,
       z = sample(100, 26))</code></pre>
<pre><code>## # A tibble: 26 x 3
##    x         y     z
##    &lt;chr&gt; &lt;int&gt; &lt;int&gt;
##  1 a         1    71
##  2 b         2    56
##  3 c         3    59
##  4 d         4    14
##  5 e         5    58
##  6 f         6    60
##  7 g         7    16
##  8 h         8     9
##  9 i         9    66
## 10 j        10    48
## # ... with 16 more rows</code></pre>
<p>
We mentioned the column names followed by the data. If you do not specify the column names, <code>tibble()</code> will supply them. Ensure that the length of each column is same.
</p>
</section>
<section id="tibble-features" class="level2">
<h2 class="anchored" data-anchor-id="tibble-features">
tibble features
</h2>
<ul>
<li>
never changes input’s types
</li>
</ul>
<p>
<code>tibble()</code> will never alter the input’s type. For example, if you supply a character vector it will not be converted to factor unlike data.frame where you need to set <code>stringsAsFactors</code> to <code>FALSE</code>.
</p>
<pre class="r"><code>tibble(x = letters,
       y = 1:26,
       z = sample(100, 26))</code></pre>
<pre><code>## # A tibble: 26 x 3
##    x         y     z
##    &lt;chr&gt; &lt;int&gt; &lt;int&gt;
##  1 a         1    62
##  2 b         2    13
##  3 c         3    75
##  4 d         4    17
##  5 e         5    82
##  6 f         6    83
##  7 g         7     9
##  8 h         8    76
##  9 i         9    19
## 10 j        10    97
## # ... with 16 more rows</code></pre>
<ul>
<li>
never adjusts variable names
</li>
</ul>
<p>
<code>tibble()</code> will never modify the column names. In the below example, you can observe that while <code>data.frame</code> adds a <code>.</code>, <code>tibble()</code> retains the column names as is.
</p>
<pre class="r"><code>names(data.frame(`order value` = 10))</code></pre>
<pre><code>## [1] "order.value"</code></pre>
<pre class="r"><code>names(tibble(`order value` = 10))</code></pre>
<pre><code>## [1] "order value"</code></pre>
<ul>
<li>
never prints all rows
</li>
</ul>
<p>
<code>tibble()</code> will never print all the rows and clutter your console. It will only print the first 10 rows and only as many columns that fit the width of the console.
</p>
<pre class="r"><code>x &lt;- 1:100
y &lt;- letters[1]
z &lt;- sample(c(TRUE, FALSE), 100, replace = TRUE)
tibble(x, y, z)</code></pre>
<pre><code>## # A tibble: 100 x 3
##        x y     z    
##    &lt;int&gt; &lt;chr&gt; &lt;lgl&gt;
##  1     1 a     TRUE 
##  2     2 a     FALSE
##  3     3 a     TRUE 
##  4     4 a     TRUE 
##  5     5 a     FALSE
##  6     6 a     TRUE 
##  7     7 a     TRUE 
##  8     8 a     FALSE
##  9     9 a     TRUE 
## 10    10 a     FALSE
## # ... with 90 more rows</code></pre>
<ul>
<li>
never recycles vector of length greater than 1
</li>
</ul>
<p>
Recycling vectors of length greater than 1 often leads to errors and as such <code>tibble()</code> will only recycle vectors of length 1.
</p>
<pre class="r"><code>x &lt;- 1:100
y &lt;- letters
z &lt;- sample(c(TRUE, FALSE), 100, replace = TRUE)
tibble(x, y, z)
Error in overscope_eval_next(overscope, expr) : object 'y' not found</code></pre>
</section>
<section id="membership-testing" class="level2">
<h2 class="anchored" data-anchor-id="membership-testing">
Membership Testing
</h2>
<p>
We can test if an object is a tibble using <code>is_tibble()</code>.
</p>
<pre class="r"><code>is_tibble(mtcars)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>is_tibble(as_tibble(mtcars))</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="tribble" class="level2">
<h2 class="anchored" data-anchor-id="tribble">
Tribble
</h2>
<p>
Another way to create tibbles is using <code>tribble()</code>:
</p>
<ul>
<li>
it is short for transposed tibbles
</li>
<li>
it is customized for data entry in code
</li>
<li>
column names start with <code>~</code>
</li>
<li>
and values are separated by commas
</li>
</ul>
<pre class="r"><code>tribble(
  ~x, ~y, ~z,
  #--|--|----
  1, TRUE, 'a',
  2, FALSE, 'b'
)</code></pre>
<pre><code>## # A tibble: 2 x 3
##       x y     z    
##   &lt;dbl&gt; &lt;lgl&gt; &lt;chr&gt;
## 1     1 TRUE  a    
## 2     2 FALSE b</code></pre>
</section>
<section id="column-names" class="level2">
<h2 class="anchored" data-anchor-id="column-names">
Column Names
</h2>
<p>
Names of the columns in tibbles need not be valid R variable names. They can contain unusual characters like a space or a smiley but must be enclosed in ticks.
</p>
<pre class="r"><code>tibble(
  ` ` = 'space',
  `2` = 'integer',
  `:)` = 'smiley'
)</code></pre>
<pre><code>## # A tibble: 1 x 3
##   ` `   `2`     `:)`  
##   &lt;chr&gt; &lt;chr&gt;   &lt;chr&gt; 
## 1 space integer smiley</code></pre>
</section>
<section id="add-rows" class="level2">
<h2 class="anchored" data-anchor-id="add-rows">
Add Rows
</h2>
<p>
Let us add data related to <strong>Safari</strong> browser to the web traffic data using <code>add_row()</code>.
</p>
<pre class="r"><code>browsers &lt;- enframe(c(chrome = 40, firefox = 20, edge = 30))
browsers</code></pre>
<pre><code>## # A tibble: 3 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 firefox    20
## 3 edge       30</code></pre>
<pre class="r"><code>add_row(browsers, name = 'safari', value = 10)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 firefox    20
## 3 edge       30
## 4 safari     10</code></pre>
<p>
If we want to add the data at a particular row, we can specify the row number using the <code>.before</code> argument. Let us add the data related to <strong>Safari</strong> browser in the second row instead of the last row.
</p>
<pre class="r"><code>add_row(browsers, name = 'safari', value = 10, .before = 2)</code></pre>
<pre><code>## # A tibble: 4 x 2
##   name    value
##   &lt;chr&gt;   &lt;dbl&gt;
## 1 chrome     40
## 2 safari     10
## 3 firefox    20
## 4 edge       30</code></pre>
</section>
<section id="add-columns" class="level2">
<h2 class="anchored" data-anchor-id="add-columns">
Add Columns
</h2>
<p>
<code>add_column()</code> adds a new column to tibbles.
</p>
<pre class="r"><code>browsers &lt;- enframe(c(chrome = 40, firefox = 20, edge = 30, safari = 10))
add_column(browsers, visits = c(4000, 2000, 3000, 1000))</code></pre>
<pre><code>## # A tibble: 4 x 3
##   name    value visits
##   &lt;chr&gt;   &lt;dbl&gt;  &lt;dbl&gt;
## 1 chrome     40   4000
## 2 firefox    20   2000
## 3 edge       30   3000
## 4 safari     10   1000</code></pre>
</section>
<section id="rownames" class="level2">
<h2 class="anchored" data-anchor-id="rownames">
Rownames
</h2>
<p>
The <a href="tibble.tidyverse.org">tibble</a> package provides a set of functions to deal with rownames. Remember, <code>tibble</code> does not have <code>rownames</code> unlike <code>data.frame</code>. To check whether a data set has rownames, use <code>has_rownames()</code>.
</p>
<pre class="r"><code>has_rownames(mtcars)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<section id="remove-rownames" class="level4">
<h4 class="anchored" data-anchor-id="remove-rownames">
Remove Rownames
</h4>
<pre class="r"><code>remove_rownames(mtcars)</code></pre>
<pre><code>##     mpg cyl  disp  hp drat    wt  qsec vs am gear carb
## 1  21.0   6 160.0 110 3.90 2.620 16.46  0  1    4    4
## 2  21.0   6 160.0 110 3.90 2.875 17.02  0  1    4    4
## 3  22.8   4 108.0  93 3.85 2.320 18.61  1  1    4    1
## 4  21.4   6 258.0 110 3.08 3.215 19.44  1  0    3    1
## 5  18.7   8 360.0 175 3.15 3.440 17.02  0  0    3    2
## 6  18.1   6 225.0 105 2.76 3.460 20.22  1  0    3    1
## 7  14.3   8 360.0 245 3.21 3.570 15.84  0  0    3    4
## 8  24.4   4 146.7  62 3.69 3.190 20.00  1  0    4    2
## 9  22.8   4 140.8  95 3.92 3.150 22.90  1  0    4    2
## 10 19.2   6 167.6 123 3.92 3.440 18.30  1  0    4    4
## 11 17.8   6 167.6 123 3.92 3.440 18.90  1  0    4    4
## 12 16.4   8 275.8 180 3.07 4.070 17.40  0  0    3    3
## 13 17.3   8 275.8 180 3.07 3.730 17.60  0  0    3    3
## 14 15.2   8 275.8 180 3.07 3.780 18.00  0  0    3    3
## 15 10.4   8 472.0 205 2.93 5.250 17.98  0  0    3    4
## 16 10.4   8 460.0 215 3.00 5.424 17.82  0  0    3    4
## 17 14.7   8 440.0 230 3.23 5.345 17.42  0  0    3    4
## 18 32.4   4  78.7  66 4.08 2.200 19.47  1  1    4    1
## 19 30.4   4  75.7  52 4.93 1.615 18.52  1  1    4    2
## 20 33.9   4  71.1  65 4.22 1.835 19.90  1  1    4    1
## 21 21.5   4 120.1  97 3.70 2.465 20.01  1  0    3    1
## 22 15.5   8 318.0 150 2.76 3.520 16.87  0  0    3    2
## 23 15.2   8 304.0 150 3.15 3.435 17.30  0  0    3    2
## 24 13.3   8 350.0 245 3.73 3.840 15.41  0  0    3    4
## 25 19.2   8 400.0 175 3.08 3.845 17.05  0  0    3    2
## 26 27.3   4  79.0  66 4.08 1.935 18.90  1  1    4    1
## 27 26.0   4 120.3  91 4.43 2.140 16.70  0  1    5    2
## 28 30.4   4  95.1 113 3.77 1.513 16.90  1  1    5    2
## 29 15.8   8 351.0 264 4.22 3.170 14.50  0  1    5    4
## 30 19.7   6 145.0 175 3.62 2.770 15.50  0  1    5    6
## 31 15.0   8 301.0 335 3.54 3.570 14.60  0  1    5    8
## 32 21.4   4 121.0 109 4.11 2.780 18.60  1  1    4    2</code></pre>
</section>
<section id="rownames-to-column" class="level4">
<h4 class="anchored" data-anchor-id="rownames-to-column">
Rownames to Column
</h4>
<pre class="r"><code>head(rownames_to_column(mtcars))</code></pre>
<pre><code>##             rowname  mpg cyl disp  hp drat    wt  qsec vs am gear carb
## 1         Mazda RX4 21.0   6  160 110 3.90 2.620 16.46  0  1    4    4
## 2     Mazda RX4 Wag 21.0   6  160 110 3.90 2.875 17.02  0  1    4    4
## 3        Datsun 710 22.8   4  108  93 3.85 2.320 18.61  1  1    4    1
## 4    Hornet 4 Drive 21.4   6  258 110 3.08 3.215 19.44  1  0    3    1
## 5 Hornet Sportabout 18.7   8  360 175 3.15 3.440 17.02  0  0    3    2
## 6           Valiant 18.1   6  225 105 2.76 3.460 20.22  1  0    3    1</code></pre>
</section>
<section id="column-to-rownames" class="level4">
<h4 class="anchored" data-anchor-id="column-to-rownames">
Column to Rownames
</h4>
<p>
To convert the first column in the data set to rownames, use <code>column_to_rownames()</code>:
</p>
<pre class="r"><code>mtcars_tbl &lt;- rownames_to_column(mtcars)
column_to_rownames(mtcars_tbl)</code></pre>
<pre><code>##                      mpg cyl  disp  hp drat    wt  qsec vs am gear carb
## Mazda RX4           21.0   6 160.0 110 3.90 2.620 16.46  0  1    4    4
## Mazda RX4 Wag       21.0   6 160.0 110 3.90 2.875 17.02  0  1    4    4
## Datsun 710          22.8   4 108.0  93 3.85 2.320 18.61  1  1    4    1
## Hornet 4 Drive      21.4   6 258.0 110 3.08 3.215 19.44  1  0    3    1
## Hornet Sportabout   18.7   8 360.0 175 3.15 3.440 17.02  0  0    3    2
## Valiant             18.1   6 225.0 105 2.76 3.460 20.22  1  0    3    1
## Duster 360          14.3   8 360.0 245 3.21 3.570 15.84  0  0    3    4
## Merc 240D           24.4   4 146.7  62 3.69 3.190 20.00  1  0    4    2
## Merc 230            22.8   4 140.8  95 3.92 3.150 22.90  1  0    4    2
## Merc 280            19.2   6 167.6 123 3.92 3.440 18.30  1  0    4    4
## Merc 280C           17.8   6 167.6 123 3.92 3.440 18.90  1  0    4    4
## Merc 450SE          16.4   8 275.8 180 3.07 4.070 17.40  0  0    3    3
## Merc 450SL          17.3   8 275.8 180 3.07 3.730 17.60  0  0    3    3
## Merc 450SLC         15.2   8 275.8 180 3.07 3.780 18.00  0  0    3    3
## Cadillac Fleetwood  10.4   8 472.0 205 2.93 5.250 17.98  0  0    3    4
## Lincoln Continental 10.4   8 460.0 215 3.00 5.424 17.82  0  0    3    4
## Chrysler Imperial   14.7   8 440.0 230 3.23 5.345 17.42  0  0    3    4
## Fiat 128            32.4   4  78.7  66 4.08 2.200 19.47  1  1    4    1
## Honda Civic         30.4   4  75.7  52 4.93 1.615 18.52  1  1    4    2
## Toyota Corolla      33.9   4  71.1  65 4.22 1.835 19.90  1  1    4    1
## Toyota Corona       21.5   4 120.1  97 3.70 2.465 20.01  1  0    3    1
## Dodge Challenger    15.5   8 318.0 150 2.76 3.520 16.87  0  0    3    2
## AMC Javelin         15.2   8 304.0 150 3.15 3.435 17.30  0  0    3    2
## Camaro Z28          13.3   8 350.0 245 3.73 3.840 15.41  0  0    3    4
## Pontiac Firebird    19.2   8 400.0 175 3.08 3.845 17.05  0  0    3    2
## Fiat X1-9           27.3   4  79.0  66 4.08 1.935 18.90  1  1    4    1
## Porsche 914-2       26.0   4 120.3  91 4.43 2.140 16.70  0  1    5    2
## Lotus Europa        30.4   4  95.1 113 3.77 1.513 16.90  1  1    5    2
## Ford Pantera L      15.8   8 351.0 264 4.22 3.170 14.50  0  1    5    4
## Ferrari Dino        19.7   6 145.0 175 3.62 2.770 15.50  0  1    5    6
## Maserati Bora       15.0   8 301.0 335 3.54 3.570 14.60  0  1    5    8
## Volvo 142E          21.4   4 121.0 109 4.11 2.780 18.60  1  1    4    2</code></pre>
</section>
</section>
<section id="glimpse" class="level2">
<h2 class="anchored" data-anchor-id="glimpse">
Glimpse
</h2>
<p>
Use <code>glimpse()</code> to get an overview of the data.
</p>
<pre class="r"><code>glimpse(mtcars)</code></pre>
<pre><code>## Rows: 32
## Columns: 11
## $ mpg  &lt;dbl&gt; 21.0, 21.0, 22.8, 21.4, 18.7, 18.1, 14.3, 24.4, 22.8, 19.2, 17...
## $ cyl  &lt;dbl&gt; 6, 6, 4, 6, 8, 6, 8, 4, 4, 6, 6, 8, 8, 8, 8, 8, 8, 4, 4, 4, 4,...
## $ disp &lt;dbl&gt; 160.0, 160.0, 108.0, 258.0, 360.0, 225.0, 360.0, 146.7, 140.8,...
## $ hp   &lt;dbl&gt; 110, 110, 93, 110, 175, 105, 245, 62, 95, 123, 123, 180, 180, ...
## $ drat &lt;dbl&gt; 3.90, 3.90, 3.85, 3.08, 3.15, 2.76, 3.21, 3.69, 3.92, 3.92, 3....
## $ wt   &lt;dbl&gt; 2.620, 2.875, 2.320, 3.215, 3.440, 3.460, 3.570, 3.190, 3.150,...
## $ qsec &lt;dbl&gt; 16.46, 17.02, 18.61, 19.44, 17.02, 20.22, 15.84, 20.00, 22.90,...
## $ vs   &lt;dbl&gt; 0, 0, 1, 1, 0, 1, 0, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1,...
## $ am   &lt;dbl&gt; 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0,...
## $ gear &lt;dbl&gt; 4, 4, 4, 3, 3, 3, 3, 4, 4, 4, 4, 3, 3, 3, 3, 3, 3, 4, 4, 4, 3,...
## $ carb &lt;dbl&gt; 4, 4, 1, 1, 2, 1, 4, 2, 2, 4, 4, 3, 3, 3, 4, 4, 4, 1, 2, 1, 1,...</code></pre>
</section>
<section id="check-column" class="level2">
<h2 class="anchored" data-anchor-id="check-column">
Check Column
</h2>
<p>
<code>has_name()</code> can be used to check if a tibble has a specific column.
</p>
<pre class="r"><code>has_name(mtcars, 'cyl')</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>has_name(mtcars, 'gears')</code></pre>
<pre><code>## [1] FALSE</code></pre>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">
Summary
</h2>
<section id="creating-tibbles-1" class="level4">
<h4 class="anchored" data-anchor-id="creating-tibbles-1">
Creating tibbles
</h4>
<ul>
<li>
use <code>tibble()</code> to create tibbles
</li>
<li>
use <code>as_tibble()</code> to coerce other objects to tibble
</li>
<li>
use <code>enframe()</code> to coerce vector to tibble
</li>
<li>
use <code>tribble()</code> to create tibble using data entry
</li>
</ul>
</section>
<section id="modifying-tibbles" class="level4">
<h4 class="anchored" data-anchor-id="modifying-tibbles">
Modifying tibbles
</h4>
<ul>
<li>
use <code>add_row()</code> to add a new row
</li>
<li>
use <code>add_column()</code> to add a new column
</li>
<li>
use <code>remove_rownames()</code> to remove rownames from data
</li>
<li>
use <code>rownames_to_colum()</code> to coerce rowname to first column
</li>
<li>
use <code>column_to_rownames()</code> to coerce first column to rownames
</li>
</ul>
</section>
<section id="testing-tibbles" class="level4">
<h4 class="anchored" data-anchor-id="testing-tibbles">
Testing tibbles
</h4>
<ul>
<li>
use <code>is_tibble()</code> to test if an object is a tibble
</li>
<li>
use <code>has_rownames()</code> to check whether a data set has rownames
</li>
<li>
use <code>has_name()</code> to check if tibble has a specific column
</li>
<li>
use <code>glimpse()</code> to get an overview of data
</li>
</ul>
</section>
</section>
<section id="references" class="level2">
<h2 class="anchored" data-anchor-id="references">
References
</h2>
<ul>
<li>
<a href="https://tibble.tidyverse.org/" class="uri">https://tibble.tidyverse.org/</a>
</li>
<li>
<a href="http://r4ds.had.co.nz/tibbles.html" class="uri">http://r4ds.had.co.nz/tibbles.html</a>
</li>
</ul>
</section>



 ]]></description>
  <category>data wrangling</category>
  <category>tibbles</category>
  <guid>https://blog.rsquaredacademy.com/posts/introduction-to-tibbles/</guid>
  <pubDate>Fri, 28 Sep 2018 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/tibble_intro.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Dataframes</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/dataframes/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-05-12-dataframes.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt about lists. In this post, we will learn about <code>dataframe</code>.
</p>
<ul>
<li>
create dataframe
</li>
<li>
select columns
</li>
<li>
select rows
</li>
<li>
utitlity functions
</li>
</ul>
</section>
<section id="create-dataframes" class="level2">
<h2 class="anchored" data-anchor-id="create-dataframes">
Create dataframes
</h2>
<p>
Use <code>data.frame</code> to create dataframes. Below is the function syntax:
</p>
<pre class="r"><code>args(data.frame)</code></pre>
<pre><code>## function (..., row.names = NULL, check.rows = FALSE, check.names = TRUE, 
##     fix.empty.names = TRUE, stringsAsFactors = default.stringsAsFactors()) 
## NULL</code></pre>
<p>
Data frames are basically lists with elements of equal lenght and as such, they are heterogeneous. Let us create a dataframe:
</p>
<pre class="r"><code>name &lt;- c('John', 'Jack', 'Jill')
age &lt;- c(29, 25, 27)
graduate &lt;- c(TRUE, TRUE, FALSE)
students &lt;- data.frame(name, age, graduate)
students</code></pre>
<pre><code>##   name age graduate
## 1 John  29     TRUE
## 2 Jack  25     TRUE
## 3 Jill  27    FALSE</code></pre>
</section>
<section id="basic-information" class="level2">
<h2 class="anchored" data-anchor-id="basic-information">
Basic Information
</h2>
<pre class="r"><code>class(students)
## [1] "data.frame"
names(students)
## [1] "name"     "age"      "graduate"
colnames(students)
## [1] "name"     "age"      "graduate"
str(students)
## 'data.frame':    3 obs. of  3 variables:
##  $ name    : chr  "John" "Jack" "Jill"
##  $ age     : num  29 25 27
##  $ graduate: logi  TRUE TRUE FALSE
dim(students)
## [1] 3 3
nrow(students)
## [1] 3
ncol(students)
## [1] 3</code></pre>
</section>
<section id="select-columns" class="level2">
<h2 class="anchored" data-anchor-id="select-columns">
Select Columns
</h2>
<ul>
<li>
<code>[]</code>
</li>
<li>
<code>[[]]</code>
</li>
<li>
<code>$</code>
</li>
</ul>
<pre class="r"><code># using [
students[1]
##   name
## 1 John
## 2 Jack
## 3 Jill

# using [[
students[[1]]
## [1] "John" "Jack" "Jill"

# using $
students$name
## [1] "John" "Jack" "Jill"</code></pre>
<section id="multiple-columns" class="level3">
<h3 class="anchored" data-anchor-id="multiple-columns">
Multiple Columns
</h3>
<pre class="r"><code>students[, 1:3]
##   name age graduate
## 1 John  29     TRUE
## 2 Jack  25     TRUE
## 3 Jill  27    FALSE

students[, c(1, 3)]
##   name graduate
## 1 John     TRUE
## 2 Jack     TRUE
## 3 Jill    FALSE</code></pre>
</section>
</section>
<section id="select-rows" class="level2">
<h2 class="anchored" data-anchor-id="select-rows">
Select Rows
</h2>
<pre class="r"><code># single row
students[1, ]
##   name age graduate
## 1 John  29     TRUE

# multiple row
students[c(1, 3), ]
##   name age graduate
## 1 John  29     TRUE
## 3 Jill  27    FALSE</code></pre>
<p>
If you have observed carefully, the column <code>names</code> has been coerced to type factor. This happens because of a default argument in <code>data.frame</code> which is <code>stringsAsFactors</code> which is set to <code>TRUE</code>. If you do not want to treat it as <code>factors</code>, set the argument to <code>FALSE</code>.
</p>
<pre class="r"><code>students &lt;- data.frame(name, age, graduate, stringsAsFactors = FALSE)</code></pre>
<p>
We will learn about wrangling dataframes in a different post.
</p>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/dataframes/</guid>
  <pubDate>Fri, 12 May 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Factors</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/introduction-to-factors/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-04-30-factors.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt about dataframes. In this post, we will learn about factors.
</p>
<ul>
<li>
create factors
</li>
<li>
order levels
</li>
<li>
specify labels
</li>
<li>
check levels
</li>
<li>
number of levels
</li>
</ul>
<p>
Categorical or qualitative data in R is treated as data type <code>factor</code>.
</p>
</section>
<section id="create-factors" class="level2">
<h2 class="anchored" data-anchor-id="create-factors">
Create Factors
</h2>
<pre class="r"><code>args(factor)</code></pre>
<pre><code>## function (x = character(), levels, labels = levels, exclude = NA, 
##     ordered = is.ordered(x), nmax = NA) 
## NULL</code></pre>
<pre class="r"><code>devices &lt;- factor(c('Mobile', 'Tablet', 'Desktop'))
devices</code></pre>
<pre><code>## [1] Mobile  Tablet  Desktop
## Levels: Desktop Mobile Tablet</code></pre>
<pre class="r"><code># number of levels
nlevels(devices)</code></pre>
<pre><code>## [1] 3</code></pre>
<pre class="r"><code># levels
levels(devices)</code></pre>
<pre><code>## [1] "Desktop" "Mobile"  "Tablet"</code></pre>
<p>
We will learn more about factors in a later post.
</p>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/introduction-to-factors/</guid>
  <pubDate>Sun, 30 Apr 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Lists</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/introduction-to-lists/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-04-18-lists.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt about matrices. In this post, we will learn about lists. Lists are very useful as they are heterogeneous i.e.&nbsp;they can contain different data types. If you remember, vectors and matrices are homogeneous i.e.&nbsp;they can contain only one type of data. If you include different data types, they will all be coerced to the same type. With lists, it is different. We learnt how to create lists briefly in the previous post while naming the rows and columns of a matrix. In this post, we will delve deeper into lists.
</p>
<ul>
<li>
how to create lists
</li>
<li>
access list elements
</li>
<li>
name list elements
</li>
<li>
coerce other R objects to list
</li>
<li>
coerce list to other R objects
</li>
</ul>
</section>
<section id="creating-lists" class="level2">
<h2 class="anchored" data-anchor-id="creating-lists">
Creating Lists
</h2>
<p>
To create a list, we use the <code>list()</code> function. Let us create a simple list to demonstrate how they can contain different data types.
</p>
<pre class="r"><code># numeric vector
vect1 &lt;- 1:10

# character vector 
vect2 &lt;- c('Jack', 'John', 'Jill')

# logical vector
vect3 &lt;- c(TRUE, FALSE)

# matrix
mat &lt;- matrix(data = 1:9, nrow = 3)

# list
l &lt;- list(vect1, vect2, vect3, mat)
l</code></pre>
<pre><code>## [[1]]
##  [1]  1  2  3  4  5  6  7  8  9 10
## 
## [[2]]
## [1] "Jack" "John" "Jill"
## 
## [[3]]
## [1]  TRUE FALSE
## 
## [[4]]
##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
<p>
If you observe the output, all the elements of the list retain their data types. Now let us learn how to access the elements of the list.
</p>
</section>
<section id="access-list-elements" class="level2">
<h2 class="anchored" data-anchor-id="access-list-elements">
Access List Elements
</h2>
<p>
You can access the elements of a list using the following operators:
</p>
<ul>
<li>
<code>[[</code>
</li>
<li>
<code>[</code>
</li>
<li>
<code><img src="https://latex.codecogs.com/png.latex?%3C/code%3E%3C/li%3E%0A%3C/ul%3E%0A%3Cp%3ELet%20us%20try%20them%20one%20by%20one.%3C/p%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3E#%20using%20%5B%5B%0Al%5B%5B1%5D%5D%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%20%5B1%5D%20%201%20%202%20%203%20%204%20%205%20%206%20%207%20%208%20%209%2010%3C/code%3E%3C/pre%3E%0A%3Cpre%20class=%22r%22%3E%3Ccode%3E#%20using%20%5B%0Al%5B1%5D%3C/code%3E%3C/pre%3E%0A%3Cpre%3E%3Ccode%3E##%20%5B%5B1%5D%5D%0A##%20%20%5B1%5D%20%201%20%202%20%203%20%204%20%205%20%206%20%207%20%208%20%209%2010%3C/code%3E%3C/pre%3E%0A%3Cp%3E%3Ccode%3E%5B%5B%3C/code%3E%20returns%20a%20vector%20while%20%3Ccode%3E%5B%3C/code%3E%20returns%20a%20list.%20The%20%3Ccode%3E"></code> operator can be used only when we have named elements in the list. Let us add names to the elements. Use the <code>names()</code> function to add names to the list.
<p></p>
<pre class="r"><code># named elements
names(l) &lt;- c('vect1', 'vect2', 'vect3', 'mat')
l</code></pre>
<pre><code>## $vect1
##  [1]  1  2  3  4  5  6  7  8  9 10
## 
## $vect2
## [1] "Jack" "John" "Jill"
## 
## $vect3
## [1]  TRUE FALSE
## 
## $mat
##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
<pre class="r"><code># use $
l$vect1</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
<pre class="r"><code># use [[
l[['vect1']]</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
<pre class="r"><code># use [
l['vect1']</code></pre>
<pre><code>## $vect1
##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
</li></ul></section> ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/introduction-to-lists/</guid>
  <pubDate>Tue, 18 Apr 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Matrices - Part 2</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/matrix-part-2/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-04-12-matrix-part-2.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt to create matrices. In this post, we will learn to:
</p>
<ul>
<li>
combining matrices
</li>
<li>
index/subset matrices
</li>
<li>
dissolve matrices
</li>
</ul>
</section>
<section id="append-data" class="level2">
<h2 class="anchored" data-anchor-id="append-data">
Append Data
</h2>
<p>
In this section, we will learn how to append data to a matrix. There are two functions that can be used for this purpose:
</p>
<ul>
<li>
<code>rbind()</code>
</li>
<li>
<code>cbind()</code>
</li>
</ul>
<p>
<code>cbind</code> will append a new column to the matrix while <code>rbind</code> will append a new row.
</p>
<section id="append-rowcolumn" class="level4">
<h4 class="anchored" data-anchor-id="append-rowcolumn">
Append Row/Column
</h4>
<pre class="r"><code># 3 x 3 matrix
mat &lt;- matrix(data = 1:9, nrow = 3)
mat</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
<pre class="r"><code># numeric vector
v &lt;- c(10, 11, 12)
v</code></pre>
<pre><code>## [1] 10 11 12</code></pre>
<pre class="r"><code># append row
rbind(mat, v)</code></pre>
<pre><code>##   [,1] [,2] [,3]
##      1    4    7
##      2    5    8
##      3    6    9
## v   10   11   12</code></pre>
<pre class="r"><code># append column
cbind(mat, v)</code></pre>
<pre><code>##             v
## [1,] 1 4 7 10
## [2,] 2 5 8 11
## [3,] 3 6 9 12</code></pre>
</section>
</section>
<section id="combine-matrices" class="level2">
<h2 class="anchored" data-anchor-id="combine-matrices">
Combine Matrices
</h2>
<p>
When you use <code>rbind</code> to combine two matrices, the number of columns must match and in case of <code>cbind</code>, the number of rows must match.
</p>
<section id="append-rowcolumn-1" class="level4">
<h4 class="anchored" data-anchor-id="append-rowcolumn-1">
Append Row/Column
</h4>
<pre class="r"><code># 3 x 3 matrix
mat1 &lt;- matrix(data = 1:9, nrow = 3)
mat2 &lt;- matrix(data = sample(9), nrow = 3)

# append rows
rbind(mat1, mat2)</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9
## [4,]    2    4    7
## [5,]    5    3    8
## [6,]    9    6    1</code></pre>
<pre class="r"><code># append columns
cbind(mat1, mat2)</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4] [,5] [,6]
## [1,]    1    4    7    2    4    7
## [2,]    2    5    8    5    3    8
## [3,]    3    6    9    9    6    1</code></pre>
</section>
</section>
<section id="subset-matrices" class="level2">
<h2 class="anchored" data-anchor-id="subset-matrices">
Subset Matrices
</h2>
<p>
In this section, we will learn to subset matrices. The <code>[]</code> operator can be used to subset matrices just like vectors but since matrices are two dimensional, we need to specify both the row number and the column number. Below are a few examples:
</p>
<pre class="r"><code># 3 x 4 matrix
mat &lt;- matrix(data = 1:12, nrow =3)
mat</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4]
## [1,]    1    4    7   10
## [2,]    2    5    8   11
## [3,]    3    6    9   12</code></pre>
<pre class="r"><code># extract element from first row, first column
mat[1, 1]</code></pre>
<pre><code>## [1] 1</code></pre>
<pre class="r"><code># extract all rows of first column
mat[, 1]</code></pre>
<pre><code>## [1] 1 2 3</code></pre>
<pre class="r"><code># extract all columns of first row
mat[1,]</code></pre>
<pre><code>## [1]  1  4  7 10</code></pre>
<pre class="r"><code># extract 2nd and 3rd row of first column
mat[c(2, 3), 1]</code></pre>
<pre><code>## [1] 2 3</code></pre>
<pre class="r"><code># extract 2nd and 3rd column of first row
mat[1, c(2, 3)]</code></pre>
<pre><code>## [1] 4 7</code></pre>
<pre class="r"><code># extract 2nd and 3rd row of first and third column
mat[c(2, 3), c(1, 3)]</code></pre>
<pre><code>##      [,1] [,2]
## [1,]    2    8
## [2,]    3    9</code></pre>
<section id="using-row-column-names" class="level4">
<h4 class="anchored" data-anchor-id="using-row-column-names">
Using Row &amp; Column Names
</h4>
<p>
In an earlier section, we learnt how to name the rows and columns of a matrix. Let us see how these names can be used to subset matrices.
</p>
<pre class="r"><code># row names
row_names &lt;- c('row_1', 'row_2', 'row_3')

# column names
col_names &lt;- c('col_1', 'col_2', 'col_3')

# matrix with row and column names
mat &lt;- matrix(data = 1:9, nrow = 3, dimnames = list(row_names, col_names))

# extract elements from first row/columns
mat['row_1', 'col_1']</code></pre>
<pre><code>## [1] 1</code></pre>
<pre class="r"><code># extract all rows of first column
mat[, 'col_1']</code></pre>
<pre><code>## row_1 row_2 row_3 
##     1     2     3</code></pre>
<pre class="r"><code># extract all columns of first row
mat['row_1',]</code></pre>
<pre><code>## col_1 col_2 col_3 
##     1     4     7</code></pre>
</section>
<section id="using-logical-expressions" class="level4">
<h4 class="anchored" data-anchor-id="using-logical-expressions">
Using Logical Expressions
</h4>
<p>
We can use logical expressions to subset matrices.
</p>
<pre class="r"><code># 3 x 4 matrix
mat &lt;- matrix(data = 1:12, nrow =3)
mat</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4]
## [1,]    1    4    7   10
## [2,]    2    5    8   11
## [3,]    3    6    9   12</code></pre>
<pre class="r"><code># elements greater than 4
mat &gt; 4</code></pre>
<pre><code>##       [,1]  [,2] [,3] [,4]
## [1,] FALSE FALSE TRUE TRUE
## [2,] FALSE  TRUE TRUE TRUE
## [3,] FALSE  TRUE TRUE TRUE</code></pre>
<pre class="r"><code># extract elements greater than 4
mat[mat &gt; 4]</code></pre>
<pre><code>## [1]  5  6  7  8  9 10 11 12</code></pre>
</section>
</section>
<section id="dissolve-matrices" class="level2">
<h2 class="anchored" data-anchor-id="dissolve-matrices">
Dissolve Matrices
</h2>
<p>
Till now we have learnt how to coerce a vector into matrix. Now let us learn how to coerce a matrix into a vector using:
</p>
<ul>
<li>
<code>c()</code>
</li>
<li>
<code>as.vector()</code>
</li>
</ul>
<pre class="r"><code># 3 x 3 matrix
mat &lt;- matrix(data = 1:9, nrow =3)
mat</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
<pre class="r"><code># using c()
c(mat)</code></pre>
<pre><code>## [1] 1 2 3 4 5 6 7 8 9</code></pre>
<pre class="r"><code># using as.vector()
as.vector(mat)</code></pre>
<pre><code>## [1] 1 2 3 4 5 6 7 8 9</code></pre>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/matrix-part-2/</guid>
  <pubDate>Wed, 12 Apr 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Matrices - Part 1</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/matrix-part-1/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-04-06-matrix.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt to index/subset vectors. In this post, we will learn to create matrices. A matrix is a regular array of data elements, arranged in rows and columns. Matrices in R are homogeneous i.e.&nbsp;they can hold a single type of data. Matrix elements are indexed by specifying the row and column index and the elements of a matrix can filled either by row or column. In the first section, we look at various ways of creating matrices in R.
</p>
</section>
<section id="creating-matrix" class="level2">
<h2 class="anchored" data-anchor-id="creating-matrix">
Creating Matrix
</h2>
<p>
The easiest way to create a matrix in R is to use the <code>matrix()</code> function. Let us look at its syntax:
</p>
<pre class="r"><code>args(matrix)</code></pre>
<pre><code>## function (data = NA, nrow = 1, ncol = 1, byrow = FALSE, dimnames = NULL) 
## NULL</code></pre>
<p>
Now that we have understood the syntax of the <code>matrix()</code> function, let us create a simple numeric matrix.
</p>
<pre class="r"><code># matrix of 3 rows filled by columns
mat &lt;- matrix(data = 1:9, nrow = 3, byrow = FALSE)
mat</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
<p>
In the above example, we created a matrix of 3 rows where the data elements are filled by columns. We need to specify either the number of <code>rows</code> or <code>columns</code> and R will automatically compute the other. The number of data elements should be equal to the product of the rows and columns, else R will return a warning.
</p>
<pre class="r"><code>matrix(data = 1:9, nrow = 2, byrow = FALSE)</code></pre>
<pre><code>## Warning in matrix(data = 1:9, nrow = 2, byrow = FALSE): data length [9] is not a
## sub-multiple or multiple of the number of rows [2]</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4] [,5]
## [1,]    1    3    5    7    9
## [2,]    2    4    6    8    1</code></pre>
<pre class="r"><code>matrix(data = 1:10, nrow = 3, byrow = FALSE)</code></pre>
<pre><code>## Warning in matrix(data = 1:10, nrow = 3, byrow = FALSE): data length [10] is not
## a sub-multiple or multiple of the number of rows [3]</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4]
## [1,]    1    4    7   10
## [2,]    2    5    8    1
## [3,]    3    6    9    2</code></pre>
<p>
We can follow some general rules to avoid the mistakes made in the previous examples:
</p>
<ul>
<li>
<p>
if the number of elements is odd, both the number of rows and columns must be odd and their product should equal the number of data elements
</p>
</li>
<li>
<p>
if the number of elements is even, either the number of rows or columns must be even. In some cases, both the rows and columns must be even
</p>
</li>
</ul>
<p>
Let us continue to explore the syntax of the <code>matrix()</code> function.
</p>
<section id="fill-data-by-row" class="level4">
<h4 class="anchored" data-anchor-id="fill-data-by-row">
Fill Data by Row
</h4>
<pre class="r"><code>matrix(data = 1:9, nrow = 3, byrow = TRUE)</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    2    3
## [2,]    4    5    6
## [3,]    7    8    9</code></pre>
</section>
<section id="fill-data-by-column" class="level4">
<h4 class="anchored" data-anchor-id="fill-data-by-column">
Fill Data by Column
</h4>
<pre class="r"><code>matrix(data = 1:9, nrow = 3, byrow = FALSE)</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    4    7
## [2,]    2    5    8
## [3,]    3    6    9</code></pre>
</section>
<section id="specify-rows" class="level4">
<h4 class="anchored" data-anchor-id="specify-rows">
Specify Rows
</h4>
<pre class="r"><code>matrix(data = 1:10, nrow = 2)</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4] [,5]
## [1,]    1    3    5    7    9
## [2,]    2    4    6    8   10</code></pre>
</section>
<section id="specify-columns" class="level4">
<h4 class="anchored" data-anchor-id="specify-columns">
Specify Columns
</h4>
<pre class="r"><code>matrix(data = 1:10, ncol = 5)</code></pre>
<pre><code>##      [,1] [,2] [,3] [,4] [,5]
## [1,]    1    3    5    7    9
## [2,]    2    4    6    8   10</code></pre>
</section>
</section>
<section id="row-column-names" class="level2">
<h2 class="anchored" data-anchor-id="row-column-names">
Row &amp; Column Names
</h2>
<p>
You can specify names for the rows and columns of a matrix. To do so, we need to use <code>list</code>. Lists can contain other data structures such as vectors, matrices and even other lists. They are heterogeneous i.e.&nbsp;they can contain different data types. We will learn more about lists in a future post, for the time being let us learn how to create a basic list using the <code>list()</code> function:
</p>
<pre class="r"><code># character vector
first_name &lt;-   c("John", "Jill", "Jack")

# numeric vector
age &lt;- c(20, 24, 32)

# list 
details &lt;- list(first_name, age)
details</code></pre>
<pre><code>## [[1]]
## [1] "John" "Jill" "Jack"
## 
## [[2]]
## [1] 20 24 32</code></pre>
<p>
Now that we know how to create a list, let us go ahead and create a matrix and name its rows and columns.
</p>
<pre class="r"><code># row names
row_names &lt;- c('row_1', 'row_2', 'row_3')

# column names
col_names &lt;- c('col_1', 'col_2', 'col_3')

# matrix with row and column names
matrix(data = 1:9, nrow = 3, dimnames = list(row_names, col_names))</code></pre>
<pre><code>##       col_1 col_2 col_3
## row_1     1     4     7
## row_2     2     5     8
## row_3     3     6     9</code></pre>
</section>
<section id="matrix-dimension" class="level2">
<h2 class="anchored" data-anchor-id="matrix-dimension">
Matrix Dimension
</h2>
<p>
Another useful function is <code>dim()</code>. It can be used to:
</p>
<ul>
<li>
check the dimension (rows and columns) of a matrix
</li>
<li>
modify the dimension of a matrix
</li>
<li>
coerce a vector to a matrix
</li>
</ul>
<section id="check-dimension-of-a-matrix" class="level4">
<h4 class="anchored" data-anchor-id="check-dimension-of-a-matrix">
Check dimension of a matrix
</h4>
<pre class="r"><code>mat &lt;- matrix(data = 1:9, nrow = 3, byrow = TRUE)
dim(mat)</code></pre>
<pre><code>## [1] 3 3</code></pre>
</section>
<section id="modify-dimension-of-a-matrix" class="level4">
<h4 class="anchored" data-anchor-id="modify-dimension-of-a-matrix">
Modify dimension of a matrix
</h4>
<pre class="r"><code># sample matrix
mat &lt;- matrix(data = 1:12, nrow = 3, byrow = TRUE)
dim(mat)</code></pre>
<pre><code>## [1] 3 4</code></pre>
<pre class="r"><code># change the dimension to 4 x 3
dim(mat) &lt;- c(4, 3)
dim(mat)</code></pre>
<pre><code>## [1] 4 3</code></pre>
</section>
<section id="coerce-vector-to-matrix" class="level4">
<h4 class="anchored" data-anchor-id="coerce-vector-to-matrix">
Coerce vector to matrix
</h4>
<pre class="r"><code># numeric vector
vect1 &lt;- 1:12
vect1</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10 11 12</code></pre>
<pre class="r"><code># coerce vect1 to a 4 x 3 matrix
dim(vect1) &lt;- c(4, 3)
vect1</code></pre>
<pre><code>##      [,1] [,2] [,3]
## [1,]    1    5    9
## [2,]    2    6   10
## [3,]    3    7   11
## [4,]    4    8   12</code></pre>
<p>
Another way to coerce an R data structure to <code>matrix</code> is to use the <code>as.matrix()</code> function. Since the only other data structure we have learnt so far is the vector, we will coerce a vector into a matric using <code>as.matrix()</code>. We will deal with the other data structures as and when we learn them.
</p>
<pre class="r"><code># numeric vector
vect1 &lt;- 1:12
vect1</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10 11 12</code></pre>
<pre class="r"><code># coerce vect1 to a matrix
as.matrix(vect1)</code></pre>
<pre><code>##       [,1]
##  [1,]    1
##  [2,]    2
##  [3,]    3
##  [4,]    4
##  [5,]    5
##  [6,]    6
##  [7,]    7
##  [8,]    8
##  [9,]    9
## [10,]   10
## [11,]   11
## [12,]   12</code></pre>
<p>
Regardless of the data type of the vector, all of them will be coerced to a matrix of dimension <code>n x 1</code> i.e.&nbsp;they will all have only one column.
</p>
</section>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/matrix-part-1/</guid>
  <pubDate>Thu, 06 Apr 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Vectors - Part 3</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/vectors-part-3/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-04-03-vectors-part-3.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt to perform simple operations on vector and handle missing values. In this post, we will learn to index/subset vectors.
</p>
</section>
<section id="index-vectors" class="level2">
<h2 class="anchored" data-anchor-id="index-vectors">
Index Vectors
</h2>
<p>
One of the most important steps in data analysis is selecting a subset of data from a bigger data set. Indexing helps in retrieving values individually or a set of values that meet a specific criteria. In this post, we look at various ways of indexing/subsetting vectors.
</p>
</section>
<section id="index-operator" class="level2">
<h2 class="anchored" data-anchor-id="index-operator">
Index Operator
</h2>
<p>
<code>[]</code> is the index operator in R. We can use various expressions within <code>[]</code> to subset data. In R, <strong>index positions begin at 1 and not 0</strong>. To begin with, let us look at values in different index positions:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  5  7  2  8  9  6 10  4  1  3</code></pre>
<pre class="r"><code># return third element
vect1[3]</code></pre>
<pre><code>## [1] 2</code></pre>
<pre class="r"><code># return seventh element
vect1[7]    </code></pre>
<pre><code>## [1] 10</code></pre>
<section id="out-of-range-index" class="level4">
<h4 class="anchored" data-anchor-id="out-of-range-index">
Out of range index
</h4>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  8  2  5  1 10  9  3  4  7  6</code></pre>
<pre class="r"><code># return value at index 0
vect1[0]</code></pre>
<pre><code>## integer(0)</code></pre>
<pre class="r"><code># length of the vector
length(vect1)</code></pre>
<pre><code>## [1] 10</code></pre>
<pre class="r"><code># out of range index
vect1[11]   </code></pre>
<pre><code>## [1] NA</code></pre>
<p>
In the first case, we specified the index as 0 and in the second case we used the index 11, which is greater than the length of the vector. R returns an empty vector in the first case and <code>NA</code> in the second case.
</p>
</section>
<section id="negative-index" class="level4">
<h4 class="anchored" data-anchor-id="negative-index">
Negative Index
</h4>
<p>
Using a negative index will delete the value in the said index position. Unlike other languages, it will not index elements from the end of the vector counting backwards. Let us look at an example to understand how negative index works in R:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  6  9  3  2 10  8  4  7  1  5</code></pre>
<pre class="r"><code># drop third element
vect1[-3]</code></pre>
<pre><code>## [1]  6  9  2 10  8  4  7  1  5</code></pre>
<pre class="r"><code># drop seventh element
vect1[-7]   </code></pre>
<pre><code>## [1]  6  9  3  2 10  8  7  1  5</code></pre>
</section>
</section>
<section id="subset-multiple-elements" class="level2">
<h2 class="anchored" data-anchor-id="subset-multiple-elements">
Subset Multiple Elements
</h2>
<p>
If we do not specify anything within <code>[]</code>, all the elements in the vector will be returned. We can specify the index elements using any expression that generates a sequence of integers. Let us look at a few examples:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  2  8  1  9  7 10  5  4  3  6</code></pre>
<pre class="r"><code># return all elements
vect1[]</code></pre>
<pre><code>##  [1]  2  8  1  9  7 10  5  4  3  6</code></pre>
<pre class="r"><code># return first 5 values
vect1[1:5]</code></pre>
<pre><code>## [1] 2 8 1 9 7</code></pre>
<pre class="r"><code># return all values from the 5th position
end &lt;- length(vect1)
vect1[5:end]</code></pre>
<pre><code>## [1]  7 10  5  4  3  6</code></pre>
<p>
If you are using the colon to generate the index positions, you will have to specify both the starting and ending position, else, R will return an error.
</p>
<p>
What if we want elements that are not in a sequence as we saw in the last example? In such cases, we have to create a vector using <code>c()</code> and use it to extract elements from the original vector. Below is an example:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  7  4 10  3  9  8  5  1  2  6</code></pre>
<pre class="r"><code># extract 2nd, 5th and 7th element
select &lt;- c(2, 5, 7)
vect1[select]</code></pre>
<pre><code>## [1] 4 9 5</code></pre>
<pre class="r"><code># extract elements in position 1 to 4, 6 and 9
select &lt;- c(1:4, 6, 9)
vect1[select]</code></pre>
<pre><code>## [1]  7  4 10  3  8  2</code></pre>
</section>
<section id="subset-named-vectors" class="level2">
<h2 class="anchored" data-anchor-id="subset-named-vectors">
Subset Named Vectors
</h2>
<p>
Vectors can be subset using the name of the elements. <strong>When using name of elements for subsetting, ensure that the names are enclosed in single or double quotations</strong>, else R will return an error. Let us look at a few examples:
</p>
<pre class="r"><code>vect1 &lt;- c(score1 = 8, score2 = 6, score3 = 9)
vect1</code></pre>
<pre><code>## score1 score2 score3 
##      8      6      9</code></pre>
<pre class="r"><code># extract score2
vect1['score2']</code></pre>
<pre><code>## score2 
##      6</code></pre>
<pre class="r"><code># extract score1 and score3
vect1[c('score1', 'score3')]</code></pre>
<pre><code>## score1 score3 
##      8      9</code></pre>
</section>
<section id="subset-using-logical-values" class="level2">
<h2 class="anchored" data-anchor-id="subset-using-logical-values">
Subset using logical values
</h2>
<p>
Logical values can be used to subset vectors. They are not very flexible but can be used for simple indexing. In all of the below examples, the logical vectors are recycled to match the length of the vector from which we subset data:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  8  1  4  5 10  9  3  6  2  7</code></pre>
<pre class="r"><code># returns all values
vect1[TRUE]</code></pre>
<pre><code>##  [1]  8  1  4  5 10  9  3  6  2  7</code></pre>
<pre class="r"><code># empty vector
vect1[FALSE]</code></pre>
<pre><code>## integer(0)</code></pre>
<pre class="r"><code># values in odd positions
vect1[c(TRUE, FALSE)]</code></pre>
<pre><code>## [1]  8  4 10  3  2</code></pre>
<pre class="r"><code># values in even positions
vect1[c(FALSE, TRUE)]</code></pre>
<pre><code>## [1] 1 5 9 6 7</code></pre>
</section>
<section id="subset-using-logical-expressions" class="level2">
<h2 class="anchored" data-anchor-id="subset-using-logical-expressions">
Subset using logical expressions
</h2>
<p>
Logical expressions can be used to extract elements that meet specific criteria. This method is most flexible and useful as we can combine multiple conditions using relational and logical operators. Before we use logical expressions, let us spend some time understanding comparison and logical operators as we will be using them extensively hereafter.
</p>
<section id="comparison-operators" class="level4">
<h4 class="anchored" data-anchor-id="comparison-operators">
Comparison Operators
</h4>
<p>
When you create an expression using a comparison operator, the output is always a logical value i.e.&nbsp;<code>TRUE</code> or <code>FALSE</code>. Let us see how we can use comparison operators to subset data:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  3  1  9 10  5  8  2  6  4  7</code></pre>
<pre class="r"><code># return elements greater than 5
vect1 &gt; 5</code></pre>
<pre><code>##  [1] FALSE FALSE  TRUE  TRUE FALSE  TRUE FALSE  TRUE FALSE  TRUE</code></pre>
<pre class="r"><code>vect1[vect1 &gt; 5]</code></pre>
<pre><code>## [1]  9 10  8  6  7</code></pre>
<pre class="r"><code># return elements greater than or equal to 5
vect1 &gt;= 5</code></pre>
<pre><code>##  [1] FALSE FALSE  TRUE  TRUE  TRUE  TRUE FALSE  TRUE FALSE  TRUE</code></pre>
<pre class="r"><code>vect1[vect1 &gt;= 5]</code></pre>
<pre><code>## [1]  9 10  5  8  6  7</code></pre>
<pre class="r"><code># return elements lesser than 5
vect1 &lt; 5</code></pre>
<pre><code>##  [1]  TRUE  TRUE FALSE FALSE FALSE FALSE  TRUE FALSE  TRUE FALSE</code></pre>
<pre class="r"><code>vect1[vect1 &lt; 5]</code></pre>
<pre><code>## [1] 3 1 2 4</code></pre>
<pre class="r"><code># return elements lesser than or equal to 5
vect1 &lt;= 5</code></pre>
<pre><code>##  [1]  TRUE  TRUE FALSE FALSE  TRUE FALSE  TRUE FALSE  TRUE FALSE</code></pre>
<pre class="r"><code>vect1[vect1 &lt;= 5]</code></pre>
<pre><code>## [1] 3 1 5 2 4</code></pre>
<pre class="r"><code># return elements equal to 5
vect1 == 5</code></pre>
<pre><code>##  [1] FALSE FALSE FALSE FALSE  TRUE FALSE FALSE FALSE FALSE FALSE</code></pre>
<pre class="r"><code>vect1[vect1 == 5]</code></pre>
<pre><code>## [1] 5</code></pre>
<pre class="r"><code># return elements not equal to 5
vect1 != 5</code></pre>
<pre><code>##  [1]  TRUE  TRUE  TRUE  TRUE FALSE  TRUE  TRUE  TRUE  TRUE  TRUE</code></pre>
<pre class="r"><code>vect1[vect1 != 5]</code></pre>
<pre><code>## [1]  3  1  9 10  8  2  6  4  7</code></pre>
</section>
</section>
<section id="logical-operators" class="level2">
<h2 class="anchored" data-anchor-id="logical-operators">
Logical Operators
</h2>
<p>
Let us combine comparison and logical operators to create expressions and use them to subset vectors:
</p>
<pre class="r"><code># random sample of 10 values
vect1 &lt;- sample(10)
vect1</code></pre>
<pre><code>##  [1]  3  2  9  7  5 10  4  8  6  1</code></pre>
<pre class="r"><code># return all elements less than 8 or divisible by 3
vect1[(vect1 &lt; 8 | (vect1 %% 3 == 0))]</code></pre>
<pre><code>## [1] 3 2 9 7 5 4 6 1</code></pre>
<pre class="r"><code># return all elements less than 7 or divisible by 2
vect1[(vect1 &lt; 7 | (vect1 %% 2 == 0))]</code></pre>
<pre><code>## [1]  3  2  5 10  4  8  6  1</code></pre>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/vectors-part-3/</guid>
  <pubDate>Mon, 03 Apr 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Vectors - Part 2</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/vectors-part-2/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-03-29-vectors-part-2.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt to create vectors of different data types. In this post, we will learn to
</p>
<ul>
<li>
coerce different data types
</li>
<li>
perform simple operations on vectors
</li>
<li>
handle missing data
</li>
<li>
index/subset vectors
</li>
</ul>
</section>
<section id="naming-vector-elements" class="level2">
<h2 class="anchored" data-anchor-id="naming-vector-elements">
Naming Vector Elements
</h2>
<p>
It is possible to name the different elements of a vector. The advantage of naming vector elements is that we can later on use these names to access the elements. Use <code>names()</code> to specify the names of a vector. You can specify the names while creating the vector or add them later.
</p>
<section id="method-1-create-vector-and-add-names-later" class="level4">
<h4 class="anchored" data-anchor-id="method-1-create-vector-and-add-names-later">
Method 1: Create vector and add names later
</h4>
<pre class="r"><code># create vector and add names later
vect1 &lt;- c(1, 2, 3)

# name the elements of the vector
names(vect1) &lt;- c("One", "Two", "Three")

# call vect1
vect1
##   One   Two Three 
##     1     2     3</code></pre>
</section>
<section id="method-2-specify-names-while-creating-vector" class="level4">
<h4 class="anchored" data-anchor-id="method-2-specify-names-while-creating-vector">
Method 2: Specify names while creating vector
</h4>
<pre class="r"><code># specify names while creating vector
vect2 &lt;- c(John = 1, Jack = 2, Jill = 3, Jovial = 4)

# call vect2
vect2
##   John   Jack   Jill Jovial 
##      1      2      3      4</code></pre>
</section>
</section>
<section id="vector-coercion" class="level2">
<h2 class="anchored" data-anchor-id="vector-coercion">
Vector Coercion
</h2>
<p>
Vectors are homogeneous i.e.&nbsp;all the elements of the vector must be of the same type. If we try to create a vector by combining different data types, the elements will be coerced to the most flexible type. The below table shows the order in which coercion occurs.
</p>
<p>
<code>character</code> data type is the most flexible while <code>logical</code> data type is the least flexible. If you try to combine any other data type with <code>character</code>, all the elements will be coerced to type <code>character</code>. In the absence of <code>character</code> data, all elements will be coerced to <code>numeric</code>. Finally, if the data does not include <code>character</code> or <code>numeric</code> types, all the elements will be coerced to <code>integer</code> type.
</p>
<section id="case-1-different-data-types" class="level4">
<h4 class="anchored" data-anchor-id="case-1-different-data-types">
Case 1: Different Data Types
</h4>
<pre class="r"><code># vector of different data types
vect1 &lt;- c(1, 1L, 'one', TRUE)

# call vect1
vect1
## [1] "1"    "1"    "one"  "TRUE"

# check data type
class(vect1)
## [1] "character"</code></pre>
</section>
<section id="case-2-numeric-integer-and-logical" class="level4">
<h4 class="anchored" data-anchor-id="case-2-numeric-integer-and-logical">
Case 2: Numeric, Integer and Logical
</h4>
<pre class="r"><code># vector of different data types
vect1 &lt;- c(1, 1L, TRUE)

# call vect1
vect1
## [1] 1 1 1

# check data type
class(vect1)
## [1] "numeric"</code></pre>
</section>
<section id="case-integer-and-logical" class="level4">
<h4 class="anchored" data-anchor-id="case-integer-and-logical">
Case : Integer and Logical
</h4>
<pre class="r"><code># vector of different data types
vect1 &lt;- c(1L, TRUE)

# call vect1
vect1
## [1] 1 1

# check data type
class(vect1)
## [1] "integer"</code></pre>
<p>
To summarize, below is the order in which coercion takes place:
</p>
</section>
</section>
<section id="vector-operations" class="level2">
<h2 class="anchored" data-anchor-id="vector-operations">
Vector Operations
</h2>
<p>
In this section, we look at simple operations that can be performed on vectors in R. Remember that the nature of the operations depends upon the type of data. Below are a few examples:
</p>
<section id="case-1-vectors-of-same-length" class="level4">
<h4 class="anchored" data-anchor-id="case-1-vectors-of-same-length">
Case 1: Vectors of same length
</h4>
<pre class="r"><code># create two vectors
vect1 &lt;- c(1, 3, 8, 4)
vect2 &lt;- c(2, 7, 1, 9)

# addition
vect1 + vect2
## [1]  3 10  9 13

# subtraction
vect1 - vect2
## [1] -1 -4  7 -5

# multiplication
vect1 * vect2
## [1]  2 21  8 36

# division
vect1 / vect2
## [1] 0.5000000 0.4285714 8.0000000 0.4444444</code></pre>
</section>
<section id="case-2-vectors-of-different-length" class="level4">
<h4 class="anchored" data-anchor-id="case-2-vectors-of-different-length">
Case 2: Vectors of different length
</h4>
<p>
In the previous case, the length i.e.&nbsp;the number of elements in the vectors were same. What happens if the length of the vectors are unequal? In such cases, the shorter vector is recycled to match the length of the longer vector. The below example should clear this concept:
</p>
<pre class="r"><code># create two vectors
vect1 &lt;- c(2, 7)
vect2 &lt;- c(1, 8, 5, 2)

# addition
vect1 + vect2
## [1]  3 15  7  9

# subtraction
vect1 - vect2
## [1]  1 -1 -3  5

# multiplication
vect1 * vect2
## [1]  2 56 10 14

# division
vect1 / vect2
## [1] 2.000 0.875 0.400 3.500</code></pre>
</section>
</section>
<section id="missing-data" class="level2">
<h2 class="anchored" data-anchor-id="missing-data">
Missing Data
</h2>
<p>
Missing data is a reality. No matter how careful you are in collecting data for your analysis, chances are always high that you end up with some missing data. In R missing values are represented by <code>NA</code>. In this section, we will focus on the following:
</p>
<ul>
<li>
test for missing data
</li>
<li>
remove missing data
</li>
<li>
exclude missing data from analysis
</li>
</ul>
<section id="detect-missing-data" class="level4">
<h4 class="anchored" data-anchor-id="detect-missing-data">
Detect missing data
</h4>
<p>
We first create a vector with missing values. After that, we will use <code>is.na()</code> to test whether the data contains missing values. <code>is.na()</code> returns a logical vector equal to the length of the vector being tested. Another function that can be used for detecting missing values is <code>complete.cases()</code>. Below is an example:
</p>
<pre class="r"><code># vector with missing values
vect1 &lt;- c(1, 3, NA, 5, 2)

# use is.na
is.na(vect1)
## [1] FALSE FALSE  TRUE FALSE FALSE

# use complete.cases
complete.cases(vect1)
## [1]  TRUE  TRUE FALSE  TRUE  TRUE</code></pre>
</section>
<section id="omit-missing-data" class="level4">
<h4 class="anchored" data-anchor-id="omit-missing-data">
Omit missing data
</h4>
<p>
In the presensce of missing data, all computations in R will return <code>NA</code>. To avoid this, we might want to remove the missing data before doing any computation. <code>na.omit()</code> will remove all missing values from the data. Let us look at an example:
</p>
<pre class="r"><code># vector with missing values
vect1 &lt;- c(1, 3, NA, 5, 2)

# call vect1
vect1
## [1]  1  3 NA  5  2

# omit missing values
na.omit(vect1)
## [1] 1 3 5 2
## attr(,"na.action")
## [1] 3
## attr(,"class")
## [1] "omit"</code></pre>
</section>
<section id="exclude-missing-data" class="level4">
<h4 class="anchored" data-anchor-id="exclude-missing-data">
Exclude missing data
</h4>
<p>
To exclude missing values from computations, use <code>na.rm</code> and set it to <code>TRUE</code>.
</p>
<pre class="r"><code># vector with missing values
vect1 &lt;- c(1, 3, NA, 5, 2)

# compute mean
mean(vect1)
## [1] NA

# compute mean by excluding missing value
mean(vect1, na.rm = TRUE)
## [1] 2.75</code></pre>
</section>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/vectors-part-2/</guid>
  <pubDate>Wed, 29 Mar 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Vectors - Part 1</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/vectors-part-1/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-03-25-vectors.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt about the basic data types in R. IN this post, we will:
</p>
<ul>
<li>
understand the concept of vectors
</li>
<li>
learn to create vectors of different data types
</li>
</ul>
</section>
<section id="vectors" class="level2">
<h2 class="anchored" data-anchor-id="vectors">
Vectors
</h2>
<p>
Vector is the most basic data structure in R. It is a sequence of elements of the same data type. If the elements are of different data types, they will be coerced to a common type that can accommodate all the elements. Vectors are generally created using <code>c()</code> (concatenate function), although depending on the data type of vector being created, other methods can be used.
</p>
</section>
<section id="numeric-vector" class="level2">
<h2 class="anchored" data-anchor-id="numeric-vector">
Numeric Vector
</h2>
<p>
We will create a numeric vector using <code>c()</code> but you can use any function that creates a sequence of numbers. After that we will use <code>is.vector()</code> to check if it is a vector and <code>class</code> to check the data type.
</p>
<pre class="r"><code># create a numeric vector
num_vect &lt;- c(1, 2, 3)

# display the vector
num_vect</code></pre>
<pre><code>## [1] 1 2 3</code></pre>
<pre class="r"><code># check if it is a vector
is.vector(num_vect)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code># check data type
class(num_vect)</code></pre>
<pre><code>## [1] "numeric"</code></pre>
<p>
Let us look at other ways to create a sequence of numbers. We leave it as an exercise to the reader to understand the functions using <code>help</code>.
</p>
<pre class="r"><code># using colon
vect1 &lt;- 1:10
vect1</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
<pre class="r"><code># using rep
vect2 &lt;- rep(1, 5)
vect2</code></pre>
<pre><code>## [1] 1 1 1 1 1</code></pre>
<pre class="r"><code># using seq
vect3 &lt;- seq(10)
vect3</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
</section>
<section id="integer-vector" class="level2">
<h2 class="anchored" data-anchor-id="integer-vector">
Integer Vector
</h2>
<p>
Creating an integer vector is similar to numeric vector except that we need to instruct R to treat the data as <code>integer</code> and not <code>numeric</code> or <code>double</code>. We will use the same methods we used for creating numeric vectors. To specify that the data is of type <code>integer</code>, we suffix the number with <code>L</code>.
</p>
<pre class="r"><code># integer vector
int_vect &lt;- c(1L, 2L, 3L)
int_vect</code></pre>
<pre><code>## [1] 1 2 3</code></pre>
<pre class="r"><code># check data type
class(int_vect)</code></pre>
<pre><code>## [1] "integer"</code></pre>
<pre class="r"><code># using colon
vect1 &lt;- 1L:10L
vect1</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
<pre class="r"><code># using rep
vect2 &lt;- rep(1L, 5)
vect2</code></pre>
<pre><code>## [1] 1 1 1 1 1</code></pre>
<pre class="r"><code># using seq
vect3 &lt;- seq(10L)
vect3</code></pre>
<pre><code>##  [1]  1  2  3  4  5  6  7  8  9 10</code></pre>
</section>
<section id="character-vector" class="level2">
<h2 class="anchored" data-anchor-id="character-vector">
Character Vector
</h2>
<p>
A character vector may contain a single character, a word or a group of words. The elements must be enclosed in single or double quotations.
</p>
<pre class="r"><code># character vector
greetings &lt;- c("hello", "good morning")
greetings</code></pre>
<pre><code>## [1] "hello"        "good morning"</code></pre>
<pre class="r"><code># check data type
class(greetings)</code></pre>
<pre><code>## [1] "character"</code></pre>
</section>
<section id="logical-vector" class="level2">
<h2 class="anchored" data-anchor-id="logical-vector">
Logical Vector
</h2>
<p>
A vector of logical values will either contain <code>TRUE</code> or <code>FALSE</code> or both.
</p>
<pre class="r"><code># logical vector
vect_logic &lt;- c(TRUE, FALSE, TRUE, TRUE, FALSE)
vect_logic</code></pre>
<pre><code>## [1]  TRUE FALSE  TRUE  TRUE FALSE</code></pre>
<pre class="r"><code># check data type
class(vect_logic)</code></pre>
<pre><code>## [1] "logical"</code></pre>
<p>
In fact, you can create an <code>integer</code> vector and coerce it to type <code>logical</code>.
</p>
<pre class="r"><code># integer vector
int_vect &lt;- rep(0L:1L, 3)
int_vect</code></pre>
<pre><code>## [1] 0 1 0 1 0 1</code></pre>
<pre class="r"><code># coerce to logical vector
log_vect &lt;- as.logical(int_vect)
log_vect</code></pre>
<pre><code>## [1] FALSE  TRUE FALSE  TRUE FALSE  TRUE</code></pre>
<pre class="r"><code># check data type
class(log_vect)</code></pre>
<pre><code>## [1] "logical"</code></pre>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/vectors-part-1/</guid>
  <pubDate>Sat, 25 Mar 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Beginners Guide to R Package Ecosystem</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/beginners-guide-to-r-package-ecosystem/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-03-13-beginners-guide-to-r-package-ecosystem.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt about getting help in R. In this post, we will learn about R packages. Packages are fundamental to R. There are approximately 15000 packages available on <a href="https://cran.r-project.org/">CRAN</a> or the Comprehensive R Archive Network.
</p>
<p>
Packages are available for different <a href="https://cran.r-project.org/web/views/">topics</a>. You should always look for a package before writing code from scratch. In case you have written your own codes for a new analysis or topic, do share it with the R community by converting the code into a package. You can learn more about building R packages from <a href="https://r-pkgs.had.co.nz/">R Packages</a>, a book written by <a href="https://hadley.nz/">Hadley Wickham</a>.
</p>
<p>
In this post, we will learn to:
</p>
<ul>
<li>
install R packages from
<ul>
<li>
CRAN
</li>
<li>
GitHub
</li>
<li>
BitBucket
</li>
<li>
Bioconductor
</li>
<li>
rForge
</li>
</ul>
</li>
<li>
install different versions of a package
</li>
<li>
load, update &amp; remove installed packages
</li>
<li>
access package documentation
</li>
</ul>
</section>
<section id="install-packages" class="level2">
<h2 class="anchored" data-anchor-id="install-packages">
Install Packages
</h2>
<section id="cran" class="level4">
<h4 class="anchored" data-anchor-id="cran">
CRAN
</h4>
<p>
Packages from CRAN can be installed using <code>install.packages()</code>. The name of the package must be enclosed in single or double quotes.
</p>
<pre class="r"><code>install.packages('ggplot2')</code></pre>
</section>
<section id="github" class="level4">
<h4 class="anchored" data-anchor-id="github">
GitHub
</h4>
<p>
Some R packages are made available on <a href="https://github.com/">GitHub</a> before releasing them on CRAN. Such packages can be installed using <code>install_github()</code> from <a href="https://cran.r-project.org/web/packages/devtools/index.html">devtools</a> or <a href="https://cran.r-project.org/web/packages/remotes/index.html">remotes</a> package. You need tp specify the name of the repository and the package. For example, to download <a href="https://ggplot2.tidyverse.org/">ggplot2</a> or <a href="https://dplyr.tidyverse.org/">dplyr</a>, below is the code:
</p>
<pre class="r"><code>devtools::install_github("tidyverse/ggplot2")
remotes::install_github("tidyverse/dplyr")</code></pre>
</section>
<section id="bitbucket" class="level4">
<h4 class="anchored" data-anchor-id="bitbucket">
BitBucket
</h4>
<p>
<a href="https://bitbucket.org/">Bitbucket</a> is similar to GitHub. You can install packages from Bitbucket using <code>install_bitbucket()</code> from devtools or remotes pacakge.
</p>
<pre class="r"><code>devtools::install_bitbucket("dannavarro/lsr-package")
remotes::install_bitbucket("dannavarro/lsr-package")</code></pre>
</section>
<section id="bioconductor" class="level4">
<h4 class="anchored" data-anchor-id="bioconductor">
Bioconductor
</h4>
<p>
<a href="https://www.bioconductor.org/">Bioconductor</a> provides tools for analysis and comprehension of high throughput genomic data. Packages hosted on Bioconductor can be installed in multiple ways:
</p>
<section id="devtools" class="level5">
<h5 class="anchored" data-anchor-id="devtools">
devtools
</h5>
<p>
Use <code>install_bioc()</code> from devtools.
</p>
<pre class="r"><code>install_bioc("SummarizedExperiment")</code></pre>
</section>
<section id="bioclite" class="level5">
<h5 class="anchored" data-anchor-id="bioclite">
biocLite
</h5>
<p>
Use <code>biocLite()</code> function.
</p>
<pre class="r"><code>source('http://bioconductor.org/biocLite.R')
biocLite('GenomicFeatures')</code></pre>
</section>
</section>
<section id="rforge" class="level4">
<h4 class="anchored" data-anchor-id="rforge">
rForge
</h4>
<p>
Many R packages are hosted at <a href="https://r-forge.r-project.org/">R-Forge</a>, a platform for development of R packages.
</p>
<pre class="r"><code>install.packages('quantstrat', repos = 'https://r-forge.r-project.org/')</code></pre>
</section>
</section>
<section id="install-different-versions" class="level2">
<h2 class="anchored" data-anchor-id="install-different-versions">
Install Different Versions
</h2>
<p>
Now that we have learnt how to install packages, let us look at installing different versions of the same package.
</p>
<pre class="r"><code>remotes::install_version('dplyr', version = 0.5.0)</code></pre>
<p>
If you want to install the latest release from GitHub, append <code><span class="citation" data-cites="*release">@*release</span></code> to the repository name. For example, to install the latest release of dplyr:
</p>
<pre class="r"><code>remotes::install_github('tidyverse/dplyr@*release')</code></pre>
</section>
<section id="installed-packages" class="level2">
<h2 class="anchored" data-anchor-id="installed-packages">
Installed Packages
</h2>
<ul>
<li>
<code>installed.packages()</code>: view currently installed packages
</li>
<li>
<code>library(‘package_name’)</code>: load packages
</li>
<li>
<code>autoload(‘function_name’, ‘package_name’)</code>: load functions and data from packages only when called in the script
</li>
<li>
<code>available.package()</code>: packages available for installation
</li>
<li>
<code>old.packages()</code>: packages which have new versions available
</li>
<li>
<code>new.packages()</code>: packages already not installed
</li>
<li>
<code>update.packages()</code>: update packages which have new versions
</li>
<li>
<code>remove.packages(‘package_name’)</code>: remove installed packages
</li>
</ul>
</section>
<section id="library-paths" class="level2">
<h2 class="anchored" data-anchor-id="library-paths">
Library Paths
</h2>
<p>
Library is a directory that contains all installed packages. Usually there will be more than one R library in your systme. You can find the location of the libraries using <code>.libPaths()</code>.
</p>
<pre class="r"><code>.libPaths()</code></pre>
<pre><code>## [1] "C:/Users/HP/Documents/R/win-library" "C:/Program Files/R/R-4.0.0/library"</code></pre>
<p>
You can use <code>lib.loc</code> when you want to install, load, update and remove packages from a particular library.
</p>
<section id="install" class="level5">
<h5 class="anchored" data-anchor-id="install">
Install
</h5>
<pre class="r"><code>install.packages('stringr', lib.loc = "C:/Program Files/R/R-3.4.1/library")</code></pre>
</section>
<section id="load" class="level5">
<h5 class="anchored" data-anchor-id="load">
Load
</h5>
<pre class="r"><code>library(lubridate, lib.loc = "C:/Program Files/R/R-3.4.1/library")</code></pre>
</section>
<section id="update-packages" class="level5">
<h5 class="anchored" data-anchor-id="update-packages">
Update Packages
</h5>
<pre class="r"><code>update.packages(lib.loc = "C:/Program Files/R/R-3.4.1/library")</code></pre>
</section>
<section id="remove-packages" class="level5">
<h5 class="anchored" data-anchor-id="remove-packages">
Remove Packages
</h5>
<pre class="r"><code>remove.packages(lib.loc = "C:/Program Files/R/R-3.4.1/library")</code></pre>
</section>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/beginners-guide-to-r-package-ecosystem/</guid>
  <pubDate>Mon, 13 Mar 2017 00:00:00 GMT</pubDate>
  <media:content url="https://blog.rsquaredacademy.com/img/package_banner.png" medium="image" type="image/png" height="89" width="144"/>
</item>
<item>
  <title>Data Types in R</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/data-types-in-r/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-02-17-data-types-in-r.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt how to create variables in R. In this post, we will learn about the following data types:
</p>
<ul>
<li>
numeric/double
</li>
<li>
integer
</li>
<li>
character
</li>
<li>
logical
</li>
<li>
date/time
</li>
</ul>
</section>
<section id="numeric" class="level2">
<h2 class="anchored" data-anchor-id="numeric">
Numeric
</h2>
<p>
In R, numbers are represented by the data type <code>numeric</code>. We will first create a variable and assign it a value. Next we will learn a few methods of checking the type of the variable.
</p>
<pre class="r"><code># create two variables
number1 &lt;- 3.5
number2 &lt;- 3

# check data type
class(number1)</code></pre>
<pre><code>## [1] "numeric"</code></pre>
<pre class="r"><code>class(number2)</code></pre>
<pre><code>## [1] "numeric"</code></pre>
<pre class="r"><code># check if data type is numeric
is.numeric(number1)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>is.numeric(number2)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
If you carefully observe, <code>integers</code> are also treated as <code>numeric/double</code>. We will learn to create integers in a while. In the meanwhile, we have introduced two new funtions in the above example:
</p>
<ul>
<li>
<code>class()</code>: returns the <code>class</code> or <code>type</code>
</li>
<li>
<code>is.numeric()</code>: tests whether the variable is of type <code>numeric</code>
</li>
</ul>
</section>
<section id="integer" class="level2">
<h2 class="anchored" data-anchor-id="integer">
Integer
</h2>
<p>
Unless specified otherwise, integers are treated as <code>numeric</code> or <code>double</code>. In this section, we will learn to create variables of the type <code>integer</code> and to convert other data types to <code>integer</code>.
</p>
<ul>
<li>
create a variable <code>number1</code> and assign it the value <code>3</code>
</li>
<li>
check the data type of <code>number1</code> using <code>class</code>
</li>
<li>
create a second variable <code>number2</code> using <code>as.integer</code> and assign it the value <code>3</code>
</li>
<li>
check the data type of <code>number2</code> using <code>class</code>
</li>
<li>
finally use <code>is.integer</code> to check the data type of both <code>number1</code> and <code>number2</code>
</li>
</ul>
<pre class="r"><code># create a variable and assign it an integer value
number1 &lt;- 3

# create another variable using as.integer
number2 &lt;- as.integer(3)

# check the data type
class(number1)</code></pre>
<pre><code>## [1] "numeric"</code></pre>
<pre class="r"><code>class(number2)</code></pre>
<pre><code>## [1] "integer"</code></pre>
<pre class="r"><code># use is.integer to check data type
is.integer(number1)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>is.integer(number2)</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="character" class="level2">
<h2 class="anchored" data-anchor-id="character">
Character
</h2>
<p>
Letters, words and group of words are represented by the data type <code>character</code>. All data of type <code>character</code> must be enclosed in single or double quotation marks. In fact any value enclosed in quotes will be treated as <code>character</code>. Let us create two variables to store the first and last name of a some random guy.
</p>
<pre class="r"><code># first name
first_name &lt;- "jovial"

# last name
last_name &lt;- 'mann'

# check data type
class(first_name)</code></pre>
<pre><code>## [1] "character"</code></pre>
<pre class="r"><code>class(last_name)</code></pre>
<pre><code>## [1] "character"</code></pre>
<pre class="r"><code># use is.charactert to check data type
is.character(first_name)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>is.character(last_name)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
You can coerce any data type to <code>character</code> using <code>as.character()</code>.
</p>
<pre class="r"><code># create variable of different data types
age &lt;- as.integer(30) # integer
score &lt;- 9.8          # numeric/double
opt_course &lt;- TRUE    # logical
today &lt;- Sys.time()   # date time

as.character(age) </code></pre>
<pre><code>## [1] "30"</code></pre>
<pre class="r"><code>as.character(score)</code></pre>
<pre><code>## [1] "9.8"</code></pre>
<pre class="r"><code>as.character(opt_course)</code></pre>
<pre><code>## [1] "TRUE"</code></pre>
<pre class="r"><code>as.character(today)</code></pre>
<pre><code>## [1] "2020-06-10 18:20:48"</code></pre>
</section>
<section id="logical" class="level2">
<h2 class="anchored" data-anchor-id="logical">
Logical
</h2>
<p>
Logical data types take only 2 values. Either <code>TRUE</code> or <code>FALSE</code>. Sich data types are created when we compare two objects in R using comparison or logical operators.
</p>
<ul>
<li>
create two variables <code>x</code> and <code>y</code>
</li>
<li>
assign them the values <code>TRUE</code> and <code>FALSE</code> respectively
</li>
<li>
use <code>is.logical</code> to check data type
</li>
<li>
use <code>as.logical</code> to coerce other data types to <code>logical</code>
</li>
</ul>
<pre class="r"><code># create variables x and y
x &lt;- TRUE
y &lt;- FALSE

# check data type
class(x)</code></pre>
<pre><code>## [1] "logical"</code></pre>
<pre class="r"><code>is.logical(y)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
The outcome of comparison operators is always <code>logical</code>. In the below example, we compare two numbers to see the outcome.
</p>
<pre class="r"><code># create two numeric variables
x &lt;- 3
y &lt;- 4

# compare x and y
x &gt; y</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code>x &lt; y</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code># store the result
z &lt;- x &gt; y
class(z)</code></pre>
<pre><code>## [1] "logical"</code></pre>
<p>
<code>TRUE</code> is represented by all numbers except <code>0</code>. <code>FALSE</code> is represented only by <code>0</code> and no other numbers.
</p>
<pre class="r"><code># TRUE and FALSE are represented by 1 and 0
as.logical(1)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>as.logical(0)</code></pre>
<pre><code>## [1] FALSE</code></pre>
<pre class="r"><code># using numbers
as.numeric(TRUE)</code></pre>
<pre><code>## [1] 1</code></pre>
<pre class="r"><code>as.numeric(FALSE)</code></pre>
<pre><code>## [1] 0</code></pre>
<pre class="r"><code># using different numbers
as.logical(-2, -1.5, -1, 0, 1, 2)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<p>
Use <code>as.logical()</code> to coerce other data types to <code>logical</code>.
</p>
<pre class="r"><code># create variable of different data types
age &lt;- as.integer(30) # integer
score &lt;- 9.8          # numeric/double
opt_course &lt;- TRUE    # logical
today &lt;- Sys.time()   # date time

as.logical(age) </code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>as.logical(score)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>as.logical(opt_course)</code></pre>
<pre><code>## [1] TRUE</code></pre>
<pre class="r"><code>as.logical(today)</code></pre>
<pre><code>## [1] TRUE</code></pre>
</section>
<section id="summary" class="level2">
<h2 class="anchored" data-anchor-id="summary">
Summary
</h2>
<ul>
<li>
numeric, integer, character, logical and date are the basic data types in R
</li>
<li>
<code>class()</code> or <code>typeof()</code> return the data type
</li>
<li>
<code>is.data_type</code> checks whether the data is of the specified data type
<ul>
<li>
<code>is.numeric()</code>
</li>
<li>
<code>is.integer()</code>
</li>
<li>
<code>is.character()</code>
</li>
<li>
<code>is.logical()</code>
</li>
<li>
<code>is.date()</code>
</li>
</ul>
</li>
<li>
<code>as.data_type</code> will coerce objects to the specified data type
<ul>
<li>
<code>as.numeric()</code>
</li>
<li>
<code>as.integer()</code>
</li>
<li>
<code>as.character()</code>
</li>
<li>
<code>as.logical()</code>
</li>
<li>
<code>as.date()</code>
</li>
</ul>
</li>
</ul>
</section>



 ]]></description>
  <category>r-introduction</category>
  <guid>https://blog.rsquaredacademy.com/posts/data-types-in-r/</guid>
  <pubDate>Fri, 17 Feb 2017 00:00:00 GMT</pubDate>
</item>
<item>
  <title>Variables in R</title>
  <dc:creator>Aravind Hebbali</dc:creator>
  <link>https://blog.rsquaredacademy.com/posts/variables-in-r/</link>
  <description><![CDATA[ 




<!-- Migrated from content/post/2017-02-05-variables.Rmd. -->
<!-- Day-1 static bundle: body reuses the pre-rendered .html fragment. -->
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">
Introduction
</h2>
<p>
In the previous post, we learnt to install RStudio. In this post, we will learn about variables and data types. You can skip this post, if you have prior experience in any other programming language.
</p>
</section>
<section id="what-is-a-variable" class="level2">
<h2 class="anchored" data-anchor-id="what-is-a-variable">
What is a variable?
</h2>
<ul>
<li>
variables are the fundamental elements of any programming language
</li>
<li>
they are used to represent values that are likely to change
</li>
<li>
they reference memory locations that store information/data
</li>
</ul>
<p>
Let us use a simple case study to understand variables. Suppose you are computing the area of a circle whose radius is 3. In R, you can do this straight away as shown below:
</p>
<pre class="r"><code>3.14 * 3 * 3</code></pre>
<pre><code>## [1] 28.26</code></pre>
<p>
But you cannot reuse the radius or the area computed in any other analysis or computation. Let us see how variables can change the above scenario and help us in reusing values and computations.
</p>
</section>
<section id="creating-variables" class="level2">
<h2 class="anchored" data-anchor-id="creating-variables">
Creating Variables
</h2>
<p>
A variable consists of 3 components:
</p>
<ul>
<li>
variable name
</li>
<li>
assignment operator
</li>
<li>
variable value
</li>
</ul>
<p>
We can store the value of the radius by creating a variable and assigning it the value. In this case, we create a variable called <code>radius</code> and assign it the value <code>3</code> using the assignment operator <code>&lt;-</code>.
</p>
<pre class="r"><code>radius &lt;- 3
radius</code></pre>
<pre><code>## [1] 3</code></pre>
<p>
Now that we have learnt to create variables, let us see how we can use them for other computations. For our case study, we will use the <code>radius</code> variable to compute the area of a circle.
</p>
</section>
<section id="using-variables" class="level2">
<h2 class="anchored" data-anchor-id="using-variables">
Using Variables
</h2>
<p>
We will create two variables, <code>radius</code> and <code>pi</code>, and use them to compute the area of a circle and store it in another variable <code>area</code>.
</p>
<pre class="r"><code># assign value 3 to variable radius
radius &lt;- 3

# assign value 3.14 to variable pi
pi &lt;- 3.14

# compute area of circle
area &lt;- pi * radius * radius

# call radius
radius</code></pre>
<pre><code>## [1] 3</code></pre>
<pre class="r"><code># call area 
area</code></pre>
<pre><code>## [1] 28.26</code></pre>
</section>
<section id="components-of-a-variable" class="level2">
<h2 class="anchored" data-anchor-id="components-of-a-variable">
Components of a Variable
</h2>
<center>
</center>
</section>
<section id="naming-conventions" class="level2">
<h2 class="anchored" data-anchor-id="naming-conventions">
Naming Conventions
</h2>
<ul>
<li>
Name must begin with a letter. Do not use numbers, dollar sign (<code>$</code>) or underscore (<code>_</code>).
</li>
<li>
The name can contain numbers or underscore. Do not use dash (<code>-</code>) or period (<code>.</code>).
</li>
<li>
Do not use the names of keywords and avoid using the names of built in functions.
</li>
<li>
Variables are case sensitive; <code>average</code> and <code>Average</code> would be different variables.
</li>
<li>
Use names that are descriptive. Generally, variable names should be nouns.
</li>
<li>
If the name is made of more than one word, use underscore to separate the words.
</li>
</ul>
</section>
<section id="recap.." class="level2">
<h2 class="anchored" data-anchor-id="recap..">
Recap..
</h2>
<ul>
<li>
Variables are the building blocks of a programming language
</li>
<li>
Variables reference memory locations that store information/data
</li>
<li>
An assignment operator <code>&lt;-</code> assigns values to a variable
</li>
<li>
A variable has the following components:
<ul>
<li>
name
</li>
<li>
memory address
</li>
<li>
value
</li>
<li>
data type
</li>
</ul>
</li>
<li>
Certain rules must be followed while naming variables
</li>
</ul>
</section>



 ]]></description>
  <category>R Programming</category>
  <guid>https://blog.rsquaredacademy.com/posts/variables-in-r/</guid>
  <pubDate>Sun, 05 Feb 2017 00:00:00 GMT</pubDate>
</item>
</channel>
</rss>
