Main idea
R work often follows this pattern:
create objects
↓
organize values
↓
import data
↓
inspect data
↓
work with data
↓
export results
The details take practice.
But this workflow will appear again and again.
Session 2
Today we will connect three important ideas:
objects → data structures → external data files
By the end, we should be able to create simple R objects, understand the basic structure of datasets, and import data from CSV and Excel files.
Most of the time, we give R a value and store it in an object.
The object name is on the left.
The value is on the right.
object_name <- value
Without an object, R gives us a result but does not save it.
But if we assign the result:
we can use it again later.
Objects allow us to build an analysis step by step.
Good object names make code easier to understand.
works, but this is more informative:
During this course, we will mostly use snake_case.
Object names:
R is case-sensitive:
are different names.
R can store different types of values.
| Type | Example |
|---|---|
| numeric | 25 |
| character | "Raccoon" |
| logical | TRUE |
| missing value | NA |
A number inside quotation marks is not a number.
Arithmetic operators return numbers.
Relational operators return TRUE or FALSE.
Functions perform specific tasks.
Most functions look like this:
function_name(object)
For example:
Functions can also use arguments:
You do not need to memorize every function.
Use the help system:
or:
Learning R is not memorizing commands.
It is learning how to find what you need.
Many functions come from packages.
The basic workflow is:
install → load → use
For example:
Switch to RStudio.
Create a script called:
session-02.R
inside your scripts/ folder.
Then create:
Use R to ask:
So far, most objects had one value.
Real datasets have many values.
R can organize values into different structures:
For this course, data frames will be the most important.
A vector stores several values in one object.
The function c() means combine.
We can use functions on vectors:
A vector can contain numbers:
or text:
or logical values:
But mixing types can change the result.
Use square brackets:
R starts counting at 1.
Relational operators work with vectors.
This returns one logical value for each element.
We can use that condition to select values:
This is the basic idea behind filtering data.
A matrix has two dimensions:
rows × columns
Example:
To select values:
Camera-trap detection histories often look like matrices.
occasion_1 occasion_2 occasion_3
CT01 0 1 0
CT02 1 0 0
CT03 0 0 1
Rows are sites or stations.
Columns are sampling occasions.
Lists can store different kinds of objects together.
Access elements with:
Many R functions return lists.
A data frame is a rectangular table.
It has:
It looks similar to an Excel spreadsheet.
But unlike a matrix, each column can have a different type.
In most ecological datasets:
rows = observations
columns = variables
For example:
| station | habitat | camera_days | active |
|---|---|---|---|
| CT01 | Forest | 30 | TRUE |
| CT02 | Forest | 28 | TRUE |
| CT03 | Pasture | 30 | FALSE |
Each row is one camera station.
Each column is one variable.
Before analyzing data, inspect it.
Ask:
Use $ to select a column:
Use [row, column] to select by position:
Later we will learn more readable tools with dplyr.
Create this data frame:
station <- c("CT01", "CT02", "CT03")
habitat <- c("Forest", "Pasture", "Forest")
detections <- c(12, 5, 0)
survey_data <- data.frame(
station,
habitat,
detections
)Then run:
Most real data will not be typed directly into R.
They will usually come from:
So R needs to know where the file is.
The working directory is the folder where R starts looking for files.
Check it with:
Because we are using an RStudio Project, the working directory should be the main project folder.
intro-r-course/
If your file is here:
intro-r-course/
│
├── data/
│ └── camera_data.csv
│
├── scripts/
└── outputs/
you can refer to it as:
data/camera_data.csv
This is better than using a full path from your own computer.
CSV means:
Comma-Separated Values
A CSV is a plain-text table.
station,species,detections,camera_days
CT01,Raccoon,12,30
CT02,Coyote,5,28
CT03,Bobcat,8,30
CSV files are simple, portable, and easy to read in R.
Base R:
Tidyverse / readr:
The important pattern is:
object <- read_function("path/to/file")
To read Excel files, use readxl.
If the file has multiple sheets:
After importing:
Do not assume the file imported correctly just because R did not show an error.
Trust, but verify.
Most import errors are simple:
Useful checks:
Download the example files from Module 4.
Place them in:
intro-r-course/data/
Then import the CSV:
Inspect it:
Now import the Excel file.
Check the sheets:
Inspect the object:
To save a CSV from R:
With readr:
A useful structure:
intro-r-course/
│
├── data/
│ └── camera_data.csv
│
├── scripts/
│ └── session-02.R
│
└── outputs/
└── camera_data_clean.csv
data/ contains original or input files.
outputs/ contains files created by your analysis.
Before the next session:
intro-r-course project.practice-02.R and practice-03.R if you have not done them yet.outputs/ folder.R work often follows this pattern:
create objects
↓
organize values
↓
import data
↓
inspect data
↓
work with data
↓
export results
The details take practice.
But this workflow will appear again and again.