number_of_cameras <- 12
species <- "Raccoon"
camera_active <- TRUEModule 3: Data Structures in R
Introduction
In the previous module, we learned that R stores information in objects.
So far, most of our objects contained only one value:
But real datasets contain much more information than a single number or word.
R allows us to organize multiple values into different data structures.
In this module, we will focus on four of the most common:
- vectors,
- matrices,
- lists,
- data frames.
Understanding the difference between these structures will make it much easier to understand what R is doing later when we start working with real datasets.
Let’s get started
Open your intro-r-course project and create a new R script.
Save it inside the scripts folder as:
module-03.R
We will use this script throughout the module.
Vectors
A vector is one of the most basic and important data structures in R.
A vector allows us to store several values inside a single object.
We can create a vector using the function:
c()
The c stands for combine.
For example:
detections <- c(5, 12, 8, 0, 3)
detections[1] 5 12 8 0 3
Instead of creating five different objects:
detection_1 <- 5
detection_2 <- 12
detection_3 <- 8
detection_4 <- 0
detection_5 <- 3we can store all five values together:
detections <- c(5, 12, 8, 0, 3)This makes it much easier to work with several related observations.
Numeric vectors
Vectors can contain numeric values:
camera_days <- c(30, 28, 30, 25, 29)
camera_days[1] 30 28 30 25 29
We can perform operations directly on the entire vector:
camera_days + 1[1] 31 29 31 26 30
or:
camera_days * 2[1] 60 56 60 50 58
R applies the operation to each element of the vector.
We can also use functions:
mean(camera_days)[1] 28.4
sum(camera_days)[1] 142
length(camera_days)[1] 5
The function length() tells us how many elements are contained in the vector.
Character vectors
Vectors can also contain text:
species <- c(
"Raccoon",
"Coyote",
"Bobcat",
"White-tailed deer"
)
species[1] "Raccoon" "Coyote" "Bobcat"
[4] "White-tailed deer"
And logical values:
camera_active <- c(TRUE, TRUE, FALSE, TRUE)
camera_active[1] TRUE TRUE FALSE TRUE
An atomic vector can contain only one basic data type.
For example, all its values can be numeric, all can be character, or all can be logical.
What happens if we mix types?
mixed_vector <- c(1, 2, "Raccoon", 4)
mixed_vector[1] "1" "2" "Raccoon" "4"
R converts all the values to a compatible type.
Check:
class(mixed_vector)[1] "character"
Because the vector contains text, R converted the numbers into character values too.
This automatic conversion is called coercion.
If an object that you expected to be numeric suddenly becomes character, check whether one of its values contains text.
Creating sequences
Sometimes we do not want to type every number individually.
For example:
1:10 [1] 1 2 3 4 5 6 7 8 9 10
We can save it as an object:
sampling_days <- 1:10
sampling_days [1] 1 2 3 4 5 6 7 8 9 10
The function seq() gives us more control. This function has several arguments: from is the starting number, to is the ending number, length.out is used to control the length of the vector, and by allows us to specify the interval of the sequence.
seq(from = 0, to = 30, by = 5)[1] 0 5 10 15 20 25 30
We can also repeat values using rep(). The times argument allows us to specify the number of times a value or vector will be repeated. When we specify each, we can control the number of times each value within the vector is repeated.
rep("Camera", times = 5)[1] "Camera" "Camera" "Camera" "Camera" "Camera"
These functions are useful when we need to generate repeated or sequential values without typing everything manually.
Selecting elements from a vector
Sometimes we only want one or a few values from a vector.
We can select elements using square brackets:
[]
Let’s create a vector:
detections <- c(5, 12, 8, 0, 3)To select the first value:
detections[1][1] 5
To select the third:
detections[3][1] 8
We can select several positions:
detections[c(1, 3, 5)][1] 5 8 3
Or a sequence of positions:
detections[1:3][1] 5 12 8
R starts counting at 1, not at 0.
The first element of a vector is:
x[1]We can also exclude positions using negative numbers:
detections[-1][1] 12 8 0 3
This returns everything except the first element.
Selecting values using a condition
Remember the relational operators from the previous module?
We can use them with vectors.
For example:
detections > 5[1] FALSE TRUE TRUE FALSE FALSE
R returns a logical value for every element:
We can use this condition inside the brackets:
detections[detections > 5][1] 12 8
R returns only the values that meet the condition:
This idea will become extremely important later when we start filtering datasets.
Create this vector:
camera_days <- c(30, 21, 29, 15, 30, 27)Then:
- Select the second value.
- Select the first three values.
- Select values 2, 4, and 6.
- Find which values are greater than 25.
- Return only the values greater than 25.
Matrices
A matrix is a two-dimensional data structure.
Unlike a vector, which has only one dimension, a matrix has:
- rows,
- columns.
We can create a matrix using the function matrix().
For example:
camera_matrix <- matrix(
1:12,
nrow = 4,
ncol = 3
)
camera_matrix [,1] [,2] [,3]
[1,] 1 5 9
[2,] 2 6 10
[3,] 3 7 11
[4,] 4 8 12
R fills matrices by column by default.
We can change this using:
camera_matrix <- matrix(
1:12,
nrow = 4,
ncol = 3,
byrow = TRUE
)
camera_matrix [,1] [,2] [,3]
[1,] 1 2 3
[2,] 4 5 6
[3,] 7 8 9
[4,] 10 11 12
Now the values are filled by row.
Matrix dimensions
We can inspect the dimensions of a matrix using:
dim(camera_matrix)[1] 4 3
This returns the number of rows and columns.
We can also use:
nrow(camera_matrix)[1] 4
and:
ncol(camera_matrix)[1] 3
Creating matrices from vectors
We can combine vectors into a matrix.
For example:
station_1 <- c(5, 2, 0)
station_2 <- c(3, 1, 4)
station_3 <- c(8, 0, 2)We can combine them by rows:
detection_matrix <- rbind(
station_1,
station_2,
station_3
)
detection_matrix [,1] [,2] [,3]
station_1 5 2 0
station_2 3 1 4
station_3 8 0 2
Or by columns:
cbind(
station_1,
station_2,
station_3
) station_1 station_2 station_3
[1,] 5 3 8
[2,] 2 1 0
[3,] 0 4 2
The functions are easy to remember:
rbind = row bind
cbind = column bind
Naming rows and columns
Matrices can have row and column names.
For example:
rownames(detection_matrix) <- c(
"CT01",
"CT02",
"CT03"
)
colnames(detection_matrix) <- c(
"Raccoon",
"Coyote",
"Bobcat"
)
detection_matrix Raccoon Coyote Bobcat
CT01 5 2 0
CT02 3 1 4
CT03 8 0 2
Now our matrix is much easier to interpret.
This type of structure will become familiar later when we work with camera-trap detection histories.
Selecting elements from a matrix
Because matrices have two dimensions, we need to specify:
[row, column]
For example:
detection_matrix[1, 2][1] 2
means:
row 1, column 2
We can select an entire row:
detection_matrix[1, ]Raccoon Coyote Bobcat
5 2 0
Or an entire column:
detection_matrix[, 2]CT01 CT02 CT03
2 1 0
We can also use names:
detection_matrix["CT01", "Coyote"][1] 2
or:
detection_matrix[, "Raccoon"]CT01 CT02 CT03
5 3 8
A useful way to remember matrix indexing is:
object[row, column]
The comma separates the two dimensions.
We can also select values from the diagonal using diag
diag(detection_matrix)[1] 5 1 2
Delete row 1 and select column 2
detection_matrix[-1,2]CT02 CT03
1 0
Delete rows 1 and 4, and select columns 1 and 3
detection_matrix[c(-1,-4), c(1,3)] Raccoon Bobcat
CT02 3 4
CT03 8 2
When we use the selectors on the left side of the object definition, we can replace the values with the ones we are defining.
# We can replace values in matrices
# Replace the second row with 100, 200, and 300
detection_matrix[2,] <- c(100,200, 300)
detection_matrix Raccoon Coyote Bobcat
CT01 5 2 0
CT02 100 200 300
CT03 8 0 2
Matrices contain one type of data
Like vectors, matrices contain only one basic type of data.
For example:
numeric_matrix <- matrix(1:9, nrow = 3)
numeric_matrix [,1] [,2] [,3]
[1,] 1 4 7
[2,] 2 5 8
[3,] 3 6 9
But if we introduce text:
mixed_matrix <- cbind(
station = c("CT01", "CT02", "CT03"),
detections = c(5, 8, 2)
)
mixed_matrix station detections
[1,] "CT01" "5"
[2,] "CT02" "8"
[3,] "CT03" "2"
R converts the numeric values into characters because the entire matrix must have a common type.
This limitation is one of the reasons why data frames are more useful for most datasets.
Lists
A list is a very flexible R data structure.
Unlike vectors and matrices, lists can contain objects of different types and structures.
For example:
camera_information <- list(
station = "CT01",
active = TRUE,
detections = 25,
species = c("Raccoon", "Coyote", "Bobcat")
)
camera_information$station
[1] "CT01"
$active
[1] TRUE
$detections
[1] 25
$species
[1] "Raccoon" "Coyote" "Bobcat"
This single list contains:
- a character value,
- a logical value,
- a numeric value,
- a character vector.
A list can even contain matrices, data frames, or other lists.
For example:
example_list <- list(
numbers = c(1, 2, 3),
matrix = matrix(1:4, nrow = 2),
message = "Hello R"
)
example_list$numbers
[1] 1 2 3
$matrix
[,1] [,2]
[1,] 1 3
[2,] 2 4
$message
[1] "Hello R"
This flexibility makes lists extremely common in R.
Many functions and statistical models return their results as lists containing several different objects.
Selecting elements from a list
If the elements have names, we can use $:
camera_information$station[1] "CT01"
camera_information$species[1] "Raccoon" "Coyote" "Bobcat"
We can also use double brackets:
camera_information[[1]][1] "CT01"
or:
camera_information[["station"]][1] "CT01"
You may also see:
camera_information[1]$station
[1] "CT01"
but there is an important difference.
returns a list containing the first element.
In contrast:
camera_information[[1]][1] "CT01"
returns the contents of the first element.
This distinction can be confusing at first, and you do not need to memorize it yet.
For now, the most important idea is:
Lists can store many different kinds of objects together.
Data frames
For the type of work we will do in this course, data frames will probably be the most important data structure.
A data frame is a rectangular table with:
- rows,
- columns.
It looks very similar to an Excel spreadsheet.
However, unlike a matrix, different columns of a data frame can contain different types of data.
For example, let’s create some information from a hypothetical camera-trap survey:
station <- c(
"CT01",
"CT02",
"CT03",
"CT04"
)
habitat <- c(
"Forest",
"Forest",
"Pasture",
"Pasture"
)
camera_days <- c(
30,
28,
30,
27
)
active <- c(
TRUE,
TRUE,
FALSE,
TRUE
)Each object is a vector.
Now we can combine them into a data frame:
camera_data <- data.frame(
station,
habitat,
camera_days,
active
)
camera_data station habitat camera_days active
1 CT01 Forest 30 TRUE
2 CT02 Forest 28 TRUE
3 CT03 Pasture 30 FALSE
4 CT04 Pasture 27 TRUE
We now have one object containing several columns.
Each column can have a different type:
class(camera_data$station)[1] "character"
class(camera_data$camera_days)[1] "numeric"
class(camera_data$active)[1] "logical"
The data frame itself has its own class:
class(camera_data)[1] "data.frame"
R returns:
In most ecological datasets:
- rows represent observations,
- columns represent variables.
Understanding this organization is extremely important when preparing data for analysis.
In our example:
station habitat camera_days active
are variables.
Each row represents one camera station.
Inspecting a data frame
Before analyzing a dataset, we should always look at its structure.
Some useful functions are:
camera_data station habitat camera_days active
1 CT01 Forest 30 TRUE
2 CT02 Forest 28 TRUE
3 CT03 Pasture 30 FALSE
4 CT04 Pasture 27 TRUE
head()
Shows the first few rows:
head(camera_data) station habitat camera_days active
1 CT01 Forest 30 TRUE
2 CT02 Forest 28 TRUE
3 CT03 Pasture 30 FALSE
4 CT04 Pasture 27 TRUE
str()
Shows the internal structure of the object:
str(camera_data)'data.frame': 4 obs. of 4 variables:
$ station : chr "CT01" "CT02" "CT03" "CT04"
$ habitat : chr "Forest" "Forest" "Pasture" "Pasture"
$ camera_days: num 30 28 30 27
$ active : logi TRUE TRUE FALSE TRUE
dim()
Returns the number of rows and columns:
dim(camera_data)[1] 4 4
nrow()
Returns the number of rows:
nrow(camera_data)[1] 4
ncol()
Returns the number of columns:
ncol(camera_data)[1] 4
names()
Returns the column names:
names(camera_data)[1] "station" "habitat" "camera_days" "active"
When you load a new dataset into R, one of the first things you should do is inspect it.
Functions such as:
head()
str()
dim()
names()can help you identify problems before you begin an analysis.
Selecting columns from a data frame
We can access a column using $:
camera_data$station[1] "CT01" "CT02" "CT03" "CT04"
camera_data$camera_days[1] 30 28 30 27
Notice what happens when we do this:
class(camera_data$camera_days)[1] "numeric"
The column itself is a vector.
This is an important idea:
A data frame is essentially a collection of equal-length vectors organized as columns.
We can also use square brackets.
For example:
camera_data[1, 2][1] "Forest"
means:
row 1, column 2
To select the first row:
camera_data[1, ] station habitat camera_days active
1 CT01 Forest 30 TRUE
To select the second column:
camera_data[, 2][1] "Forest" "Forest" "Pasture" "Pasture"
We can also use column names:
camera_data[, "habitat"][1] "Forest" "Forest" "Pasture" "Pasture"
Or select several columns:
camera_data[, c("station", "camera_days")] station camera_days
1 CT01 30
2 CT02 28
3 CT03 30
4 CT04 27
Later we will learn easier and more readable ways to select rows and columns using dplyr.
For now, it is important to understand how data frames work in base R.
Adding a column
We can add a new column using $.
For example:
camera_data$elevation <- c(
105,
112,
98,
101
)
camera_data station habitat camera_days active elevation
1 CT01 Forest 30 TRUE 105
2 CT02 Forest 28 TRUE 112
3 CT03 Pasture 30 FALSE 98
4 CT04 Pasture 27 TRUE 101
The new column must have the correct number of values.
Our data frame has four rows, so the new column needs four values.
We can verify the number of rows with:
nrow(camera_data)[1] 4
Changing values
We can also replace specific values.
For example:
camera_data$camera_days[2] <- 30
camera_data station habitat camera_days active elevation
1 CT01 Forest 30 TRUE 105
2 CT02 Forest 30 TRUE 112
3 CT03 Pasture 30 FALSE 98
4 CT04 Pasture 27 TRUE 101
Or using row and column positions:
camera_data[2, "camera_days"] <- 30This can be useful, but we should be careful when modifying raw data manually.
Whenever possible, avoid manually changing your original data file.
It is usually better to make corrections through code so that the changes are documented and reproducible.
Different structures, different purposes
Let’s summarize the four structures we have seen.
| Structure | Dimensions | Same data type? | Common use |
|---|---|---|---|
| Vector | 1 | Yes | Store a sequence of values |
| Matrix | 2 | Yes | Numeric or structured rectangular data |
| List | 1 | No | Store different objects together |
| Data frame | 2 | Different types allowed by column | Store datasets |
The structure we choose depends on the type of information we want to store.
Throughout the rest of the course:
- we will constantly work with vectors,
- most of our datasets will be data frames,
- we will occasionally encounter matrices,
- and many R functions will return results as lists.
When we eventually work with camtrapR, all four of these ideas will appear again.
Converting between data structures
R also provides functions for converting objects from one structure to another.
For example:
as.data.frame(detection_matrix) Raccoon Coyote Bobcat
CT01 5 2 0
CT02 100 200 300
CT03 8 0 2
converts a matrix into a data frame.
We can save the result:
detection_data <- as.data.frame(detection_matrix)
class(detection_data)[1] "data.frame"
Similarly:
as.matrix(camera_data) station habitat camera_days active elevation
[1,] "CT01" "Forest" "30" "TRUE" "105"
[2,] "CT02" "Forest" "30" "TRUE" "112"
[3,] "CT03" "Pasture" "30" "FALSE" " 98"
[4,] "CT04" "Pasture" "27" "TRUE" "101"
attempts to convert a data frame into a matrix.
However, remember that a matrix can contain only one basic data type.
Because camera_data contains text, numbers, and logical values, R will need to convert everything to a common type.
Try:
camera_matrix_2 <- as.matrix(camera_data)
camera_matrix_2 station habitat camera_days active elevation
[1,] "CT01" "Forest" "30" "TRUE" "105"
[2,] "CT02" "Forest" "30" "TRUE" "112"
[3,] "CT03" "Pasture" "30" "FALSE" " 98"
[4,] "CT04" "Pasture" "27" "TRUE" "101"
and then:
class(camera_matrix_2)[1] "matrix" "array"
This is another good example of why understanding data structures matters.
A function may technically convert an object, but the result may not have the structure or data type that you expected.
Practice at home
This activity should take approximately 15–20 minutes.
Open your intro-r-course project and create a new script called:
practice-03.R
Save it inside the scripts folder.
Part 1: Vectors
Create these objects:
stations <- c("CT01", "CT02", "CT03", "CT04", "CT05")
detections <- c(12, 5, 0, 18, 7)
camera_days <- c(30, 28, 30, 25, 29)Then:
- Return the third element of
stations. - Return the first three values of
detections. - Determine which values of
detectionsare greater than 10. - Return only the detections greater than 10.
- Calculate the mean number of camera-days.
Part 2: Data frame
Create a data frame called survey_data using:
stations[1] "CT01" "CT02" "CT03" "CT04" "CT05"
detections[1] 12 5 0 18 7
camera_days[1] 30 28 30 25 29
Then:
- Display the data frame.
- Use
str()to inspect its structure. - Use
dim()to find its dimensions. - Use
names()to obtain its column names. - Select only the
detectionscolumn. - Select the information from the second camera station.
- Add a new column called
habitatcontaining:
c("Forest", "Forest", "Pasture", "Forest", "Pasture")[1] "Forest" "Forest" "Pasture" "Forest" "Pasture"
- Use
str()again and look at the type of each column.
# Practice for Module 3
# ---------------------------
# Part 1: Vectors
# ---------------------------
stations <- c(
"CT01",
"CT02",
"CT03",
"CT04",
"CT05"
)
detections <- c(
12,
5,
0,
18,
7
)
camera_days <- c(
30,
28,
30,
25,
29
)
# Third station
stations[3][1] "CT03"
# First three detection values
detections[1:3][1] 12 5 0
# Which values are greater than 10?
detections > 10[1] TRUE FALSE FALSE TRUE FALSE
# Return only detections greater than 10
detections[detections > 10][1] 12 18
# Mean camera-days
mean(camera_days)[1] 28.4
# ---------------------------
# Part 2: Data frame
# ---------------------------
survey_data <- data.frame(
stations,
detections,
camera_days
)
# Display the data
survey_data stations detections camera_days
1 CT01 12 30
2 CT02 5 28
3 CT03 0 30
4 CT04 18 25
5 CT05 7 29
# Inspect structure
str(survey_data)'data.frame': 5 obs. of 3 variables:
$ stations : chr "CT01" "CT02" "CT03" "CT04" ...
$ detections : num 12 5 0 18 7
$ camera_days: num 30 28 30 25 29
# Dimensions
dim(survey_data)[1] 5 3
# Column names
names(survey_data)[1] "stations" "detections" "camera_days"
# Select detections
survey_data$detections[1] 12 5 0 18 7
# Select the second camera station
survey_data[2, ] stations detections camera_days
2 CT02 5 28
# Add habitat
survey_data$habitat <- c(
"Forest",
"Forest",
"Pasture",
"Forest",
"Pasture"
)
# Inspect the updated data frame
str(survey_data)'data.frame': 5 obs. of 4 variables:
$ stations : chr "CT01" "CT02" "CT03" "CT04" ...
$ detections : num 12 5 0 18 7
$ camera_days: num 30 28 30 25 29
$ habitat : chr "Forest" "Forest" "Pasture" "Forest" ...
The important thing is that you understand the difference between:
- a vector,
- a matrix,
- a list,
- and a data frame,
and that you can recognize how rows, columns, and positions are used to access information.