Module 3: Data Structures in R

Introduction

In the previous module, we learned that R stores information in objects.

So far, most of our objects contained only one value:

number_of_cameras <- 12

species <- "Raccoon"

camera_active <- TRUE

But real datasets contain much more information than a single number or word.

R allows us to organize multiple values into different data structures.

In this module, we will focus on four of the most common:

  • vectors,
  • matrices,
  • lists,
  • data frames.

Understanding the difference between these structures will make it much easier to understand what R is doing later when we start working with real datasets.

Let’s get started

Open your intro-r-course project and create a new R script.

Save it inside the scripts folder as:

module-03.R

We will use this script throughout the module.

Vectors

A vector is one of the most basic and important data structures in R.

A vector allows us to store several values inside a single object.

We can create a vector using the function:

c()

The c stands for combine.

For example:

detections <- c(5, 12, 8, 0, 3)

detections
[1]  5 12  8  0  3

Instead of creating five different objects:

detection_1 <- 5
detection_2 <- 12
detection_3 <- 8
detection_4 <- 0
detection_5 <- 3

we can store all five values together:

detections <- c(5, 12, 8, 0, 3)

This makes it much easier to work with several related observations.

Numeric vectors

Vectors can contain numeric values:

camera_days <- c(30, 28, 30, 25, 29)

camera_days
[1] 30 28 30 25 29

We can perform operations directly on the entire vector:

camera_days + 1
[1] 31 29 31 26 30

or:

camera_days * 2
[1] 60 56 60 50 58

R applies the operation to each element of the vector.

We can also use functions:

mean(camera_days)
[1] 28.4
sum(camera_days)
[1] 142
length(camera_days)
[1] 5

The function length() tells us how many elements are contained in the vector.

Character vectors

Vectors can also contain text:

species <- c(
  "Raccoon",
  "Coyote",
  "Bobcat",
  "White-tailed deer"
)

species
[1] "Raccoon"           "Coyote"            "Bobcat"           
[4] "White-tailed deer"

And logical values:

camera_active <- c(TRUE, TRUE, FALSE, TRUE)

camera_active
[1]  TRUE  TRUE FALSE  TRUE
Vectors contain one type of data

An atomic vector can contain only one basic data type.

For example, all its values can be numeric, all can be character, or all can be logical.

What happens if we mix types?

mixed_vector <- c(1, 2, "Raccoon", 4)

mixed_vector
[1] "1"       "2"       "Raccoon" "4"      

R converts all the values to a compatible type.

Check:

class(mixed_vector)
[1] "character"

Because the vector contains text, R converted the numbers into character values too.

This automatic conversion is called coercion.

Tip

If an object that you expected to be numeric suddenly becomes character, check whether one of its values contains text.

Creating sequences

Sometimes we do not want to type every number individually.

For example:

1:10
 [1]  1  2  3  4  5  6  7  8  9 10

We can save it as an object:

sampling_days <- 1:10

sampling_days
 [1]  1  2  3  4  5  6  7  8  9 10

The function seq() gives us more control. This function has several arguments: from is the starting number, to is the ending number, length.out is used to control the length of the vector, and by allows us to specify the interval of the sequence.

seq(from = 0, to = 30, by = 5)
[1]  0  5 10 15 20 25 30

We can also repeat values using rep(). The times argument allows us to specify the number of times a value or vector will be repeated. When we specify each, we can control the number of times each value within the vector is repeated.

rep("Camera", times = 5)
[1] "Camera" "Camera" "Camera" "Camera" "Camera"

These functions are useful when we need to generate repeated or sequential values without typing everything manually.

Selecting elements from a vector

Sometimes we only want one or a few values from a vector.

We can select elements using square brackets:

[]

Let’s create a vector:

detections <- c(5, 12, 8, 0, 3)

To select the first value:

detections[1]
[1] 5

To select the third:

detections[3]
[1] 8

We can select several positions:

detections[c(1, 3, 5)]
[1] 5 8 3

Or a sequence of positions:

detections[1:3]
[1]  5 12  8
Important

R starts counting at 1, not at 0.

The first element of a vector is:

x[1]

We can also exclude positions using negative numbers:

detections[-1]
[1] 12  8  0  3

This returns everything except the first element.

Selecting values using a condition

Remember the relational operators from the previous module?

We can use them with vectors.

For example:

detections > 5
[1] FALSE  TRUE  TRUE FALSE FALSE

R returns a logical value for every element:

We can use this condition inside the brackets:

detections[detections > 5]
[1] 12  8

R returns only the values that meet the condition:

This idea will become extremely important later when we start filtering datasets.

Try it yourself

Create this vector:

camera_days <- c(30, 21, 29, 15, 30, 27)

Then:

  1. Select the second value.
  2. Select the first three values.
  3. Select values 2, 4, and 6.
  4. Find which values are greater than 25.
  5. Return only the values greater than 25.

Matrices

A matrix is a two-dimensional data structure.

Unlike a vector, which has only one dimension, a matrix has:

  • rows,
  • columns.

We can create a matrix using the function matrix().

For example:

camera_matrix <- matrix(
  1:12,
  nrow = 4,
  ncol = 3
)

camera_matrix
     [,1] [,2] [,3]
[1,]    1    5    9
[2,]    2    6   10
[3,]    3    7   11
[4,]    4    8   12

R fills matrices by column by default.

We can change this using:

camera_matrix <- matrix(
  1:12,
  nrow = 4,
  ncol = 3,
  byrow = TRUE
)

camera_matrix
     [,1] [,2] [,3]
[1,]    1    2    3
[2,]    4    5    6
[3,]    7    8    9
[4,]   10   11   12

Now the values are filled by row.

Matrix dimensions

We can inspect the dimensions of a matrix using:

dim(camera_matrix)
[1] 4 3

This returns the number of rows and columns.

We can also use:

nrow(camera_matrix)
[1] 4

and:

ncol(camera_matrix)
[1] 3

Creating matrices from vectors

We can combine vectors into a matrix.

For example:

station_1 <- c(5, 2, 0)
station_2 <- c(3, 1, 4)
station_3 <- c(8, 0, 2)

We can combine them by rows:

detection_matrix <- rbind(
  station_1,
  station_2,
  station_3
)

detection_matrix
          [,1] [,2] [,3]
station_1    5    2    0
station_2    3    1    4
station_3    8    0    2

Or by columns:

cbind(
  station_1,
  station_2,
  station_3
)
     station_1 station_2 station_3
[1,]         5         3         8
[2,]         2         1         0
[3,]         0         4         2

The functions are easy to remember:

rbind = row bind

cbind = column bind

Naming rows and columns

Matrices can have row and column names.

For example:

rownames(detection_matrix) <- c(
  "CT01",
  "CT02",
  "CT03"
)

colnames(detection_matrix) <- c(
  "Raccoon",
  "Coyote",
  "Bobcat"
)

detection_matrix
     Raccoon Coyote Bobcat
CT01       5      2      0
CT02       3      1      4
CT03       8      0      2

Now our matrix is much easier to interpret.

This type of structure will become familiar later when we work with camera-trap detection histories.

Selecting elements from a matrix

Because matrices have two dimensions, we need to specify:

[row, column]

For example:

detection_matrix[1, 2]
[1] 2

means:

row 1, column 2

We can select an entire row:

detection_matrix[1, ]
Raccoon  Coyote  Bobcat 
      5       2       0 

Or an entire column:

detection_matrix[, 2]
CT01 CT02 CT03 
   2    1    0 

We can also use names:

detection_matrix["CT01", "Coyote"]
[1] 2

or:

detection_matrix[, "Raccoon"]
CT01 CT02 CT03 
   5    3    8 
Tip

A useful way to remember matrix indexing is:

object[row, column]

The comma separates the two dimensions.

We can also select values from the diagonal using diag

diag(detection_matrix)
[1] 5 1 2

Delete row 1 and select column 2

detection_matrix[-1,2]
CT02 CT03 
   1    0 

Delete rows 1 and 4, and select columns 1 and 3

detection_matrix[c(-1,-4), c(1,3)]
     Raccoon Bobcat
CT02       3      4
CT03       8      2

When we use the selectors on the left side of the object definition, we can replace the values with the ones we are defining.

# We can replace values in matrices

# Replace the second row with 100, 200, and 300
detection_matrix[2,] <- c(100,200, 300)
detection_matrix
     Raccoon Coyote Bobcat
CT01       5      2      0
CT02     100    200    300
CT03       8      0      2

Matrices contain one type of data

Like vectors, matrices contain only one basic type of data.

For example:

numeric_matrix <- matrix(1:9, nrow = 3)

numeric_matrix
     [,1] [,2] [,3]
[1,]    1    4    7
[2,]    2    5    8
[3,]    3    6    9

But if we introduce text:

mixed_matrix <- cbind(
  station = c("CT01", "CT02", "CT03"),
  detections = c(5, 8, 2)
)

mixed_matrix
     station detections
[1,] "CT01"  "5"       
[2,] "CT02"  "8"       
[3,] "CT03"  "2"       

R converts the numeric values into characters because the entire matrix must have a common type.

This limitation is one of the reasons why data frames are more useful for most datasets.

Lists

A list is a very flexible R data structure.

Unlike vectors and matrices, lists can contain objects of different types and structures.

For example:

camera_information <- list(
  station = "CT01",
  active = TRUE,
  detections = 25,
  species = c("Raccoon", "Coyote", "Bobcat")
)

camera_information
$station
[1] "CT01"

$active
[1] TRUE

$detections
[1] 25

$species
[1] "Raccoon" "Coyote"  "Bobcat" 

This single list contains:

  • a character value,
  • a logical value,
  • a numeric value,
  • a character vector.

A list can even contain matrices, data frames, or other lists.

For example:

example_list <- list(
  numbers = c(1, 2, 3),
  matrix = matrix(1:4, nrow = 2),
  message = "Hello R"
)

example_list
$numbers
[1] 1 2 3

$matrix
     [,1] [,2]
[1,]    1    3
[2,]    2    4

$message
[1] "Hello R"

This flexibility makes lists extremely common in R.

Many functions and statistical models return their results as lists containing several different objects.

Selecting elements from a list

If the elements have names, we can use $:

camera_information$station
[1] "CT01"
camera_information$species
[1] "Raccoon" "Coyote"  "Bobcat" 

We can also use double brackets:

camera_information[[1]]
[1] "CT01"

or:

camera_information[["station"]]
[1] "CT01"

You may also see:

camera_information[1]
$station
[1] "CT01"

but there is an important difference.

returns a list containing the first element.

In contrast:

camera_information[[1]]
[1] "CT01"

returns the contents of the first element.

This distinction can be confusing at first, and you do not need to memorize it yet.

For now, the most important idea is:

Lists can store many different kinds of objects together.

Data frames

For the type of work we will do in this course, data frames will probably be the most important data structure.

A data frame is a rectangular table with:

  • rows,
  • columns.

It looks very similar to an Excel spreadsheet.

However, unlike a matrix, different columns of a data frame can contain different types of data.

For example, let’s create some information from a hypothetical camera-trap survey:

station <- c(
  "CT01",
  "CT02",
  "CT03",
  "CT04"
)

habitat <- c(
  "Forest",
  "Forest",
  "Pasture",
  "Pasture"
)

camera_days <- c(
  30,
  28,
  30,
  27
)

active <- c(
  TRUE,
  TRUE,
  FALSE,
  TRUE
)

Each object is a vector.

Now we can combine them into a data frame:

camera_data <- data.frame(
  station,
  habitat,
  camera_days,
  active
)

camera_data
  station habitat camera_days active
1    CT01  Forest          30   TRUE
2    CT02  Forest          28   TRUE
3    CT03 Pasture          30  FALSE
4    CT04 Pasture          27   TRUE

We now have one object containing several columns.

Each column can have a different type:

class(camera_data$station)
[1] "character"
class(camera_data$camera_days)
[1] "numeric"
class(camera_data$active)
[1] "logical"

The data frame itself has its own class:

class(camera_data)
[1] "data.frame"

R returns:

Rows and columns usually represent different things

In most ecological datasets:

  • rows represent observations,
  • columns represent variables.

Understanding this organization is extremely important when preparing data for analysis.

In our example:

station      habitat      camera_days      active

are variables.

Each row represents one camera station.

Inspecting a data frame

Before analyzing a dataset, we should always look at its structure.

Some useful functions are:

camera_data
  station habitat camera_days active
1    CT01  Forest          30   TRUE
2    CT02  Forest          28   TRUE
3    CT03 Pasture          30  FALSE
4    CT04 Pasture          27   TRUE

str()

Shows the internal structure of the object:

str(camera_data)
'data.frame':   4 obs. of  4 variables:
 $ station    : chr  "CT01" "CT02" "CT03" "CT04"
 $ habitat    : chr  "Forest" "Forest" "Pasture" "Pasture"
 $ camera_days: num  30 28 30 27
 $ active     : logi  TRUE TRUE FALSE TRUE

dim()

Returns the number of rows and columns:

dim(camera_data)
[1] 4 4

nrow()

Returns the number of rows:

nrow(camera_data)
[1] 4

ncol()

Returns the number of columns:

ncol(camera_data)
[1] 4

names()

Returns the column names:

names(camera_data)
[1] "station"     "habitat"     "camera_days" "active"     
Tip

When you load a new dataset into R, one of the first things you should do is inspect it.

Functions such as:

head()
str()
dim()
names()

can help you identify problems before you begin an analysis.

Selecting columns from a data frame

We can access a column using $:

camera_data$station
[1] "CT01" "CT02" "CT03" "CT04"
camera_data$camera_days
[1] 30 28 30 27

Notice what happens when we do this:

class(camera_data$camera_days)
[1] "numeric"

The column itself is a vector.

This is an important idea:

A data frame is essentially a collection of equal-length vectors organized as columns.

We can also use square brackets.

For example:

camera_data[1, 2]
[1] "Forest"

means:

row 1, column 2

To select the first row:

camera_data[1, ]
  station habitat camera_days active
1    CT01  Forest          30   TRUE

To select the second column:

camera_data[, 2]
[1] "Forest"  "Forest"  "Pasture" "Pasture"

We can also use column names:

camera_data[, "habitat"]
[1] "Forest"  "Forest"  "Pasture" "Pasture"

Or select several columns:

camera_data[, c("station", "camera_days")]
  station camera_days
1    CT01          30
2    CT02          28
3    CT03          30
4    CT04          27

Later we will learn easier and more readable ways to select rows and columns using dplyr.

For now, it is important to understand how data frames work in base R.

Adding a column

We can add a new column using $.

For example:

camera_data$elevation <- c(
  105,
  112,
  98,
  101
)

camera_data
  station habitat camera_days active elevation
1    CT01  Forest          30   TRUE       105
2    CT02  Forest          28   TRUE       112
3    CT03 Pasture          30  FALSE        98
4    CT04 Pasture          27   TRUE       101

The new column must have the correct number of values.

Our data frame has four rows, so the new column needs four values.

We can verify the number of rows with:

nrow(camera_data)
[1] 4

Changing values

We can also replace specific values.

For example:

camera_data$camera_days[2] <- 30

camera_data
  station habitat camera_days active elevation
1    CT01  Forest          30   TRUE       105
2    CT02  Forest          30   TRUE       112
3    CT03 Pasture          30  FALSE        98
4    CT04 Pasture          27   TRUE       101

Or using row and column positions:

camera_data[2, "camera_days"] <- 30

This can be useful, but we should be careful when modifying raw data manually.

Important

Whenever possible, avoid manually changing your original data file.

It is usually better to make corrections through code so that the changes are documented and reproducible.

Different structures, different purposes

Let’s summarize the four structures we have seen.

Structure Dimensions Same data type? Common use
Vector 1 Yes Store a sequence of values
Matrix 2 Yes Numeric or structured rectangular data
List 1 No Store different objects together
Data frame 2 Different types allowed by column Store datasets

The structure we choose depends on the type of information we want to store.

Throughout the rest of the course:

  • we will constantly work with vectors,
  • most of our datasets will be data frames,
  • we will occasionally encounter matrices,
  • and many R functions will return results as lists.

When we eventually work with camtrapR, all four of these ideas will appear again.

Converting between data structures

R also provides functions for converting objects from one structure to another.

For example:

as.data.frame(detection_matrix)
     Raccoon Coyote Bobcat
CT01       5      2      0
CT02     100    200    300
CT03       8      0      2

converts a matrix into a data frame.

We can save the result:

detection_data <- as.data.frame(detection_matrix)

class(detection_data)
[1] "data.frame"

Similarly:

as.matrix(camera_data)
     station habitat   camera_days active  elevation
[1,] "CT01"  "Forest"  "30"        "TRUE"  "105"    
[2,] "CT02"  "Forest"  "30"        "TRUE"  "112"    
[3,] "CT03"  "Pasture" "30"        "FALSE" " 98"    
[4,] "CT04"  "Pasture" "27"        "TRUE"  "101"    

attempts to convert a data frame into a matrix.

However, remember that a matrix can contain only one basic data type.

Because camera_data contains text, numbers, and logical values, R will need to convert everything to a common type.

Try:

camera_matrix_2 <- as.matrix(camera_data)

camera_matrix_2
     station habitat   camera_days active  elevation
[1,] "CT01"  "Forest"  "30"        "TRUE"  "105"    
[2,] "CT02"  "Forest"  "30"        "TRUE"  "112"    
[3,] "CT03"  "Pasture" "30"        "FALSE" " 98"    
[4,] "CT04"  "Pasture" "27"        "TRUE"  "101"    

and then:

class(camera_matrix_2)
[1] "matrix" "array" 

This is another good example of why understanding data structures matters.

A function may technically convert an object, but the result may not have the structure or data type that you expected.

Practice at home

Tip

This activity should take approximately 15–20 minutes.

Open your intro-r-course project and create a new script called:

practice-03.R

Save it inside the scripts folder.

Part 1: Vectors

Create these objects:

stations <- c("CT01", "CT02", "CT03", "CT04", "CT05")

detections <- c(12, 5, 0, 18, 7)

camera_days <- c(30, 28, 30, 25, 29)

Then:

  1. Return the third element of stations.
  2. Return the first three values of detections.
  3. Determine which values of detections are greater than 10.
  4. Return only the detections greater than 10.
  5. Calculate the mean number of camera-days.

Part 2: Data frame

Create a data frame called survey_data using:

stations
[1] "CT01" "CT02" "CT03" "CT04" "CT05"
detections
[1] 12  5  0 18  7
camera_days
[1] 30 28 30 25 29

Then:

  1. Display the data frame.
  2. Use str() to inspect its structure.
  3. Use dim() to find its dimensions.
  4. Use names() to obtain its column names.
  5. Select only the detections column.
  6. Select the information from the second camera station.
  7. Add a new column called habitat containing:
c("Forest", "Forest", "Pasture", "Forest", "Pasture")
[1] "Forest"  "Forest"  "Pasture" "Forest"  "Pasture"
  1. Use str() again and look at the type of each column.
# Practice for Module 3

# ---------------------------
# Part 1: Vectors
# ---------------------------

stations <- c(
  "CT01",
  "CT02",
  "CT03",
  "CT04",
  "CT05"
)

detections <- c(
  12,
  5,
  0,
  18,
  7
)

camera_days <- c(
  30,
  28,
  30,
  25,
  29
)

# Third station
stations[3]
[1] "CT03"
# First three detection values
detections[1:3]
[1] 12  5  0
# Which values are greater than 10?
detections > 10
[1]  TRUE FALSE FALSE  TRUE FALSE
# Return only detections greater than 10
detections[detections > 10]
[1] 12 18
# Mean camera-days
mean(camera_days)
[1] 28.4
# ---------------------------
# Part 2: Data frame
# ---------------------------

survey_data <- data.frame(
  stations,
  detections,
  camera_days
)

# Display the data
survey_data
  stations detections camera_days
1     CT01         12          30
2     CT02          5          28
3     CT03          0          30
4     CT04         18          25
5     CT05          7          29
# Inspect structure
str(survey_data)
'data.frame':   5 obs. of  3 variables:
 $ stations   : chr  "CT01" "CT02" "CT03" "CT04" ...
 $ detections : num  12 5 0 18 7
 $ camera_days: num  30 28 30 25 29
# Dimensions
dim(survey_data)
[1] 5 3
# Column names
names(survey_data)
[1] "stations"    "detections"  "camera_days"
# Select detections
survey_data$detections
[1] 12  5  0 18  7
# Select the second camera station
survey_data[2, ]
  stations detections camera_days
2     CT02          5          28
# Add habitat
survey_data$habitat <- c(
  "Forest",
  "Forest",
  "Pasture",
  "Forest",
  "Pasture"
)

# Inspect the updated data frame
str(survey_data)
'data.frame':   5 obs. of  4 variables:
 $ stations   : chr  "CT01" "CT02" "CT03" "CT04" ...
 $ detections : num  12 5 0 18 7
 $ camera_days: num  30 28 30 25 29
 $ habitat    : chr  "Forest" "Forest" "Pasture" "Forest" ...

The important thing is that you understand the difference between:

  • a vector,
  • a matrix,
  • a list,
  • and a data frame,

and that you can recognize how rows, columns, and positions are used to access information.