
Interview-Based Catch Estimation
Source:vignettes/interview-estimation.Rmd
interview-estimation.RmdIntroduction
This vignette extends the effort estimation workflow covered in “Getting Started with tidycreel” to interview-based catch and harvest estimation. The complete workflow involves five steps:
- Design: Define your survey calendar and stratification
- Counts: Attach instantaneous count observations
- Interviews: Attach complete trip interview data
- CPUE: Estimate catch per unit effort
- Total Catch: Combine effort and CPUE estimates
The package uses ratio-of-means estimation for catch per unit effort (CPUE), which is appropriate for access point surveys with complete trip interviews. This estimator accounts for the correlation between catch and effort within each interview.
Survey Design and Count Data
We start with the same design and count data workflow from the “Getting Started” vignette:
library(tidycreel)
# Load example calendar and counts
data(example_calendar)
data(example_counts)
# Create design with counts
design <- creel_design(example_calendar, date = date, strata = day_type)
design <- add_counts(design, example_counts)
#> Warning in svydesign.default(ids = psu_formula, strata = strata_formula, : No
#> weights or probabilities supplied, assuming equal probability
# Estimate total effort
effort_est <- estimate_effort(design)
#> Warning: Instantaneous counts were expanded without a period length.
#> ℹ No `period_length_col` was supplied to `add_counts()`, so the estimate is the
#> count column summed over days.
#> ! If that column holds an instantaneous angler count, the result is in
#> angler-days, not angler-hours.
#> ℹ Supply the period each count was randomised within: `add_counts(design,
#> counts, period_length_col = <col>)`.
#> This warning is displayed once per session.
print(effort_est)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 7
#> estimate se se_between se_within ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 372. 13.2 13.2 0 344. 401. 14The effort estimate provides the total angler-hours for the survey period, which we’ll combine with CPUE to estimate total catch.
Adding Interview Data
Next, we attach interview data containing catch and effort for complete trips:
# Load example interview data
data(example_interviews)
head(example_interviews)
#> date hours_fished catch_total catch_kept trip_status trip_duration
#> 1 2024-06-01 2.0 5 2 complete 2.0
#> 2 2024-06-01 3.5 8 5 complete 3.5
#> 3 2024-06-02 1.5 2 1 complete 1.5
#> 4 2024-06-02 2.0 3 2 incomplete 1.0
#> 5 2024-06-03 2.5 6 3 complete 2.5
#> 6 2024-06-03 4.0 12 8 complete 4.0
#> interview_id angler_type angler_method species_sought n_anglers refused
#> 1 1 bank bait walleye 2 FALSE
#> 2 2 boat artificial walleye 1 FALSE
#> 3 3 bank bait bass 3 FALSE
#> 4 4 bank fly panfish 2 FALSE
#> 5 5 boat artificial walleye 1 FALSE
#> 6 6 boat bait bass 4 FALSE
# Attach interviews to the design
design <- add_interviews(design, example_interviews,
catch = catch_total,
effort = hours_fished,
n_anglers = n_anglers,
harvest = catch_kept,
trip_status = trip_status,
trip_duration = trip_duration
)
#> ℹ Added 22 interviews: 17 complete (77%), 5 incomplete (23%)
print(design)
#>
#> ── Creel Survey Design ─────────────────────────────────────────────────────────
#> Type: "instantaneous"
#> Date column: date
#> Strata: day_type
#> Calendar: 14 days (2024-06-01 to 2024-06-14)
#> day_type: 2 levels
#> Counts: 14 observations
#> PSU column: date
#> Count column: effort_hours
#> Party-size term: not carried
#> Count type: "instantaneous"
#> Survey: <survey.design2> (constructed)
#> Interviews: 22 observations
#> Type: "access"
#> Catch: catch_total
#> Effort: hours_fished
#> Harvest: catch_kept
#> Trip status: 17 complete, 5 incomplete
#> Party size: n_anglers
#> Survey: <survey.design2> (constructed)
#> Sections: "none"The add_interviews() function maps the interview data
columns to the design structure. The design now shows both count and
interview data attached. Interview data is treated as a parallel data
stream to count data—the two datasets do not need to align on specific
dates.
Estimating CPUE
With interview data attached, we can estimate catch per unit effort:
# Estimate CPUE
cpue_est <- estimate_catch_rate(design)
#> ℹ Using complete trips for CPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for CPUE estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.
print(cpue_est)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Ratio-of-Means CPUE
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Unit: fish/angler-hour
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 0.967 0.110 0.752 1.18 17The CPUE estimate uses a ratio-of-means estimator: total catch divided by total effort across all interviews. This estimator is appropriate for access point surveys because it accounts for the correlation between catch and effort within each trip. The variance is computed using the delta method, accounting for the covariance between numerator and denominator.
For details on the ratio-of-means formula and variance calculation,
see ?estimate_catch_rate.
Estimating Total Catch
We can combine the effort and CPUE estimates to compute total catch:
# Estimate total catch
total_catch_est <- estimate_total_catch(design)
#> Warning: `estimate_total_catch()` is pooling over domains the counts do not classify:
#> angler_method and species_sought.
#> ! The rate differs across their levels in these interviews (angler_method:
#> artificial 1.121, bait 0.8, fly 1.31 and species_sought: bass 1.018, panfish
#> 1.2, walleye 0.933), so the total depends on the interview sample's mix over
#> those domains.
#> ℹ Without the domain in the counts the total is `E_total * rate_pooled`,
#> weighted by the interview mix rather than the effort mix. Interview selection
#> is not proportional to effort by construction (Malvestuto 1996).
#> ℹ This is a risk, not an error: the counts carry no composition to check
#> against, so it cannot be verified from the data.
#> ℹ Classifying angler_method and species_sought in the count data removes the
#> assumption -- the total becomes `sum(E_h * rate_h)`.
#> This warning is displayed once per session.
print(total_catch_est)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total Catch (Effort × CPUE)
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 364. 47.9 262. 466. 17The total catch estimate multiplies the effort and CPUE estimates. The variance is computed using the delta method with the formula:
Var(E × C) = E² Var(C) + C² Var(E)
This formula assumes independence between the count and interview
data streams, which is appropriate since they are collected through
separate sampling processes. See ?estimate_total_catch for
more details on the variance propagation.
Estimating Harvest
The package distinguishes between total catch (all fish caught) and harvest (fish kept). We can estimate harvest per unit effort (HPUE) and total harvest:
# Estimate HPUE
hpue_est <- estimate_harvest_rate(design)
#> ℹ Filtering to complete trips for HPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for harvest estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.
print(hpue_est)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Ratio-of-Means HPUE
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Unit: fish/angler-hour
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 0.595 0.0616 0.475 0.716 17
# Estimate total harvest
total_harvest_est <- estimate_total_harvest(design)
#> Warning: `estimate_total_harvest()` is pooling over a domain the counts do not classify:
#> angler_method.
#> ! The rate differs across its levels in these interviews (angler_method:
#> artificial 0.667, bait 0.483, fly 0.897), so the total depends on the
#> interview sample's mix over that domain.
#> ℹ Without the domain in the counts the total is `E_total * rate_pooled`,
#> weighted by the interview mix rather than the effort mix. Interview selection
#> is not proportional to effort by construction (Malvestuto 1996).
#> ℹ This is a risk, not an error: the counts carry no composition to check
#> against, so it cannot be verified from the data.
#> ℹ Classifying angler_method in the count data removes the assumption -- the
#> total becomes `sum(E_h * rate_h)`.
#> This warning is displayed once per session.
print(total_harvest_est)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total Harvest (Effort × HPUE)
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 225. 28.1 165. 284. 17Harvest estimation uses the same ratio-of-means approach as CPUE, but
with the harvest column (catch_kept) as the numerator. As
expected, total harvest is lower than total catch since not all caught
fish are kept.
Grouped Estimation
Like effort estimation, CPUE and total catch functions support
grouped estimation using the by parameter:
# Estimate CPUE by day_type (not evaluated - example data too small)
cpue_by_day <- estimate_catch_rate(design, by = day_type)
# Estimate total catch by day_type
total_catch_by_day <- estimate_total_catch(design, by = day_type)Note: The example data has insufficient interview sample sizes for grouped estimation (weekend n=9, weekday n=13). Ratio estimation requires at least 10 observations per group for stability. In real surveys, aim for at least 30 interviews per stratum for reliable grouped estimates.
Variance Methods
Like effort estimation, CPUE estimation supports multiple variance methods. The default is Taylor linearization, but bootstrap and jackknife are also available:
# Bootstrap variance estimation
set.seed(42) # For reproducibility
cpue_boot <- estimate_catch_rate(design, variance = "bootstrap")
#> ℹ Using complete trips for CPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for CPUE estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.
# Jackknife variance estimation
cpue_jk <- estimate_catch_rate(design, variance = "jackknife")
#> ℹ Using complete trips for CPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for CPUE estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.All three methods produce similar results for these data. Use bootstrap or jackknife when you want to verify Taylor linearization assumptions or when working with complex grouped estimates.
When to use each method:
- Taylor linearization (default): Computationally efficient and appropriate for most smooth statistics. This is the recommended default.
- Bootstrap: Use when working with non-smooth statistics or when you want to verify Taylor linearization assumptions. More computationally intensive.
- Jackknife: Alternative resampling method that is deterministic (unlike bootstrap). Useful for verification or when bootstrap is too slow.
Age Structure Estimation
Age data collected during interviews can be used to estimate a
pressure-weighted age distribution and design-weighted mean age for the
catch or harvest. Attach age records with add_ages(), then
call est_age_distribution() followed by
est_mean_age().
data(example_ages)
data(example_catch)
head(example_ages)
#> interview_id species age age_type
#> 1 1 walleye 4 harvest
#> 2 1 walleye 3 harvest
#> 3 1 walleye 5 harvest
#> 4 2 bass 2 harvest
#> 5 2 bass 3 harvest
#> 6 6 walleye 6 harvest
# Ages are read from a subsample of the catch, so the age-class totals are
# scaled onto the reported catch. Grouping by species needs this table, which
# is the only record of catch per species.
design <- add_catch(
design,
example_catch,
catch_uid = interview_id,
interview_uid = interview_id,
species = species,
count = count,
catch_type = catch_type
)
# Attach age data — one row per aged fish
design <- add_ages(
design,
example_ages,
age_uid = interview_id,
interview_uid = interview_id,
species = species,
age = age,
age_type = age_type
)
# Estimate weighted age frequency by species
ad <- est_age_distribution(design, by = species)
#> Warning: ! Age totals were rescaled onto the reported catch.
#> ℹ Measured fish (weighted): 18; reported: 93 -- a factor of 5.17.
#> ℹ estimate, se and the confidence bounds describe the REPORTED catch, estimated
#> from the measured subsample. Shares (percent) are unaffected.
print(ad)
#> species age estimate se ci_lower ci_upper percent cumulative_percent
#> 1 bass 2 10.00000 4.233202 1.7030763 18.296924 40.0 40.0
#> 2 bass 3 15.00000 6.524569 2.2120798 27.787920 60.0 100.0
#> 3 panfish 0 3.25000 1.899150 -0.4722657 6.972266 25.0 25.0
#> 4 panfish 1 6.50000 4.023369 -1.3856588 14.385659 50.0 75.0
#> 5 panfish 2 3.25000 3.991201 -4.5726107 11.072611 25.0 100.0
#> 6 walleye 3 12.22222 4.069220 4.2466972 20.197747 22.2 22.2
#> 7 walleye 4 18.33333 6.136191 6.3066208 30.360046 33.3 55.6
#> 8 walleye 5 12.22222 4.069220 4.2466972 20.197747 22.2 77.8
#> 9 walleye 6 12.22222 8.491868 -4.4215332 28.865978 22.2 100.0
#> n
#> 1 3
#> 2 3
#> 3 2
#> 4 2
#> 5 2
#> 6 3
#> 7 3
#> 8 3
#> 9 3Ages come from a subsample of the catch, so
est_age_distribution() reports a two-phase estimate: the
age composition among aged fish, scaled onto the design-estimated
reported catch. The call warns when it rescales, naming the factor.
Shares (percent) are unaffected by how many fish were
aged.
Each row is one integer age class. The estimate column
is the survey-design-weighted count of fish at that age,
percent is the within-species proportion, and
cumulative_percent accumulates from the youngest age class
up. The same variance, type, and
conf_level arguments available in
est_length_distribution() are accepted here.
# Compute weighted mean age from the distribution object
est_mean_age(ad)
#> species mean_age mean_age_se mean_age_ci_lower mean_age_ci_upper
#> 1 bass 2.600000 0.1456703 2.3144914 2.885509
#> 2 panfish 1.000000 0.3400005 0.3336113 1.666389
#> 3 walleye 4.444444 0.2706522 3.9139759 4.974913est_mean_age() applies the ratio estimator (Σ a N̂_a / N̂)
to the weighted age counts returned by
est_age_distribution() and propagates variance via the
delta method.
Complete Workflow Example
The full pipeline in one workflow is:
# Load data
data(example_calendar)
data(example_counts)
data(example_interviews)
# Build design and attach data
complete_design <- creel_design(example_calendar, date = date, strata = day_type) |>
add_counts(example_counts) |>
add_interviews(example_interviews,
catch = catch_total,
effort = hours_fished,
n_anglers = n_anglers,
harvest = catch_kept,
trip_status = trip_status,
trip_duration = trip_duration
)
#> Warning in svydesign.default(ids = psu_formula, strata = strata_formula, : No
#> weights or probabilities supplied, assuming equal probability
#> ℹ Added 22 interviews: 17 complete (77%), 5 incomplete (23%)
# Compute all estimates
effort <- estimate_effort(complete_design)
cpue <- estimate_catch_rate(complete_design)
#> ℹ Using complete trips for CPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for CPUE estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.
hpue <- estimate_harvest_rate(complete_design)
#> ℹ Filtering to complete trips for HPUE estimation
#> (n=17, 77.3% of 22 interviews) [default]
#> Warning: Small sample size for harvest estimation.
#> ! Sample size is 17. Ratio estimates are more stable with n >= 30.
#> ℹ Variance estimates may be unstable with n < 30.
total_catch <- estimate_total_catch(complete_design)
total_harvest <- estimate_total_harvest(complete_design)
# Print key results
print(effort)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 7
#> estimate se se_between se_within ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 372. 13.2 13.2 0 344. 401. 14
print(total_catch)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total Catch (Effort × CPUE)
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 364. 47.9 262. 466. 17
print(total_harvest)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Total Harvest (Effort × HPUE)
#> Variance: Taylor linearization
#> Confidence level: 95%
#> Effort target: sampled_days
#>
#> # A tibble: 1 × 5
#> estimate se ci_lower ci_upper n
#> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 225. 28.1 165. 284. 17This demonstrates the complete v0.2.0 workflow from survey design through total catch and harvest estimation.
Working with Incomplete Trips
This vignette demonstrates the default workflow using complete trips, which is the recommended approach following Pollock et al. (1994) roving-access design principles.
For situations with incomplete-trip interviews where you want to
examine using them for estimation, see the Incomplete Trip
Estimation vignette
(vignette("incomplete-trips", package = "tidycreel")). That
vignette covers:
- When incomplete trip estimation is scientifically valid
- How to validate incomplete trip estimates using TOST equivalence testing
- Step-by-step workflow with
validate_incomplete_trips() - Examples of passing and failing validation scenarios
- Why you should NEVER pool complete and incomplete trips
Important: The package defaults to complete trips
only. Incomplete trip estimation requires explicit opt-in via the
use_trips parameter in estimate_catch_rate()
and should only be used after validation with
validate_incomplete_trips().
Next Steps
For more details on interview-based estimation functions, see:
-
?add_interviews- Attach interview data to a design -
?estimate_catch_rate- Estimate catch per unit effort -
?estimate_harvest_rate- Estimate harvest per unit effort -
?estimate_total_catch- Estimate total catch -
?estimate_total_harvest- Estimate total harvest -
?add_ages- Attach age data to a design -
?est_age_distribution- Estimate weighted age frequency -
?est_mean_age- Compute design-weighted mean age -
?example_interviews- Example interview dataset -
?example_ages- Example age dataset
For the effort estimation workflow, see the “Getting Started with tidycreel” vignette.