Zero, Missing, or Unsampled? Reading Creel Diagnostics
What a validation report can tell us about real creel data

In the previous post, we combined counts and interview rates to estimate 3,072 fish caught. We used a small, deliberately tidy example so we could follow the calculation. Real creel data usually bring more questions: a blank catch entry, a selected day with no count, or a group of anglers represented in the counts but absent from the interviews.
Those gaps do not all mean the same thing. Zero catch is an observed outcome. A missing catch value is unknown. An unsampled day can be an intentional part of a probability sample. Treating all three as zero changes the data and can change the estimate.
tidycreel provides diagnostics to help locate these problems. I want to work through what those diagnostics tell us and where we still need to look at the sampling plan or the clerk’s notes before making a decision.
In brief
- Check raw tables before attaching them to a design.
- Compare observed records with planned visits, as well as the full calendar.
- Check every expected interview stratum, including strata with no records.
- Keep valid zero-catch trips and investigate missing values.
- Record the decision behind each correction or exclusion.
This post uses tidycreel 7.0.0, “Goldeye”, and the simulated June 1–16 example from the last two installments. You can download the complete R example to recreate the design and run every check below. The code deliberately introduces problems into copies of the data, leaving the original teaching example intact. The snippets below start with the calendar, counts, interviews, and design objects created in that script.
Start with the tables before building the design
validation_report() checks supplied count and interview tables. It does not need a creel_design object. That makes it useful when an export first arrives, before the design object supplies the sampling context.
Suppose an imported count is negative, another count has a July date in a June survey, and one interview has a missing catch value:
counts_import <- counts
interviews_import <- interviews
counts_import$n_anglers[1] <- -1
counts_import$date[2] <- as.Date("2024-07-05")
interviews_import$catch_total[2] <- NA_real_
raw_report <- validation_report(
counts = counts_import,
interviews = interviews_import,
na_threshold = 0,
date_range = range(calendar$date)
)
raw_reportThe report flags the negative n_anglers value as a failure, the count date as outside the supplied range, and catch_total as having missing data. The summary’s n_pass, n_warn, and n_fail count columns checked, not bad records. Its detail column identifies which columns need attention.
For the underlying column-level results, use validate_creel_data():
field_report <- validate_creel_data(
counts = counts_import,
interviews = interviews_import,
na_threshold = 0,
date_range = range(calendar$date)
)
subset(as.data.frame(field_report), status != "pass")| Table | Column | Check | Result |
|---|---|---|---|
| Counts | date |
Date range | One value outside June 1–16 |
| Counts | n_anglers |
Negative values | One negative value |
| Interviews | catch_total |
Missing-value rate | One of 40 values missing |
These are review prompts, not instructions to replace values. A negative count might be a database sentinel, a transcription mistake, or a problem in an import. The field sheet or database documentation determines the correction. Preserve the imported table and make corrections in a separate, documented step.
The 7.0.0 table checks also have limits. A type check reports a column’s class; it does not establish that the class matches your intended schema. A date-range check applies to a Date column, so parse imported date strings and check parsing failures first. Table-level checks also cannot establish that linked records describe the same trip. The keys introduced in the linked-tables post still need attention.
A passing threshold can still contain missing data
The default na_threshold is 0.10. A column warns when its missing fraction exceeds that threshold. One missing catch value among forty interviews is 2.5%, so this column passes the default missing-value check even though one trip still has an unknown catch value.
Here I set na_threshold = 0 to surface any missing values in the teaching example. That is a review setting, not a universal rule that every column must be complete. An optional comment field and the fishing-time denominator have different consequences for analysis.
Likewise, supplying the survey’s actual date range is more useful than accepting the broad default range of 1970–2100. Neither a threshold nor a green status establishes that an observation can safely enter an estimator. Check the fields that the intended calculation requires.
Compare calendar gaps with the actual sampling plan
The sampling-calendar post distinguished available days from selected days. Our calendar contains sixteen days; the survey sampled eight. Those eight intentionally unsampled days remain in the calendar because they belong to the population of days represented by the effort estimate.
coverage <- check_completeness(design)
coverage$missing_daysIn this example, $missing_days contains eight dates. In the 7.0.0 implementation, this component identifies calendar dates with no count records. It does not compare them with the selected schedule. The label alone therefore cannot tell us which dates represent missed fieldwork.
To answer that question, compare the planned visits with the observed records:
planned_dates <- data.frame(date = sample_dates)
observed_dates <- unique(counts["date"])
missing_visits <- dplyr::anti_join(
planned_dates, observed_dates, by = "date"
)
missing_visitsThere are no missing planned visits in the original example. Now remove the June 13 count to simulate a lost record:
counts_missing <- subset(counts, date != as.Date("2024-06-13"))
missing_visits <- dplyr::anti_join(
planned_dates, unique(counts_missing["date"]), by = "date"
)
missing_visitsThis check identifies June 13, a selected day. It needs follow-up: did the clerk miss the visit, did the record fail to upload, or was the survey schedule changed? An observed count of zero would mean the count occurred and found no anglers. An absent record does not establish that outcome.
This example has one count per selected day. With multiple shifts, count rounds, or sites, compare the full visit key, such as date, shift, site, and round. One count on a date does not establish that every scheduled visit occurred. Keep the full population calendar; shrinking it to observed dates would change the population represented by the expansion.
Look for strata with few interviews and strata with none
check_completeness() also flags interview strata below n_min, which defaults to ten. Our original sample has ten weekday and thirty weekend interviews. For illustration, a review threshold of fifteen flags the weekday stratum:
review <- check_completeness(design, n_min = 15L)
review$low_n_strataThe result has key = "weekday" and n_observed = 10. Fifteen is an illustrative review threshold here, not a sample-size recommendation. This component counts interview rows; it does not assess their independence, weights, or eligibility for a particular estimator. Enough rows to pass a threshold do not guarantee useful precision.
In 7.0.0, this low-n check counts strata appearing in the interview table. A stratum with no rows can be absent from the result. To illustrate why that matters, we will keep only the weekday interviews, then join their row counts to the strata in the calendar. Starting with the calendar makes the missing weekend interviews visible:
expected_strata <- unique(calendar["day_type"])
weekday_only <- subset(interviews, day_type == "weekday")
observed_n <- dplyr::count(
weekday_only, day_type, name = "n_interviews"
)
interview_coverage <- dplyr::left_join(
expected_strata, observed_n, by = "day_type"
)
interview_coverage$n_interviews[
is.na(interview_coverage$n_interviews)
] <- 0L
interview_coverage| Stratum | Interview rows |
|---|---|
| Weekday | 10 |
| Weekend | 0 |
Replacing the missing row count with zero here is appropriate: there are no weekend interviews in this deliberately reduced table. It does not assign a catch rate of zero. The weekend effort estimate still needs a supported rate before it can contribute to a fish total.
Repeat the coverage check for the observations eligible for your estimator. For the access-point example, inspect complete trips with recorded catch and positive fishing time. Roving interviews need the protocol and trip treatment described in the roving-interview post. Use all the relevant stratum columns when your design has more than day type.
Keep valid zero-catch trips in the denominator
Our forty interviews include eight trips with zero catch. They contribute sixteen angler-hours even though they contribute no fish. Removing those trips raises the sample’s apparent success rate:
| Interview sample | Fish reported | Angler-hours | Pooled fish per angler-hour |
|---|---|---|---|
| All forty trips | 140 | 80 | 1.7500 |
| Thirty-two successful trips | 140 | 64 | 2.1875 |
That is a 25% increase caused entirely by dropping unsuccessful trips. These pooled sample ratios illustrate the effect of deletion; the previous post’s full-period total still requires rates matched to weekday and weekend effort.
A missing catch value needs a different decision. Setting it to zero asserts that the angler caught nothing. Deleting the interview removes fishing time and changes the represented sample. Recover the value where possible; otherwise, document how missing outcomes affect the analysis and its assumptions. Using na.rm = TRUE does not resolve that question, especially when fish and hours then come from different sets of trips.
We are expanding tidycreel’s error checking
These examples also help explain why we are expanding tidycreel’s error checking. The review documented in #388 used a real 2025 reservoir season to show that a report can identify problems without stopping them from reaching an estimator. In 7.0.0, validate_creel_data() returns the report; it does not stop, remove, or change the supplied data. Calling it is an analyst’s choice, and a fail in the report does not itself stop the pipeline.
The proposed approach has three parts: stop on impossible values or missing fields required by the design when data are attached; warn about suspicious values with counts and examples; and retain a report for broader review. I want a problem such as a negative angler count to be visible where the data enter the analysis, with enough information to find and correct the record.
Related work addresses the coverage questions we just walked through: distinguishing unsampled days from missed scheduled visits (#391), deciding how to handle interviews on unsampled days (#454), and checking counts and interviews against declared shifts (#426). Recent development changes already check count dates and strata against the calendar (#449); those changes are beyond the 7.0.0 release used in this example.
The broader work is still in development. For now, the checks above give us a way to inspect the records, identify the questions, and document the decisions before calculating a total.
Carry the decisions into the report
I would keep a short record of each issue, the evidence used to resolve it, and the resulting action. That might be a correction supported by the original field sheet, a confirmed unsampled date, or an unresolved missing interview outcome that limits a reported total.
Before reporting this example, I would check that:
- The calendar still covers June 1–16 and the total uses
target = "period_total". - Each selected visit has a count or a documented explanation for its absence.
- Every effort stratum has interviews eligible for its rate estimator.
- Valid zero-catch trips remain in the analysis and required missing values have a documented treatment.
- Linked records, party size, and units agree across the count and interview components.
- The report states the estimator, package version, standard error, and confidence interval.
The series introduction described reproducibility as being able to explain how survey decisions become reported results. For me, that includes being able to explain why we corrected one record, kept a zero-catch trip, or left an unsampled day in the calendar. The real creel data, sampling plan, and clerk’s notes supply the evidence for those decisions.
The next step is to carry those checks, decisions, and estimates into a reproducible Quarto report that a colleague or manager can follow.