Population denominators: bridged, single-race, and legacy
Source:vignettes/population-denominators.Rmd
population-denominators.RmdA mortality rate is death counts over a matching population, and
“matching” is the whole problem. NCHS has coded race three incompatible
ways over the series, so narcan ships three denominator schemes and
routes all of them through one guarded join,
add_pop_counts(). The schemes are NOT comparable – you
cannot chain them into one trend – so set race_scheme
deliberately.
| Scheme | race_scheme |
Source | Coverage | Use for |
|---|---|---|---|---|
| Legacy (historical NCHS) |
"legacy" (default) |
frozen pop_est (single-race-alone PEP) |
1979-2020 | reproducing prior narcan-based analyses byte-for-byte |
| Single-race | "single" |
Census PEP single-race (OMB 1997) | 2000-2024 | 2021+ single-race deaths (codes 101-106) |
| SEER bridged | "bridged" |
SEER U.S. Population Data (Vintage 2024) | 1969-2024 (population) | coherent bridged-race trends, deaths through 2020 |
(PEP = Census Population Estimates Program; OMB 1997 = the 1997 OMB race/ethnicity standards; SEER = NCI’s Surveillance, Epidemiology, and End Results program.)
All data here are public (bundled Census/SEER estimates, or a small bundled county fixture) plus synthetic death counts; no restricted NCHS records are used, and no chunk downloads anything.
Across every scheme, add_pop_counts() joins on
by_vars – default
c("year", "age", "sex", "race"), which returns a national
denominator; adding state_fips or county_fips
to by_vars routes the same call to sub-national
denominators.
add_pop_counts(), scheme by scheme
Legacy (the default)
race_scheme = "legacy" joins the frozen
pop_est and reproduces narcan’s historical behavior
byte-for-byte: unmatched keys warn and leave pop = NA.
pop_est is a pieced-together series whose race labels are
white/black/other/nhw
(not the four bridged categories) and whose 2000-2020 denominators are
single-race-alone Census estimates. It exists to reproduce prior
narcan-based analyses exactly – it is not the official NCHS
bridged-race series. For a coherent bridged-race denominator use
"bridged"; for 2000+ single-race deaths use
"single".
legacy <- data.frame(year = 1999L, age = 25L, sex = "male", race = "white")
add_pop_counts(legacy) # default scheme; adds `pop`
#> year age sex race pop
#> 1 1999 25 male white 7289220For 2000-2020 the legacy denominator is single-race-alone, which
undercounts the multiple-race population and so inflates race-specific
rates (especially for the AIAN and API groups);
add_pop_counts() nudges you once per session in that span
toward "bridged" or "single".
Single-race
race_scheme = "single" joins single-race denominators
(2000-2024) for deaths coded 101-106
(white_only/black_only/american_indian_only/asian_only/
nhopi_only/multiracial; nhopi =
Native Hawaiian or Other Pacific Islander). It is strict: out-of-domain
age/sex/race and any unmatched key hard-error, so a denominator is never
silently NA.
deaths <- expand.grid(
year = 2024L, age = seq(0, 85, 5), sex = c("male", "female"),
race = "asian_only", stringsAsFactors = FALSE
)
deaths$deaths <- rep(c(1, 2, 5, 12, 30, 55), length.out = nrow(deaths))
head(deaths)
#> year age sex race deaths
#> 1 2024 0 male asian_only 1
#> 2 2024 5 male asian_only 2
#> 3 2024 10 male asian_only 5
#> 4 2024 15 male asian_only 12
#> 5 2024 20 male asian_only 30
#> 6 2024 25 male asian_only 55add_pop_counts() adds the matched pop
column:
rated <- add_pop_counts(deaths, race_scheme = "single")
head(rated)
#> year age sex race deaths pop pop_scheme
#> 1 2024 0 male asian_only 1 600728 single
#> 2 2024 5 male asian_only 2 678658 single
#> 3 2024 10 male asian_only 5 659666 single
#> 4 2024 15 male asian_only 12 666097 single
#> 5 2024 20 male asian_only 30 760220 single
#> 6 2024 25 male asian_only 55 826163 singleFrom the matched pop, the rate pipeline
(add_std_pop() -> calc_asrate_var() ->
calc_stdrate_var()) turns these counts into an
age-standardized rate. That pipeline – including the standard-population
choice and confidence intervals – is walked through step by step in
vignette("age-standardized-rates").
Bridged (SEER-uniform, era-ragged)
race_scheme = "bridged" joins the SEER bridged
denominators (1969-2024). It is strict and era-ragged (the set of valid
race and origin categories changes by year, so coverage is uneven across
eras): SEER resolves AIAN/API (American Indian or Alaska Native / Asian
or Pacific Islander) and Hispanic origin only from 1990 (pre-1990 is
white/black/other), so year MUST be a join key.
The “1969-2024” coverage is population availability. On the
death side, NCHS bridged-race coding ends with data year 2020 (2021 is
reserved; single-race codes 101-106 begin in 2022), so there is no
bridged-race-coded death series to pair with the 2021-2024 population –
use "single" for 2021+ deaths. Treat "bridged"
as a coherent race series through 2020.
api <- data.frame(year = 2019L, age = 40L, sex = "female", race = "api")
add_pop_counts(api, race_scheme = "bridged",
by_vars = c("year", "age", "sex", "race"))
#> year age sex race pop pop_scheme
#> 1 2019 40 female api 874749 bridgedOmitting year from by_vars is an error –
the valid race set is era-dependent:
add_pop_counts(api, race_scheme = "bridged",
by_vars = c("age", "sex", "race"))
#> Error:
#> ! add_pop_counts(): `df` carries population-dimension column(s) `year` not in `by_vars`; under a strict race_scheme ("single"/"bridged") the denominator would be silently summed over them. Add them to `by_vars`, or drop them from `df` to aggregate that dimension.And an api row before 1990 errors, because SEER has no
AIAN/API split yet:
old_api <- data.frame(year = 1985L, age = 40L, sex = "female", race = "api")
add_pop_counts(old_api, race_scheme = "bridged",
by_vars = c("year", "age", "sex", "race"))
#> Error:
#> ! add_pop_counts(): race value(s) 'api' are not denominable under race_scheme = "bridged" before 1990 (SEER pre-1990 race = white/black/other only; the AIAN/API split begins in 1990). Valid pre-1990: white, black, other (or "total"). Restrict to year >= 1990, or collapse to `other`.Strictness guards
Legacy and bridged share the labels white/black/other/total, so a
mis-set race_scheme cannot always be detected. But
single-race labels under a non-single scheme always can be – passing
them to the default is a hard error, a guard against forgetting
race_scheme = "single":
bad <- data.frame(year = 2024L, age = 25L, sex = "male", race = "asian_only")
add_pop_counts(bad) # default is "legacy" -> error
#> Error:
#> ! add_pop_counts(): `race` holds single-race values (101-106 or *_only/multiracial) but race_scheme is not "single". Pass race_scheme = "single" to use single-race denominators.A frame already summed over a dimension uses a reserved token –
race = "total", sex = "both",
hispanic_origin = "all". The matching marginal is
synthesized on demand (finest cells summed; no marginal is stored, so it
cannot double-count):
both_sex <- data.frame(year = 2024L, age = 65L, sex = "both", race = "asian_only")
add_pop_counts(both_sex, race_scheme = "single")$pop # male + female
#> [1] 1042672State and county grains
Include state_fips (or county_fips) in
by_vars to route to sub-national denominators. The state
single-race table is bundled, so a state join needs no extra package.
See vignette("geography-fips") for how to derive harmonized
state_fips/county_fips columns from raw NCHS
fields.
ca <- data.frame(state_fips = "06", year = 2024L, age = 40L, sex = "female",
race = "asian_only", deaths = 20)
add_pop_counts(ca, race_scheme = "single",
by_vars = c("state_fips", "year", "age", "sex", "race"))
#> state_fips year age sex race deaths pop pop_scheme
#> 1 06 2024 40 female asian_only 20 269078 singleA geography column that is in the frame but absent from
by_vars is an error, never a silent national join:
add_pop_counts(ca, race_scheme = "single",
by_vars = c("year", "age", "sex", "race"))
#> Error:
#> ! add_pop_counts(): `df` carries population-dimension column(s) `state_fips` not in `by_vars`; under a strict race_scheme ("single"/"bridged") the denominator would be silently summed over them. Add them to `by_vars`, or drop them from `df` to aggregate that dimension.Public MCOD carries state/county geographic detail only through data year 2004 (2005+ county detail needs restricted NCHS files). The county denominators are distributed as a downloadable parquet; here we point narcan at the small bundled Wyoming fixture instead of a live download.
This example needs the optional duckdb package
(install.packages("duckdb")); with it absent, the chunk
below simply produces no output.
fx <- system.file("extdata", "pop_singlerace_county_fixture.parquet",
package = "narcan")
wy_deaths <- data.frame(
county_fips = "56013", year = 1999:2004, age = 40L, sex = "male",
race = "white_only", deaths = c(1, 2, 4, 3, 5, 2)
)
wy_deaths
#> county_fips year age sex race deaths
#> 1 56013 1999 40 male white_only 1
#> 2 56013 2000 40 male white_only 2
#> 3 56013 2001 40 male white_only 4
#> 4 56013 2002 40 male white_only 3
#> 5 56013 2003 40 male white_only 5
#> 6 56013 2004 40 male white_only 2
options(narcan.pop_single_county_parquet = fx) # normally a downloaded parquet
add_pop_counts(subset(wy_deaths, year >= 2000), race_scheme = "single",
by_vars = c("county_fips", "year", "age", "sex", "race"))
#> county_fips year age sex race deaths pop pop_scheme
#> 1 56013 2000 40 male white_only 2 1198 single
#> 2 56013 2001 40 male white_only 4 1195 single
#> 3 56013 2002 40 male white_only 3 1201 single
#> 4 56013 2003 40 male white_only 5 1180 single
#> 5 56013 2004 40 male white_only 2 1106 singleSingle-race denominators start in 2000, so the 1999 row is dropped before the join above – death-side coverage (public MCOD geography) and denominator-side coverage (single-race) can differ, so join only where both exist.
Coverage and the frozen slice
The national single-race table is bundled at every
covered year: the frozen 0.5.0 pop_singlerace (2020-2024)
plus pop_singlerace_full (2000-2024, dependency-free
.rda), so a national pre-2020 request needs no download.
The 2000-2019 backfill is additive – it does not change those five
years. The state and county backfills are the ones
distributed as downloadable single-race *_full parquets
(tag-pinned GitHub Release assets, cached on first use). So
get_pop_state()/get_pop_county() serve
2020-2024 from the bundled frozen table, but a pre-2020
years request at state/county grain routes to that
downloadable parquet – or to a local copy via the
narcan.pop_single_<grain>_parquet option, as in the
county chunk above. A request for a year the resolved data does not
cover hard-errors – never a silent 0-row or NA:
get_pop_state(scheme = "single", years = 2025)
#> Error:
#> ! get_pop_state(): single-race state denominators cover 2020-2024; year(s) 2025 were requested (pre-2020 years route to the backfill; there is no coverage past 2024).Provenance
pop_sources() prints the manifest – dataset, scheme,
grain, vintage, and coverage for every single-race and bridged dataset
(the frozen legacy pop_est predates this download-manifest
system). Check a bundled dataset’s vintage before mixing it with a
freshly downloaded file.
invisible(utils::capture.output(src <- pop_sources()))
knitr::kable(src[, c("dataset", "scheme", "grain", "vintage",
"year_min", "year_max")])| dataset | scheme | grain | vintage | year_min | year_max |
|---|---|---|---|---|---|
| pop_singlerace | single | national | V2024 | 2020 | 2024 |
| pop_singlerace_state | single | state | V2024 | 2020 | 2024 |
| pop_singlerace_full | single | national | int2000/int2010/V2024 | 2000 | 2024 |
| pop_bridged | bridged | national | seer_1969_2024 | 1969 | 2024 |
| pop_singlerace_county_full | single | county | int2000/int2010/V2024 | 2000 | 2024 |
| pop_singlerace_state_full | single | state | int2000/int2010/V2024 | 2000 | 2024 |
| pop_bridged_state | bridged | state | seer_1969_2024 | 1969 | 2024 |
| pop_bridged_county | bridged | county | seer_1969_2024 | 1969 | 2024 |
Hispanic-stratified denominators
Both strict schemes carry a Hispanic-origin axis. For an
origin-specific denominator, add a hispanic_origin column
(from add_hispanic_origin()) to by_vars; for
the all-origin denominator, drop the column entirely. The two answer
different questions – and origin-unknown deaths belong only in the
all-origin numerator, because there is no “unknown-origin” population to
divide by.
# All-origin: one row, every death counts (the 3 origin-unknown deaths are here).
all_origin <- data.frame(year = 2024L, age = 30L, sex = "female",
race = "white_only", deaths = 12 + 210 + 3)
add_pop_counts(all_origin, race_scheme = "single",
by_vars = c("year", "age", "sex", "race"))$pop
#> [1] 8373814
# Origin-stratified: the non-denominable unknowns are dropped; origin is a key.
# ('denominable' means it has a matching population to divide by.)
strat <- data.frame(
year = 2024L, age = 30L, sex = "female", race = "white_only",
hispanic_origin = c("hispanic", "non_hispanic"), deaths = c(12, 210)
)
add_pop_counts(strat, race_scheme = "single",
by_vars = c("year", "age", "sex", "race", "hispanic_origin")
)[, c("hispanic_origin", "deaths", "pop")]
#> hispanic_origin deaths pop
#> 1 hispanic 12 2199747
#> 2 non_hispanic 210 6174067The two stratified denominators sum to the all-origin one; the
difference is only which deaths sit in the numerator. Leaving a
hispanic_origin column in the frame but out of
by_vars is a hard error (it would be silently summed over),
as is an "unknown"/NA origin in a stratified
join. The detailed-vs-binary recode distinction, the pre-1990 mixed-era
trap, and the misclassification caveats are covered in
vignette("hispanic-origin").
A note on hand-joins
Always join through add_pop_counts() (or the guarded
rate helpers). The descriptive accessors
get_pop_state()/get_pop_county() return
population rows but do NOT guard a hand-join – if you
left_join() a slice yourself, or query the parquet with raw
SQL, you own verifying that the slice is unique on your keys and that no
key is silently dropped. The guarded path exists precisely to make those
mistakes impossible.
See also
-
vignette("getting-started")– the package overview and where population denominators fit in the workflow. -
vignette("age-standardized-rates")– feeds these denominators through a full age-standardized rate pipeline end to end.