Hi there. I am attempting to use the CPS basic monthly survey to examine employment trends by occupation. When calculating standard errors for monthly employment by occupation, I tried using the BLS’s Generalized Variance Formulas and borrowed parameters from the broad SOC categories they provide parameters for. Unfortunately, the calculated standard errors were much larger than is reasonable (in many cases 30-50 times larger than the estimate), presumably because the GVF factors I borrowed were are supposed to be used on broader occupational categories with many more observations.
Is there a standard practice for calculating these standard errors or a clear way I should proceed?
Following CPS Technical Paper 77 (chapter 2-4), we typically recommend using replicate weights to estimate variances with CPS microdata (see our detailed CPS replicate weight user guide). However, replicate weights are only available for CPS supplements and are not provided for monthly BMS data. Since you reference borrowing GVF factors for variance estimation, I assume that you have reviewed the BLS instructions for calculating standard errors and confidence intervals. This guide states that “when considering multiple series to borrow from, using the 𝛼 and 𝛽 parameters that generate the highest standard error is generally advised”, though I understand that having standard errors that are 30-50 times larger than the estimates may not be particularly helpful.
While the PSU and strata sample design parameters are not released publicly, Davern et al. (2007) showed that specifying the lowest level of identifiable geography (sequentially as INDIVIDCC, COUNTY, METFIPS, and STATEFIP) as the strata, and household SERIAL (only unique to each household in a given survey month and year) as the cluster, performed reasonably well at estimating standard errors when compared to using the internal sample design data.
I’m not sure if you’re able to answer these types of questions, but in case you are:
Davern et al.’s approach focuses on estimating standard errors for ASEC samples. Would you also expect the approach to provide approximate standard errors for non-ASEC samples, like ORG samples?
I’m conducting one analysis with pooled data from the 1962-2025 ASEC samples, and another analysis with pooled data from the 1982-2025 ORG samples. However, the individcc, county, and metfips codes are inconsistent across those years. Thus, might it be a reasonable approach to specify “statefip” as the strata variable (and “serial” as the PSU), rather than using Davern’s lowest geography indicator for the strata? If not, do you have any alternative suggestions, preferably conservative ones?
I am not an expert on this, but from my knowledge of the CPS sampling methodology nothing specific comes to my mind that would cause Davern et al.'s approach to estimating standard errors using a survey design-based estimator (with a stratum and clustering variable) to not work for non-ASEC samples. With that said, my recommendation is to consult with colleagues or review the literature for guidance on where this approach has been tried. Something to consider is that monthly samples are generally smaller than the ASEC and are therefore likely to have greater variance with both the internal and the public use data.
Should you choose to proceed with this approach, note that selecting a coarser final stratum, such as the state, does work around changing codes for metro areas but will likely tradeoff for less accurate standard errors. IPUMS CPS just recently released the variable PLACEFIPS, which provides consistent coding of central and principal cities from September 1995-onwards. Additionally, PLACECENSUS provides consistent coding for samples from October 1985 - May 1995 (metropolitan areas are not identified in June - August 1995). Bridging only two or three different coding systems might hopefully be much easier than the larger number of vintages in METFIPS.
Please be aware that SERIAL is only unique within a sample month. A combination of YEAR, MONTH, and SERIAL provides a unique identifier for every household in IPUMS CPS. The exception to this are households in the March BMS samples, which also appear in the ASEC samples.
@Ivan_Strahof Thank you for this detailed guidance! I too am hoping to create reasonable confidence intervals for (1) total employment by occupation and (2) unemployment rates by occupation for monthly BMS data. However, I’d also like to create quarterly, half-year, and annual estimates as well.
Based on the guidance you’ve already provided, I’m planning to use YEAR + MONTH + SERIAL combinations as PSUs; for strata, I plan to use STATEFIP + PLACEFIPS (where PLACEFIPS is available) and STATEFIP (when no PLACEFIPS info is available). (If this isn’t an advisable approach, please let me know.)
I do have a few follow-up questions:
Just to confirm, it wouldn’t be accurate to use only the original person-level weights (WTFINL) to create confidence intervals, correct?
When using PLACEFIPS values as part of my strata definitions, would I also need to account for changes in codes over time in order to make my confidence intervals more accurate? [It seems like I won’t actually need to worry about this, since my analyses are from 2010 onward, but I thought I would ask in case I end up incorporating other variables into my strata whose values do change.]
In order to produce quarterly, half-year, and annual employment totals by occupation, I’m currently dividing each WTFINL value by the number of months present within each timeframe (e.g. 3 for most quarters, 12 for most years, etc). I would then feed those divided WTFINL values, rather than the original ones, into my calculations. However, would I need to adjust my confidence intervals somehow for the fact that the same respondents get surveyed multiple times in a given year? Or would using YEAR+MONTH+SERIAL as my PSU definition adjust for these repeated samples?
Sorry for all the questions–I’m pretty new to the CPS and want to make sure I’m not seriously underestimating my confidence intervals!
(My current calculations, if you’re curious, are available at this dashboard–but the confidence intervals shown were calculated using weight values only, with no inclusion of strata or PSUs. So I’m guessing they are narrower than they should be. The source code for my calculations is available here.)
WTFINL is the correct weight to use for (un)employment estimates by occupation and the weight should be divided by the total number of months in your analytical sample.
The approach for strata identification you have outlined is consistent with the recommendation from Davern et al. (2007). I will note that when using the smallest identifiable geographic unit (PLACEFIPS) as the stratum, many households will end up in the STATEFIP remainder. You might therefore add COUNTY as an intermediate stratum identifier. Households without an identified PLACEFIPS but with an identified COUNTY get their own stratum and only those with neither would fall back to state. Note that PLACEFIPS should always be combined with STATEFIP to uniquely identify cities. Additionally, in a cascading approach like this, it may make sense to restrict strata to only COUNTY and PLACEFIPS codes that appear across all vintages. A code that gets dropped or added later splits the same area across the three tiers, which can leave a stratum with too few PSUs and understate the error term. Since PLACEFIPS codes (in combination with STATEFIP) are unique across vintages (i.e., the same place should have the same code across vintages), there is no change in codes over time that needs to be accounted for.
On the choice of PSU, I would revise my recommendation and suggest using the unique household identifier CPSID rather than YEAR + MONTH + SERIAL in order to incorporate the rotating panel structure of the CPS. This adjusts intervals for the repeated sampling that you have noted.
I have now implemented your suggested methodology within my codebase. In order to ensure that all strata were present for all year/month combinations, I took the following approach:
I created an initial set of strata by combining PLACEFIPS, COUNTY, and STATEFIP values, in that order, into a single value (e.g., 44000_6037_california).
I then checked whether each stratum appeared within all year/month pairs. Any strata with a non-0 PLACEFIPS value that did not appear within all year/month pairs would have that value replaced with a 0. That way, these values would get consolidated into the ‘generic’ strata for that state-county pair. (For example, if ‘44000_6037_california’ didn’t appear within every year/month pair, it would get revised to ‘0_6037_california’.)
I then identified strata within this revised group that still didn’t appear within all year/month pairs. Any such strata with a non-0 COUNTY value would have that value replaced with a 0, thus making them equal to their ‘generic’ state-level strata values. (For instance, ‘0_6037_california’ would get revised to ‘0_0_california’.)
At this point, nearly all of my strata appeared within all year-month pairs. The one exception was Delaware, as its ‘generic’ statewide stratum (‘0_0_delaware’) only appeared within a small proportion of my year-month pairs. (There were 3 other strata, 0_10001_delaware, 0_10003_delaware, and 0_10005_delaware, that appeared in all year/month pairs.) I resolved this issue by making all Delaware strata equal to 0_0_delaware, but I’m not sure this was the right approach, given the resulting loss of precision. Perhaps it would make more sense to merge the statewide values into one of the county-level strata? (I might need to consult a map to figure out which of the three county-level strata these values should be merged into.)
The Python code I used to implement these updates, in case anyone is curious, is available here (and released under the MIT license).
Here’s a look at the different confidence intervals this method produced for different time periods: (Source)
I’m now wondering whether the annual values appear too small, given that the number of unique respondents is much smaller than the number of responses; however, I did use CPSID as my PSU value, which, as you explained, will compensate for the repeated samples.
Anyway, thanks again for your help–and I’d love your insights on what to do about Delaware!
It looks like you were able to make significant progress in graphing these estimates with the survey design confidence intervals! In general, we leave analytical decisions to individual researchers so I’ll try to answer your questions within the scope of what IPUMS User Support can provide guidance on.
To address your question about Delaware, I reviewed county identification for the state and was surprised to see cases with COUNTY = 0, especially in samples where all three Delaware counties are identified. I’ve reached out to my colleague who works with CPS geography to take a closer look at what might be causing this.
To your other question, the person-level weight (WTFINL) allows each individual sample month to be representative of the US non-institutional civilian resident population, with the annual estimate being the average across these 12 months. Note that October 2025 BMS data was never collected (see our blog post for more info). Therefore 2025 CPS monthly data include a smaller number of samples than other years, and this missing sample must be taken into account when calculating the number of samples included in pooled analyses. Setting CPSID as your PSU corrects for repeated sampling in the panel and allows you to get survey design-based confidence intervals.