# How can I use replicate weights to create standard errors in R?

**URL:** <https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593>\
**Category:** CPS\
**Created:** [October 2, 2018, 9:25pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593 "2018-10-02T21:25:00Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![bmccomas](https://avatars.discourse-cdn.com/v4/letter/b/b9bd4f/32.png) [@bmccomas](https://forum.ipums.org/u/bmccomas)\
**Post date:** [October 2, 2018, 9:25pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/1 "2018-10-02T21:25:00Z")

</div>

I am using R to analyze CPS data on household income and would like to use the replicate weights to create standard errors.

I am aware that such a code exists in STATA and other statistical software but am having issues translating this to R.

---

<div class="post-metadata">

**Author:** ![gfellis](https://avatars.discourse-cdn.com/v4/letter/g/3d9bf3/32.png) [@gfellis](https://forum.ipums.org/u/gfellis)\
**Post date:** [October 5, 2018, 3:09pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/2 "2018-10-05T15:09:00Z")

</div>

_ **Post edited 12/23/2020 to correct a typo** _

Note that the CPS weighting system has changed a little bit this year, and not all of our documentation has been updated. There used to be just one variable name across all supplements (WTSUPP), but now the variable name depends on which supplement you are using. Here are some examples using ASEC data (and so use ASECWT), but you can see the chart here to see what variable you should use:

[https://cps.ipums.org/cps/weights\_ren…](https://cps.ipums.org/cps/weights_renaming_2018.shtml)

And here are some examples using first the survey package, and then the srvyr (which is based on survey, but uses dplyr syntax).

library(ipumsr)

library(dplyr)

# Read data and some light data formatting

data \<- read\_ipums\_micro(“cps\_00021.xml”)

#\> Use of data from IPUMS-CPS is subject to conditions including that users should

#\> cite the data appropriately. Use command `ipums_conditions()` for more details.

data \<- data %\>%

mutate(

AGE = as.numeric(AGE),

SEX = as\_factor(SEX),

INCTOT = as.numeric(lbl\_na\_if(INCTOT, ~.val \>= 99999990))

)

## R (survey package) -----

# If not installed already: install.packages(“survey”)

library(survey)

svy \<- svrepdesign(data = data, weight = ~ASECWT, repweights = “REPWTP[0-9]+”, type = “JK1”, scale = 4/160, rscales = rep(1, 160), mse = TRUE)

# Calculate mean of INCTOT

svymean(~INCTOT, svy, na.rm = TRUE)

#\> mean SE

#\> INCTOT 42526 383.64

# Calculate a mean of INCTOT, on the subset of people aged 25-64

svy\_subset \<- subset(svy, AGE \>=25 & AGE \< 65)

svymean(~INCTOT, svy\_subset, na.rm = TRUE)

#\> mean SE

#\> INCTOT 51407 496.95

# Calculate the mean of INCTOT by SEX

svyby(~INCTOT, ~SEX, svy, svymean, na.rm = TRUE)

#\> SEX INCTOT se

#\> Male Male 53196.41 637.2199

#\> Female Female 32456.95 325.3275

# R (srvyr package - uses dplyr-like syntax) -----

# If not installed already: install.packages(“srvyr”)

library(srvyr)

svy \<- as\_survey(data, weight = ASECWT, repweights = matches(“REPWTP[0-9]+”), type = “JK1”, scale = 4/160, rscales = rep(1, 160), mse = TRUE)

# Calculate mean of INCTOT

svy %\>%

summarize(mn = survey\_mean(INCTOT, na.rm = TRUE))

#\> # A tibble: 1 x 2

#\> mn mn\_se

#\> \<dbl\> \<dbl\>

#\> 1 42526. 384.

# Calculate a mean of INCTOT, on the subset of people aged 25-64

svy %\>%

filter(AGE \>= 25 & AGE \< 65) %\>%

summarize(mn = survey\_mean(INCTOT, na.rm = TRUE))

#\> # A tibble: 1 x 2

#\> mn mn\_se

#\> \<dbl\> \<dbl\>

#\> 1 51407. 497.

# Calculate the mean of INCTOT by SEX

svy %\>%

group\_by(SEX) %\>%

summarize(mn = survey\_mean(INCTOT, na.rm = TRUE))

#\> # A tibble: 2 x 3

#\> SEX mn mn\_se

#\> \<fct\> \<dbl\> \<dbl\>

#\> 1 Male 53196. 637.

#\> 2 Female 32457. 325.

---

<div class="post-metadata">

**Author:** ![Jamaal\_Green](https://avatars.discourse-cdn.com/v4/letter/j/3e96dc/32.png) [@Jamaal\_Green](https://forum.ipums.org/u/Jamaal_Green)\
**Post date:** [September 23, 2019, 5:16pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/3 "2019-09-23T17:16:38Z")

</div>

I’m working with the ASEC file to estimate TANF participation. I found this forum for specifying the survey design in R, but when looking here on Anthony Damico’s site on complex survey design: [http://asdfree.com/current-population-survey-basic-monthly-cpsbasic.html](http://asdfree.com/current-population-survey-basic-monthly-cpsbasic.html) the type and row parameters are different. Is this a mistake on Damico’s part? Should I follow this approach? Additionally are there are any reference tables to make sure estimates are correct? The ACS PUMS provide state level estimates to check to make sure your survey design is correct. Is this available for CPS?

---

<div class="post-metadata">

**Author:** ![JeffBloem](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.ipums.org/jeffbloem/32/40_2.png) [@JeffBloem](https://forum.ipums.org/u/JeffBloem)\
**Post date:** [September 24, 2019, 2:23pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/4 "2019-09-24T14:23:01Z")

</div>

One key difference between the approach detailed on the linked webpage and the approach noted on the IPUMS Forum above, is that the later integrates the replicate weights provided with the CPS data while the former does not. These are two distinct ways of calculating standard errors. More information about replicate weights is [available here](https://cps.ipums.org/cps/repwt.shtml). Regarding any previously calculated statistics using the CPS data, I’ll direct you to the [BLS website](https://www.bls.gov/cps/tables.htm). They list a number of tables with published statistics using the CPS data.

---

<div class="post-metadata">

**Author:** ![Jamaal\_Green](https://avatars.discourse-cdn.com/v4/letter/j/3e96dc/32.png) [@Jamaal\_Green](https://forum.ipums.org/u/Jamaal_Green)\
**Post date:** [September 24, 2019, 6:43pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/5 "2019-09-24T18:43:32Z")

</div>

Excuse me. I meant to send the ASEC link that does include the replicate weights. I just want to make sure the approach linked in the above post is appropriate for calculating proper standard errors with the survey package? And I was curious if there are any reference files to double check the estimates like one can do with the ACS values?

---

<div class="post-metadata">

**Author:** ![JeffBloem](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.ipums.org/jeffbloem/32/40_2.png) [@JeffBloem](https://forum.ipums.org/u/JeffBloem)\
**Post date:** [September 25, 2019, 3:54pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/6 "2019-09-25T15:54:19Z")

</div>

Yes, the R packages [survey](https://cran.r-project.org/web/packages/survey/survey.pdf) and [srvyr](https://cran.r-project.org/web/packages/srvyr/index.html) can help facilitate specification of sample design with the ipumsr package. Regarding any previously calculated statistics using the CPS data, I’ll direct you to the [BLS website](https://www.bls.gov/cps/tables.htm). They list a number of tables with published statistics using the CPS data.

---

<div class="post-metadata">

**Author:** ![Philippe\_Lemoine](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.ipums.org/philippe_lemoine/32/300_2.png) [@Philippe\_Lemoine](https://forum.ipums.org/u/Philippe_Lemoine)\
**Post date:** [December 12, 2019, 4:07pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/7 "2019-12-12T16:07:14Z")

</div>

This is very helpful, but in the call to svrepdesign/as\_survey, shouldn’t the value of scale be 4/160 instead of 4/60? According to [IPUMS CPS](https://cps.ipums.org/cps/repwt.shtml), the multiplier in front of the sum of squared deviations in the formula for the standard error is 4/160, not 4/60. Am I missing something?

---

<div class="post-metadata">

**Author:** ![Molly\_Richard](https://avatars.discourse-cdn.com/v4/letter/m/3be4f8/32.png) [@Molly\_Richard](https://forum.ipums.org/u/Molly_Richard)\
**Post date:** [December 21, 2020, 10:20pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/8 "2020-12-21T22:20:30Z")

</div>

Do you have any similar sample code for using replicate waits within these packages for IPUMS USA (ACS files)?

---

<div class="post-metadata">

**Author:** ![Matthew\_Bombyk](https://avatars.discourse-cdn.com/v4/letter/m/34f0e0/32.png) [@Matthew\_Bombyk](https://forum.ipums.org/u/Matthew_Bombyk)\
**Post date:** [December 23, 2020, 7:21pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/9 "2020-12-23T19:21:12Z")

</div>

You’re correct that there was a typo in the earlier post. I have edited that post to use the correct denominator of 160 in the survey design specification step. Thanks for pointing this out, and sorry for the year-late follow up, your post must have slipped through the cracks!

---

<div class="post-metadata">

**Author:** ![Matthew\_Bombyk](https://avatars.discourse-cdn.com/v4/letter/m/34f0e0/32.png) [@Matthew\_Bombyk](https://forum.ipums.org/u/Matthew_Bombyk)\
**Post date:** [December 23, 2020, 7:37pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/10 "2020-12-23T19:37:38Z")

</div>

The code for ACS samples should be nearly the same. The IPUMS USA page on replicate weights gives details on the calculations: [https://usa.ipums.org/usa/repwt.shtml](https://usa.ipums.org/usa/repwt.shtml)

I’ll refer you to the earlier post in this thread by @gfellis giving example code for using replicate weights with ASEC data. Apart from specific variables used in analysis, the code will work more or less unchanged for ACS, with a few modifications to the survey design specification. Below I’ve highlighted the things you would need to change, depending on whether you’re using -survey- or -srvyr- packages:

## Using -survey-

### ASEC:

```
svy <- svrepdesign(data = data, weight = ~ASECWT, repweights = “REPWTP[0-9]+”, 
      type = “JK1”, scale = 4/160, rscales = rep(1, 160), mse = TRUE)
```

### ACS:

```
svy <- svrepdesign(data = data, weight = ~ **PERWT** , repweights = “REPWTP[0-9]+”,
      type = “JK1”, scale = 4/ **80** , rscales = rep(1, **80** ), mse = TRUE)
```

## Using -srvyr-

### ASEC:

```
svy <- as_survey(data, weight = ASECWT, repweights = matches(“REPWTP[0-9]+”), 
      type = “JK1”, scale = 4/160, rscales = rep(1, 160), mse = TRUE)
```

### ACS:

```
svy <- as_survey(data, weight = **PERWT** , repweights = matches(“REPWTP[0-9]+”), 
      type = “JK1”, scale = 4/ **80** , rscales = rep(1, **80** ), mse = TRUE)
```

---

<div class="post-metadata">

**Author:** ![Cela](https://avatars.discourse-cdn.com/v4/letter/c/f19dbf/32.png) [@Cela](https://forum.ipums.org/u/Cela)\
**Post date:** [October 21, 2021, 9:34pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/12 "2021-10-21T21:34:56Z")

</div>

@Matthew_Bombyk (or other IPUMS Staff :D)

If I can revive an old thread… Can I confirm that I am repurposing the code correctly for weighting at the household level? For instance, if my goal is to aggregate data at the household level for CPS ASEC, I should be using:

-survey package-

```auto
svy <- svrepdesign(data = data, weight = ~ ASECWTH, repweights = “ REPWT[0-9]+”, 
      type = “JK1”, scale = 4/160, rscales = rep(1, 160), mse = TRUE)

```

Of note is using REPWT over REPWTP, and using ASECWTH over ASECWT. My understanding is we don’t need to change the parameters in type, scale, or scales?

Thanks so much!

---

<div class="post-metadata">

**Author:** ![Matthew\_Bombyk](https://avatars.discourse-cdn.com/v4/letter/m/34f0e0/32.png) [@Matthew\_Bombyk](https://forum.ipums.org/u/Matthew_Bombyk)\
**Post date:** [October 22, 2021, 1:53pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/13 "2021-10-22T13:53:26Z")

</div>

That looks right to me.

---

<div class="post-metadata">

**Author:** ![Cela](https://avatars.discourse-cdn.com/v4/letter/c/f19dbf/32.png) [@Cela](https://forum.ipums.org/u/Cela)\
**Post date:** [October 22, 2021, 2:01pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/14 "2021-10-22T14:01:58Z")

</div>

Awesome. Thanks @Matthew_Bombyk . 🙂

---

<div class="post-metadata">

**Author:** ![Jana\_Sessler](https://avatars.discourse-cdn.com/v4/letter/j/f19dbf/32.png) [@Jana\_Sessler](https://forum.ipums.org/u/Jana_Sessler)\
**Post date:** [January 4, 2022, 2:37pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/15 "2022-01-04T14:37:07Z")

</div>

I tried to follow exactly the steps explained above, but always get this error:

Error in if (combined.weights & probably.not.combined.weights) warning(paste(“Data do not look like combined weights: mean replication weight is”, :  
missing value where TRUE/FALSE needed

does anybody have an idea what’s wrong ?

Many thanks in advance !

---

<div class="post-metadata">

**Author:** ![Matthew\_Bombyk](https://avatars.discourse-cdn.com/v4/letter/m/34f0e0/32.png) [@Matthew\_Bombyk](https://forum.ipums.org/u/Matthew_Bombyk)\
**Post date:** [January 7, 2022, 11:10pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/16 "2022-01-07T23:10:27Z")

</div>

This may be a problem with your -survey- package. I recommend installing the latest version of the survey package. You can type:

```auto
install.packages("survey")

```

If you’re using RStudio, try this in base R first and see if the problem is fixed.

---

<div class="post-metadata">

**Author:** ![Katie\_Savin](https://avatars.discourse-cdn.com/v4/letter/k/f1d935/32.png) [@Katie\_Savin](https://forum.ipums.org/u/Katie_Savin)\
**Post date:** [May 18, 2023, 2:57pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/18 "2023-05-18T14:57:51Z")

</div>

Thanks for this code and instruction, it is very helpful.  
I am trying to use CPS data for descriptive statistics on SSI recipients in California and their rates of SNAP receipt in particular.  
I am trying to use the ASECWT for my code in Rstudio, however when I use the syntax you provided after installing the survey package, I am receiving this error messages I can’t clear:  
Error in UseMethod(“as\_survey”) :  
no applicable method for ‘as\_survey’ applied to an object of class “function”  
Any advice?  
Thanks in advance,  
Katie

---

<div class="post-metadata">

**Author:** ![Ivan\_Strahof](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.ipums.org/ivan_strahof/32/626_2.png) [@Ivan\_Strahof](https://forum.ipums.org/u/Ivan_Strahof)\
**Post date:** [May 24, 2023, 3:02pm UTC](https://forum.ipums.org/t/how-can-i-use-replicate-weights-to-create-standard-errors-in-r/2593/19 "2023-05-24T15:02:01Z")

</div>

It appears that RStudio is trying to apply the UseMethod function to the as\_survey function, however this is not possible since as\_survey is also a function. Would you be able to share where you are getting code to use replicate weights? IPUMS CPS provides this code on [the replicate weights user guide](https://cps.ipums.org/cps/repwt.shtml). In RStudio, you first want to install the srvyr package:

install.packages(“srvyr”)  
library(“srvyr”)

Then, run the as\_survey function:

svy ← as\_survey(data, weight = ASECWT, repweights = matches(“REPWTP[1-160]+”))
