Showing posts with label Ask CRRC. Show all posts
Showing posts with label Ask CRRC. Show all posts

Wednesday, June 01, 2011

Ask CRRC | Population Sizes and Sample Sizes

Q: The 2010 Caucasus Barometer includes about 2,000 completed interviews in each country: Armenia, Azerbaijan, and Georgia. However, the three countries vary in size; the population of Armenia is just under 3 million, Georgia has a population of about 4.6 million, and the population of Azerbaijan is about 8.4 million (according to the CIA World Factbook). How can the same or a similar sample size be appropriate for each country?


A: Great question! Contrary to popular belief, the total population size has little effect on the necessary sample size. Necessary sample size is more dependent on the amount of variability between members of a population. Only one person would need to be sampled if there were no variability in a population and every member would give identical answers.


Let’s use a physical example to make this more clear:



The two populations above have the same average height, but the members of Population B have much more variability in height than the members of Population A. Thus, if you were sampling Population B you would need a much larger sample size in order to reach the same level of certainty about the population’s average height than you were if you were sampling Population A. In short, the greater the amount of variability in the population, the larger your sample size needs to be in order to capture that variability.

Other issues that affect sample size include how accurate you want conclusions drawn from the sample to be and how certain you want those conclusions to be. In making a precise statement, you could say, for example, that “from the 2010 Caucasus Barometer, our best estimate of the proportion of Tbilisi residents who have travelled to another country is 18.5%, and we are 95% sure that the true value is between 15.5% and 21.5%.” Technically speaking, 95% is our confidence level and our margin of error is 3%. Therefore, we are 95% sure that the true value lies within the range of our best estimate plus or minus 3%. To increase your level of confidence or reduce the margin of error, you would need a larger sample size -- and more money to pay for the extra interviews.

Here is one more thing worth knowing about sampling. Imagine a country of 5 million people, and a village of 500 inhabitants (both with the same amount of variability). Let’s say you require a sample of 200 from the country to reach a 95% level of confidence and a 3% margin of error. How many inhabitants of the village should be sampled to reach that same level of confidence and the same margin of error? Take a guess.

Done? The number is surprisingly high: we still need to sample one hundred and forty three inhabitants from the village. So while the country is 10,000 times the size of the village, it only requires an extra 57 people in the sample to achieve the same margin of error at the same level of confidence. In other words, one entirely counter-intuitive aspect about sampling is that small populations may still require a large proportion to be sampled to get representative findings.

In summary, while population size is one of the four factors that influence the necessary sample size for any survey (and even more factors have to be considered for complex surveys like the CB), its influence is relatively negligible.

Do you have further questions? Write a comment and let us know.

Friday, March 25, 2011

Ask CRRC | Sampling Weights I

Q: In the posting on representativeness, you said that every member of the population must have some chance of being selected for the sample. In the next posting about sample size, your Rustavi example had every member of the population with an equal chance of being selected. What if everyone has a chance, but not an equal chance? In this case, is it possible to make a sample be representative of the population?

A: This is very important question! The short answer is yes—the sample can be representative of the population, but you need to do a little extra work. Let’s use a simple example:

Suppose we are interested in comparing the experiences of male and female students in an engineering program. The program has 800 men and 200 women. If we randomly select a sample of 200 students (20% of the total student population in the engineering program), then we should expect only about 40 women in our sample. Suppose we randomly select 100 men and then randomly select 100 women. This means that every man has an equal chance of being selected for the sample and every woman has an equal chance of being selected, but every student did not. If we want to use the responses of the men to say something only about male students or the responses of women to say something only about female students, then we can do this using some simple formulas from statistics. However, what if we want to use of all of the information that we have to say something about the entire population of students?

In this case, different members of the population have different chances of being selected. Every man has a 1 in 8 chance of being selected, while every woman has a 1 in 2 chance. We can turn this around and say that every man who is interviewed represents 8 people including himself and every woman who is interviewed represents 2 people including herself. This is what is known as a sampling weight – every man in the sample has a sampling weight of 8, while every woman in the sample has a sampling weight of 2:

We need to utilize sampling weights when making estimates about an entire population. This means that we need to use different statistical formulas than the simple ones used above. We also need to use a computer program that has built-in functions to make estimates about populations using data with sampling weights (e.g., SPSS for estimates or STATA for estimates and associated margins of error). As long as we do that, then our sample is still representative of our population even though every member of the population did not have the same chance of being selected for an interview.

Wednesday, March 02, 2011

Ask CRRC | Sample Size

Q: In the last posting you said that in order for the sample to be representative of the entire population, every member of the population had to have some chance of being selected for the sample. However, you didn’t say anything about sample size. Doesn’t sample size matter?

A: As long as the sample size is not tiny, then the sample can be representative of the population – having 200 respondents or 2,000 respondents does not make a difference in whether you can call the sample representative of the population. Where sample size does make a difference is in how accurate your conclusions about the population of interest will be. Let’s explain what that means with an example:

Suppose we are interested in the population of voters in Rustavi and that we are interested in the proportion of residents who find the availability of gas to be an important local issue. We take a list of the 98,492 registered voters in Rustavi and randomly select a sample for interview. Now, let’s imagine two different scenarios: In the first, we randomly select 200 respondents and interview them. In the second, we randomly select 2,000 respondents and interview them. Now, imagine that in the first scenario, 64 respondents mentioned the availability of gas as an important local issue and 138 did not. Imagine that in the second scenario 640 respondents mentioned it and 1,380 did not. Because 64/200=0.32 and 640/2,000=0.32, in both scenarios exactly 32% of the respondents said that the availability of gas is an important local issue.

Both of these samples are representative of the population of Rustavi because every resident had a chance to be in the sample. In both cases, our best estimate of the proportion of Rustavi residents who consider the availability of gas to be a major issue is the same. This is the proportion that we encountered in each sample: 32%.

However, the two different sample sizes allow us to say two different things about the greater population of Rustavi. This is because in general the larger the sample size, the smaller the margin of error. The margin of error tells us how wide the range is within which we are sure that the true value for the entire population lies. For example, in the first scenario, using statistical formulas we can calculate that there is a 95% chance that the proportion of the entire population of 98,492 registered voters that considers the availability of gas to be an important issue is between 25.5% and 38.5%. However, in the second scenario, our calculations will tell us that we can be 95% confident that the proportion is between 30% and 34%.

That is, in the first scenario, we were 95% confident that the proportion was between 32% - 6.5% and 32% + 6.5%. In the second scenario, we were 95% confident that the proportion was between 32% - 2% and 32% + 2%. In other words, in the first scenario, the margin of error is 6.5% and in second scenario the margin of error is 2%. To conclude, different sample sizes can still be representative of a population. However, the margin of error varies with respect to the sample size and can tell us how accurate conclusions are about the population of interest.

Saturday, February 12, 2011

Ask CRRC | Representative Sample

Q: When conducting a survey, how do you select a sample that is representative of an entire population?

A: In order for a sample to be representative of an entire population, every member of the population must have some chance of being randomly selected. In reality, there are segments of a population that can and cannot be interviewed. Therefore, we need to understand the nature of the population for which each survey is representative.

Take the Caucasus Barometer (CB) as an example. First, we randomly select voting precincts from a list of all voting precincts that contain members of the population. Thus, every precinct has a chance of being selected.


The beginning of a long list of voting precincts from which precincts are randomly selected for sampling.

Second, CRRC randomly selects households within each of the selected voting precincts. Then, interviewers conduct a “random walk” in order to randomly select households. This random walk gives each household a chance of being selected for an interview.


This CRRC interviewer has a map of households in a selected voting precinct to assist her in her “random walk” household selection. Photo by Paul Stephens.

Third, an adult household member is randomly selected for an interview within each randomly selected household. Interviewers make a list of all adult (18 years and older) household members and randomly select one of those members for an interview. The interviewer uses a kind of random number table called a “Kish table” to randomly select one of those household members to interview. Using these three steps above, each member of the population has a chance of being selected for an interview.



This household consists of a 25 year old man, a 61 year old woman and a 24 year old woman. The 25 year old man has been randomly selected for interview.

As in any country, logistical realities mean that some segments of the country’s adult population do not have a chance of being sampled. For example, some voting precincts could not be sampled even if they were randomly selected (e.g., special voting precincts for military personnel). Also, some people might not be able to be surveyed even if they were randomly selected. This includes people who do not speak the language in which the survey is conducted or those who are not physically able to be interviewed. Other excluded groups of the population include people in prisons or hospitals. The impact of losing some of these groups is relatively little since such groups are usually so small that they are within the margin of error.

By understanding which groups of the population can and cannot be included in the sample, CRRC takes all of the steps above to ensure that samples are representative of the entire population. In addition, CRRC prints questionnaires in minority languages and recruits interviewers who speak those languages so that the CB can be described as representative of the adult population of the Republic of Georgia.

Wednesday, October 06, 2010

Ask CRRC | Survey vs Census

Q: What’s the difference between a survey and a census?

A: In short – census takers attempt to contact all members of a population, while surveyors select a sample of people from the population and use the responses of those people to draw conclusions about the proportions of people in the greater population holding various opinions.
There are many advantages to conducting a survey rather than a census, and here are some key examples: 
Firstly, results can be produced much more quickly with a survey than with a census. Imagine that you want to gauge Georgian political opinion just before an election. How much time would it take you to interview every adult Georgian? How many interviewers would you need to train in order to conduct all of the interviews in the month before the elections? A political opinion survey conducted by CRRC immediately before the May 2010 elections employed 100 interviewers to attempt 3,284 interviews. The adult population of Georgia is approximately 3.5 million persons, meaning that a census would require roughly 106,577 interviewers.

Secondly, the far smaller number of interviews conducted in a survey means that you can allocate more of your resources towards ensuring quality. Would you want to spend your money providing a competitive salary to 100 quality interviewers and training them well, or would you rather spend your money paying a minimal wage to 106,577 interviewers and training them insufficiently? In short, a survey allows for more resources to be allocated to other aspects of the process. CRRC invests resources in ensuring quality throughout the survey process, including performing checks to ensure interviewer integrity and entering the data from each interview into the database twice in order to catch data entry errors.

Thirdly, with a survey you can spend your time and money making sure that you collect information on all members of your sample. You can revisit houses where you didn’t find people at home the first time. This is important because certain parts of the population are harder to reach than others. For example, women, older people, and unemployed people are all more likely to be at home when an interviewer visits. These demographic groups may have different answers to survey questions than their counterparts, and a sample that over-represents them may be biased. CRRC interviewers randomly select a respondent in each selected household. If that household member isn’t home, the interviewer schedules a re-visit to the household, and makes a total of three visits to attempt to find that household member at home. This ensures that the sample contains a representative mix of men and women, young and old, employed and unemployed.

The reasons listed above are all interrelated – time, money, and manpower are always limited, and conducting a survey allows an organization to gain as much information as possible for the resources that they expend. However, in some cases the situation is even more extreme – in some cases, the object of measurement has to be destroyed in order to be measured. Think of how a manufacturer measures the number of calories per cookie: they burn a cookie in a machine called a bomb calorimeter, shown in the figure above. The number of calories in the cookie is a measure of how much heat the cookie produces when burned. Not every cookie is identical, so manufacturers take a sample of cookies. They burn each one in a bomb calorimeter, and report the average number of calories generated per cookie in the sample. If they performed a census on the population of cookies and burned every cookie, there would be nothing left to sell.