Showing posts with label Sampling Techniques. Show all posts
Showing posts with label Sampling Techniques. Show all posts

Saturday, April 12, 2008

Laying the Analytics Foundation II: Designing a Questionnaire

Designing a successful questionnaire often involves balancing two dueling objectives: (1) getting all the information you need, and (2) persuading your survey participant to provide the information. Ideally, your questionnaire should not have too many questions or be too difficult to fill out; least your participant becomes frustrated and quits, or worse, provides bogus information. At the same time, if you do not get the information you need, the entire exercise becomes a waste of time and resources. Prior to designing the questionnaire and carrying out the survey, it is assumed that you have tried to obtain the information from secondary sources and failed. Gathering information by primary sources, such as a survey, is almost always more expensive than obtaining it through secondary sources.

Before creating the questionnaire, determine exactly what information you need for your analyses. Use short and simple questions to query the information. Avoid using difficult or ambiguous language. The rule of thumb (what I've been told) is that an 8th grade student should be able to read and completely understand the questions. Make a good faith effort to limit the number of questions to what you absolutely need. Provide enough space for respondents to be able to write out the answers.

The next step is arranging the questions according to some logical sequence, to not confuse the participant. If you look at the example below (Figure 2.1), the questionnaire was designed to obtain information on USPS packages transported by rail vans. We decided gathering information on the packages was not enough for our analyses; we also needed information on the rail vans and plants. Accordingly, we divided the questionnaire into three parts. The first part pertained to the rail plant, because that's what the data collector would first encounter. Once the data collector had entered the plant and filled out the necessary data, the next step was finding a rail van. Hence, the second part of the questionnaire involved gathering data on the rail van. Finally, after locating the rail van, the data collector would be able to find the mail packages and fill out the third and final part of the questionnaire.

Carefully decide the type of question to include in your questionnaire. Your questionnaire can consist of Structured questions or Unstructured questions, or both. For structured questions, you essentially know the answers of the questions and force the participant to provide a specific answer. It may be multiple choice, binary (i.e., Yes/No), or inquire for a specific type of information (i.e., Mail Code). The common predicament with structured questions is that you have to know the potential answers in advance. Rarely, you get unknown information with a purely structured questionnaire. Unstructured questions, on the other hand, provides the participant with a free form to volunteer information. Unstructured questionnaires allow you to uncover new details about your test subject. However, you might not get the information you need for your analyses. Although structured questions are great for analytics purposes, use some unstructured questions to provide some flexibility in your questionnaire (See Question #24 in the sample questionnaire).

In our sample questionnaire, you'll see that I tried to fit everything on two sides of a single page. The front page contained the questions, while the back page contained the instructions for answering the questions. This was intentionally done to simplify the job of printing and distributing the questionnaires. I provided a brief purpose so that the data collectors had a broad overview of why we were gathering the information. Each question number in the front page had a corresponding number in the back page that provided the instructions on how to collect the data. Exceptions were highlighted in bold or underlined to draw attention. I also provided hints on where the data collector could find the necessary information. If a data collector - after reading the instructions - had any questions about the survey or procedures, I provided the name and phone number of a contact person to help him/her out. For contingencies where data collectors had to record a lot more data than expected, I provided supplemental questionnaires.

On the bottom right corner of the front page, you'll see a space for processing code. The processing code is used to tag the completed questionnaire after you receive it. It is good practice to save the original paper copies. During later stages of data processing and analyses - if you ever stumble on data that makes no sense - you can use the processing code or tag number to pull up the original questionnaire and see how it was filled out.

Finally, test your questionnaire once it is completed. Give copies to people you know and ask to fill them out. This will allow you to identify and fix any wrinkles you may have overlooked.


Figure 2.1: Both sides of a sample questionnaire

Monday, November 12, 2007

Laying the Analytics Foundation I: Designing a Sample

On my blog, I have tried to focus on my experience with analytical methods and techniques applied in the corporate work environment, which may be different than what we were taught in school. Very often, companies are eager to obtain data as quickly and cheaply as possible, and do not apply the same rigor that would be expected when writing an academic research paper. In most circumstances, we need not worry as long as the obtained data has some decision making value. Nevertheless, there will be times when you may be called to defend your analyses to senior management or a regulatory agency. A question that I am asked very frequently in such situations is, “How much confidence do you have in your numbers?” In these cases, the more you are able to follow the academic concepts and theories picked up in school, the more you are able reduce your exposure to criticism. After all, nobody in your audience is going to dispute Cochran's formulas.

First, every effort should be made to obtain reliable secondary data before contemplating a sampling study. However, if the necessary data is scarce or nonexistent, and the costs for conducting a study on the population is too prohibitive, you will need to extract a sample from your population. But before you decide on the size of your sample, you would need to accurately determine your primary sampling unit. Your primary sampling unit is the smallest indivisible unit of your population that you intend to sample. All elements of what you deem to be your primary sampling unit should have identical or similar characteristics. The results of your study could vary greatly if you are not careful when determining your primary sampling unit. Unfortunately, this is also one of the most neglected aspects of sampling that I've witnessed in many corporate environments. The U.S. Postal Service, for instance, has over 500 bulk mail processing plants that vary in size. Smaller plants have 1-3 AFCS sorting machines, medium sized plants have between 5-7 and larger plants more than 10 AFCS sorting machines. The characteristics of a plant with one sorting machine is very different than that with 10 sorting machines. If you pull a random sample of 50 plants from a list of the 500 hundred plants; you could end up with all small plants, and your study would not have any representation of medium and large sized plants. Nonsensical? I have seen it happen. The results are not pretty.

The next step is ensuring the stability of your sampling frame, another aspect of sampling that's often overlooked in the corporate environment. Your sampling frame is the population list of primary sampling units from which you are to choose your sample. I have seen good statisticians analyzing the frame to ensure that the list doesn't grow or retract in subsequent periods. So what if your sampling frame fluctuates greatly from one period to another? Easy. You don't (or rather can't) do a sample. In such circumstances, if you have no other choice, you can pull your sample from the most current frame. But you shouldn't put too much 'confidence in your numbers'.

Now you are ready to pull your sample. Most corporations that I've worked at commonly use a simple rule of thumb to determine sample size, which is 10 percent of the population. However, those who want to follow a more scientific method, the formula for determining sample size is given below:



The assumption behind these formulas are that the more variation between your sample and population means, the larger should be your sample size. You can obtain your population parameters by (1) doing a pilot study, (2) using that of a previous study of a similar population, and/or (3) taking an initial sample and using the mean from that sample. There are also ways to introduce more precision to the above formulas, if needed.

Once you have determined your sample size, there are several ways to pick your sample from your sampling frame. Some of the more popular methods are:

Simple Random Sampling
Simple random sampling is, by far, the most popular method used by businesses to construct their samples. In this method, you
randomly select the sampling units - equal to the number of your sample size - from your sampling frame without any bias or restrictions. In other words, every item in your sampling frame has an equal chance (or same probability) of being included in your sample. One way to achieve this objective would be to assign every sampling unit a unique number, and then use a random number generator to select a set of numbers - equal to your sample size - from that pool of sampling units.

Stratified Sampling
In this method, you group your population into various homogeneous groups or strata based on some broad characteristic shared by the units in each group or stratum, but not by the others. You then randomly select your sampling items from each stratum, all of which should equal your sample size. In the simple random sampling method explained previously, you risk excluding certain sampling units, whose characteristics you wish to include in your analyses.
This method ensures all the characteristics of your population you wish to study are represented in your sample. In the example of the USPS bulk mail processing plants that I provided, the study was flawed because we had conducted a simple random sample. This mistake could have been averted if we had carefully evaluated our sampling methodology and conducted a stratified sample instead.

Interval Sampling
In this method,also known as Systematic Sampling, you divide your population by your sample size to obtain a factor. From your sampling frame, you then select your sampling items using an interval that equals the factor calculated. For example, if your population size is 1,000 units and your sample size is 100; your factor would be 10 (1,000/100). Using this method, you have to select every 10th item from your sampling frame to construct your sample.

Judgment Sampling
This is a biased sampling method where the choice of selecting the sampling items rests exclusively on the judgment of the analyst(s) carrying out the study. If a sample of 10 students, for instance, has to be selected from a class of 100 students; the analyst chooses the 10 students that he/she thinks represents the class best. Judgment sampling is good for quick and dirty studies where the business can't afford to spend too much money.

Using the approaches discussed above, you should now be able to extract a sample from your population.