Sampling: Theory and Practice for Representing a Population

A representative sample reflects a population in order to produce reliable results. Historically defined as a scaled-down model of the population, it is now based on a probabilistic principle: each individual must have a known probability of being selected. In practice, two main approaches coexist: probabilistic methods, which have a solid theoretical framework, and non-probabilistic methods, including Quota Sampling, which are widely used in France by private survey firms, provided they are implemented rigorously to limit bias.

Why has the concept of a representative sample changed?

The question question An inquiry posed in the same way to all participants in the same survey, with the goal of obtaining answers of the same nature. We distinguish closed questions whose responses are chosen from the options provided (scales, yes / no, classification options, etc.), from open-ended questions for which no answer categories are offered. In the latter case, the spontaneous answer of the interviewee is gathered. Find out more of representativeness has been at the heart of surveys since the late 19th century. Initially, a sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more was considered representative representative A sample is said to be representative when it has been generated via a representative draw, meaning that each individual in the population had a non-zero probability of being selected. Representativeness relates to the sample and not to a unit of the population. It is not synonymous with a proportional sample or a "reduced model" of the population studied. In survey theory, it has been demonstrated that an optimal sampling strategy consists of organising the sample according to the variables that are most correlated to the subject of the study. Find out more when it accurately reflected the structure of the population population The universe of a survey composed of basic statistical units. These units may be physical persons, households, companies, municipalities, etc. The population serves as a sampling frame for selecting a sample, and as a basis for calculation of the extrapolations from the sample. Find out more , with each category present in the same proportions. This view was based on an intuitive logic: that of a scale model.

A major turning point came in 1934 with Jerzy Neyman’s article “On Two Different Aspects of the Representative Method: The Method of Stratified Sampling and the Method of Purposive Sampling.” Representativeness is no longer defined as a descriptive resemblance, but as a property of the selection process, based on the calculation of probabilities, thereby providing surveys with a rigorous mathematical foundation.

The definition of representativeness then became: each unit of the population has a known, non-zero probability probability The probability of an event is a real number between 0 and 1 that reflects the chance of that event occurring. An impossible event has a probability of 0; a certain event has a probability equal to 1. The theory and calculation of probabilities are the basis of mathematical statistics and polls. Find out more of being included in the sample.

Image
Etude livre bibliothèque
Image
Comment fonctionnent les méthodes probabilistes ?

How do probabilistic methods work?

Probabilistic Probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more methods are based on probability probability The probability of an event is a real number between 0 and 1 that reflects the chance of that event occurring. An impossible event has a probability of 0; a certain event has a probability equal to 1. The theory and calculation of probabilities are the basis of mathematical statistics and polls. Find out more theory and provide a theoretical framework for sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more design.

They assume that, for each individual, it is possible to determine the probability of being selected. This requires an exhaustive Sampling Frame Sampling Frame List of individuals or statistical units forming a population within which one wishes to draw a sample, and from which the information required will be gathered. For example, the entire group of telephone subscribers may constitute the sampling frame of certain studies. Find out more as well as a Sampling Plan, which allows the probability of each unit being selected for the sample to be calculated.

In this framework, the validity of the results rests on a strong assumption: all selected units must respond. In practice, however, non-response non-response People solicited to participate in a statistical survey do not always respond - in this case they are called non-responses. “Total non-response” cases (when a person contacted for a survey does not respond to any question at all) are distinguished from “partial non-response” cases (when the interviewee responds to some questions, but not others). Find out more is common. To maintain theoretical validity, it must therefore be demonstrated that this non-response is random, meaning that it is not related to the variables under study study Any research done across a population, line of business, or products (programme content, programme schedules, etc.) A study relies most often on the data coming out of a survey (or surveys). Find out more . If this condition is not met, the absence of bias bias The bias of a statistical result is the difference between the result obtained and the exact value that one is seeking. There are three distinct types of bias: sampling bias (for example: using an unsuitable sampling frame), observational bias (for example: poor wording of a question or answer grid), and estimate bias (for example: not fully taking into account the categories for drawing the sample). Find out more can no longer be guaranteed.

Why are non-probabilistic methods used?

In the absence of a comprehensive Sampling Frame Sampling Frame List of individuals or statistical units forming a population within which one wishes to draw a sample, and from which the information required will be gathered. For example, the entire group of telephone subscribers may constitute the sampling frame of certain studies. Find out more , probabilistic probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more methods become impossible to implement. Non-probabilistic methods therefore offer a practical alternative.

They rely on an empirical approach aimed at approximating a random selection, even though no selection probability probability The probability of an event is a real number between 0 and 1 that reflects the chance of that event occurring. An impossible event has a probability of 0; a certain event has a probability equal to 1. The theory and calculation of probabilities are the basis of mathematical statistics and polls. Find out more is formally defined. In this context, the role of the interviewer becomes central: they follow precise instructions to limit selection bias bias The bias of a statistical result is the difference between the result obtained and the exact value that one is seeking. There are three distinct types of bias: sampling bias (for example: using an unsuitable sampling frame), observational bias (for example: poor wording of a question or answer grid), and estimate bias (for example: not fully taking into account the categories for drawing the sample). Find out more .

These methods offer significant practical advantages. They are simpler to implement, less costly, and allow for rapid sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more collection.

On the other hand, they do not guarantee, in theory, that every individual has a nonzero and known probability of being selected.

Their validity therefore depends largely on the rigor of the process implemented, particularly through repeated contact contact In media planning, a contact refers to an individual's exposure to an advertisement. In radio parlance, this is Opportunity to Hear (OTH) and for TV, Internet, press, cinema and poster advertisements, it is called Opportunity to See (OTS). Find out more attempts, the neutrality of the Pitch, and, more generally, the ability to avoid selection bias among individuals.

Image
Pourquoi les méthodes non probabilistes sont-elles utilisées ?
Image
La méthode des quotas

Quota Sampling: An Operational Approach to Representativeness

Quota Sampling Quota Sampling A sampling method that involves taking a "considered" selection of the sample according to characteristics (gender, age, region, etc.) of the units interviewed, in order to adhere to a predefined structure. Find out more is the most widely used non-probabilistic method in France by private survey survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more firms.

The principle behind this method is to replicate the structure of the survey’s reference population population The universe of a survey composed of basic statistical units. These units may be physical persons, households, companies, municipalities, etc. The population serves as a sampling frame for selecting a sample, and as a basis for calculation of the extrapolations from the sample. Find out more based on certain criteria (often sociodemographic and geographic). The sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more is thus constructed to reflect known proportions, such as the male-to-female distribution distribution This refers to the distribution of a variable; “structure” is also used with this meaning. In statistics, for each x value of a quantitative variable, the distribution function gives the proportion of individuals having a variable value that is less than or equal to x. Find out more .

However, this approach does not allow for the determination of a selection probability probability The probability of an event is a real number between 0 and 1 that reflects the chance of that event occurring. An impossible event has a probability of 0; a certain event has a probability equal to 1. The theory and calculation of probabilities are the basis of mathematical statistics and polls. Find out more for each individual. Since the Sampling Plan is unknown, it does not fall within the Probabilistic Probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more theoretical framework.

Despite this, the method is based on solid foundations.

If the quota quota The share allocated to a category of individuals when forming a sample. Find out more criteria accurately account for the variables under study study Any research done across a population, line of business, or products (programme content, programme schedules, etc.) A study relies most often on the data coming out of a survey (or surveys). Find out more , the estimates can be effective—that is, with limited bias bias The bias of a statistical result is the difference between the result obtained and the exact value that one is seeking. There are three distinct types of bias: sampling bias (for example: using an unsuitable sampling frame), observational bias (for example: poor wording of a question or answer grid), and estimate bias (for example: not fully taking into account the categories for drawing the sample). Find out more . Furthermore, this method allows for the incorporation of useful auxiliary information as early as the sample design phase.

In general, a quota survey produces more precise results than a probabilistic survey, in the sense of lower Variance Variance The variance of a statistical distribution is the average of the squares of the differences between the values of the distribution and their average. A standard deviation is always determined by first finding the variance, then taking the square root. Find out more .

This is why, for small samples (up to 1,000 or 2,000 respondents), it is often considered that Quota Sampling offers the best trade-off between bias and Variance.

However, Quota Sampling is effective only if it is applied rigorously.
To do so, it is necessary to choose quotas appropriate to the Survey’s subject matter, select respondents as randomly as possible, and ensure that individuals from the entire Population are surveyed.

Otherwise, the sample may be affected by selection bias, as certain categories of people may be over- or underrepresented, or even entirely absent. This risk also exists in probabilistic surveys, but it is generally better controlled through random selection.

Key Points on Representative Samples and Quota Sampling

Outside of official statistics statistics Statistics or Data Science is a field of mathematics that studies phenomena through data collection, processing, analysis, graphical representation and visualisation (Data Visualisation), as well as the interpretation of results. Find out more , research institutes often use non-probabilistic methods, particularly Quota Sampling Quota Sampling A sampling method that involves taking a "considered" selection of the sample according to characteristics (gender, age, region, etc.) of the units interviewed, in order to adhere to a predefined structure. Find out more . Although these approaches do not fall within the theoretical framework of probabilistic probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more surveys, they can produce reliable results when the quota quota The share allocated to a category of individuals when forming a sample. Find out more criteria are relevant, the field field All the operations that structure the work of interviewers: description of the interviewees, instructions for choosing the interviewees, quotas, administering the questionnaire, quality controls, etc. The entire physical system (rooms, personal computers, telephones, etc.) used for telephone surveys, and the related organisation is called the “telephone field”. Find out more is rigorous, and selection biases are controlled.