Data Source Hybridisation: Principles, Methods and Challenges for Audience Measurement

In the face of the growing volume of data and the fragmentation of usage patterns, measurement methods are evolving. La hybridisation des sources est en train d’émerger comme une approche clé pour le cross-referencing, l’enrichissement et la fiabilisation des mesures, et pour mieux répondre aux demandes croissantes en matière de précision et de compréhension des usages.

What is data hybridisation?

Data Data Digital data Find out more source hybridisation hybridisation In statistics, this is an approach that involves mixing two data sources which differ both in nature and in level in order to create a third, richer or more detailed one. Find out more involves combining multiple complementary sources to improve the quality, accuracy, or richness of a measurement. This approach builds on statistical methods based on probabilistic probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more samples, which throughout the 20th century formed the foundation for the production of reliable quantitative information in many fields, ranging from public statistics statistics Statistics or Data Science is a field of mathematics that studies phenomena through data collection, processing, analysis, graphical representation and visualisation (Data Visualisation), as well as the interpretation of results. Find out more to epidemiology, economics, and the social sciences. 

Their strength lies in a rigorous theoretical framework, which enables the production of unbiased estimates accompanied by a measure of precision. However, in the face of growing demand for more frequent, more detailed, and more granular information, approaches relying solely on samples are now showing certain limitations

These limitations are particularly evident in the production of local statistics, where direct estimates from surveys suffer from high Variance Variance The variance of a statistical distribution is the average of the squares of the differences between the values of the distribution and their average. A standard deviation is always determined by first finding the variance, then taking the square root. Find out more and may even become unusable. To address these challenges, many statistical institutes, companies, and research organizations have developed hybrid approaches that combine survey survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more data, administrative sources, Census Census A technique for gathering information across the entire population, in contrast to a survey, which by definition only applies to a sample. Normally, a census involves an exhaustive count during which some additional socio-demographic questions are asked. Example: the national population census in France by INSEE which publishes results annually. Find out more data, and big data big data Big data (or megadata) refers to the enormous quantities of data produced by all digital activities: private or professional, human or machine. Technological advances have made it possible to use this data for purposes other than its main purpose, regardless of how it is structured. Data sets with characteristics (e.g. volume, velocity, variety, variability, veracity) which, for a particular domain problem, at a given time, cannot be efficiently processed to derive value with existing technologies and techniques. The term Big Data is commonly used in a variety of ways, for example, as the name of the scalable technology used to deal with large datasets. Find out more .

Audience Measurement Audience Measurement Quantitative study of the frequency of use of media. Find out more is no exception to this profound evolution in methods: the fragmentation of usage patterns and the proliferation of ways to access content challenge institutions to overcome the limitations of survey-based approaches by leveraging the depth of big data now available.

Image
Qu'est-ce que l'hybridation de données
Image
Pourquoi l'hybridation des données s'inscrit dans la théorie des sondages

Why is data hybridisation part of Survey theory?

Although the term “hybridisation” has appeared relatively recently in the context of Statistics Statistics Statistics or Data Science is a field of mathematics that studies phenomena through data collection, processing, analysis, graphical representation and visualisation (Data Visualisation), as well as the interpretation of results. Find out more , the concept is an old one and has its origins in the seminal work on Survey Survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more theory. The principle is simple: when auxiliary information is available, one should seek to use it.  Deville & Särndal formalized the framework for survey calibration estimators, which aim to improve the estimation of Total Population Total Population This refers to household populations as well as some communities (workers' hostels, retirement homes, university residences, prisons, etc.), the homeless, seamen and people living in mobile homes. Find out more by using known auxiliary information. 

Hybridisation Hybridisation In statistics, this is an approach that involves mixing two data sources which differ both in nature and in level in order to create a third, richer or more detailed one. Find out more methods have become more sophisticated as the nature of the data data Digital data Find out more processed and the power power Indicator for evaluating formats in the context of a media plan. This is the prioritisation of formats according to their target audience. Find out more of the tools capable of processing processing Means any operation or set of operations which is performed on personal data or on sets of personal data, whether or not by automated means, such as collection, recording, organisation, structuring, storage, adaptation or alteration, retrieval, consultation, use, disclosure by transmission, dissemination or otherwise making available, alignment or combination, restriction, erasure or destruction Find out more it have evolved, but the philosophy remains the same: to use multiple data sources to create a new, more refined or richer dataset.

The logic behind Hybrid Measurement Hybrid Measurement In statistics, this is an approach that involves mixing two data sources which differ both in nature and in level in order to create a third, richer or more detailed one. Find out more is based on four key fundamental principles.

- Complementarity: no single source single source A system that enables data collection about several topics or areas of study from a single panel of individuals or households. Find out more is sufficient on its own. The goal is to fill the gaps in one source with the strengths of another.

- Temporal and structural alignment: To combine data, it must be aligned in time (for example, covering the same period) and in its definitions (comparable units of measurement).

- Modeling: Hybridisation involves building a statistical model to combine the different data sources.

- Data source governance: A Hybrid Measurement requires a clear framework, and in particular, the highest level of transparency regarding shared data.

What are the main methods for hybridisation of data sources?

There are many statistical methods for reconciling data data Digital data Find out more from different sources, including combinations of several statistical methods. We can distinguish several major families, each addressing distinct needs.

 

  1. Statistical matching Statistical matching The statistical matching method is used to reconcile data collected from different surveys. It consists of complementing a so-called "recipient" sample, where certain information is not present, with a so-called "donor" sample relating to other individuals. To do this, the two samples must contain common information, called bridging variables. Statistical matching is a particular way of handling non-responses. Find out more : matching multiple databases using probabilistic probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more methods

Based on imputation techniques, this method involves matching multiple databases to create an enriched, consistent, and more comprehensive dataset. It is used when no single source single source A system that enables data collection about several topics or areas of study from a single panel of individuals or households. Find out more contains all the relevant variables of interest and the different data sources cannot be directly matched using a common identifier.

The goal is to reconstruct a database database All the homogeneous data collected over the course of several surveys and whose structure provides the potential for tailor-made consultations and inquiries. Find out more similar to what would have been obtained if all variables had been collected on the same individuals. The linking of databases relies on the similarity of individuals across a set of variables common to the different databases: a distance is defined between individuals from the different databases, and they are then matched based on their similarity.

The merger helps reduce the response burden* on panelists or interviewees and reconstruct a comprehensive database similar to the original data—one that is therefore easily usable in standard data analysis tools. However, variables specific to each database are never observed together. The quality of the merger therefore depends heavily on the explanatory power power Indicator for evaluating formats in the context of a media plan. This is the prioritisation of formats according to their target audience. Find out more of the common variables over the specific variables. In the absence of relevant common variables, the merger will be nearly random.


*The response burden (response burden in English) refers to the burden (time, memory effort, complexity, etc.) imposed on survey survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more respondents. It is a central concept in survey methodology because it directly influences the quality of participation and the reliability of responses.

 

Image
fusion-statistique.

2. Calibration : improving the accuracy of a sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more based on a comprehensive data data Digital data Find out more source

This approach is used when a data source is available from a sample or a panel and another source of comprehensive survey . In this case, we want to use the information from the exhaustive survey, which corresponds to a known total for the entire population , to improve the accuracy statistics , reduce the variability of results, or correct a selection bias in the sample or Panel. The approach consists of introducing additional calibration constraints into the Adjustment of the sample or panel.

The adjustment ensures consistency between the two data sources without having to modify their structure. Furthermore, it does not require access to the raw data from the exhaustive survey—which is often very voluminous—but only to the totals for the alignment variables. However, the various data sources must be perfectly comparable, which is not always the case by default. Preprocessing may therefore be necessary to ensure consistency between the measured scopes and the calculated indicators.

Image
calage

3. Profiling: Enriching a comprehensive measurement using a qualified Panel Panel A sample from which information is collected over time. It can be questioned several times at regular intervals. This method is pertinent for studying changes in behaviour. The sample can be questioned continuously: - either: each individual in the panel is asked to fill in a daily questionnaire relating to the subject of the study. This is the case for the Radio panel. - or the information is recorded and returned regularly. This is the case for the Médiamat panel. This method is pertinent for understanding behavioural patterns and their changes if the panel lasts long enough. Find out more

This method is used when a highly qualified data data Digital data Find out more source—typically derived from a sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more or panel—is available alongside another comprehensive measurement source, and the goal is to enrich the comprehensive measurement using the often highly detailed information from the other source. In fact, the exhaustive data allows us to observe uses that are still rare or occasional—ones that a sample cannot measure accurately.

The approach involves building a statistical qualification model based on the sample or Panel data and then applying it to the exhaustive data to enrich it.

It improves our understanding of emerging or rare uses without having to significantly increase Sample Size Sample Size The number of statistical units constituting a sample. The accuracy of a survey depends greatly on the sample size. Find out more . However, since comprehensive data is generally collected in silos, the explanatory variables available for modeling are relatively limited, which restricts the capacity of a model to reliably estimate profiles.

Image
calage

4. Synthetic Population Population The universe of a survey composed of basic statistical units. These units may be physical persons, households, companies, municipalities, etc. The population serves as a sampling frame for selecting a sample, and as a basis for calculation of the extrapolations from the sample. Find out more Generation: Creating a Foundation for Integrating Multiple Sources

Synthetic population generation is not unique to hybridisation hybridisation In statistics, this is an approach that involves mixing two data sources which differ both in nature and in level in order to create a third, richer or more detailed one. Find out more , but it can be used as a preliminary step in integrating different data sources. It is particularly useful when at least one data source comes from a comprehensive survey survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more . This approach was originally used for fine-grained spatial analysis. It involves constructing a comprehensive and Representative sample sample A subset of the population studied, selected in accordance with a sampling plan, and subject to the collection of information. Find out more of the population onto which survey or panel panel A sample from which information is collected over time. It can be questioned several times at regular intervals. This method is pertinent for studying changes in behaviour. The sample can be questioned continuously: - either: each individual in the panel is asked to fill in a daily questionnaire relating to the subject of the study. This is the case for the Radio panel. - or the information is recorded and returned regularly. This is the case for the Médiamat panel. This method is pertinent for understanding behavioural patterns and their changes if the panel lasts long enough. Find out more results and data . This redistribution may use deterministic deterministic The term deterministic data denotes a process for identifying an individual based on a unique identifier, which may, for example, be an identification number of a mobile device or a telephone operator subscription number. Find out more methods when a common identifier is available across the different sources, or stochastic or probabilistic probabilistic The term probabilistic data refers to a process of identifying an individual based on a probabilistic model rather than on an identifier considered to be “infallible” thanks to its uniqueness. Find out more methods otherwise. The fusion or qualification techniques detailed above can be applied to a synthetic population, as can the techniques described at Probabilisation .

Thus, synthetic population generation makes it possible to combine different sources, even when they do not pertain to universes that are strictly comparable, while facilitating the preservation of the characteristics of the original data. It also makes it possible to consider the large-scale use of highly granular individual data without encountering associated privacy issues.

The quality and validity of the synthetic population depend heavily on the quantity and granularity of the information provided by national statistics statistics Statistics or Data Science is a field of mathematics that studies phenomena through data collection, processing, analysis, graphical representation and visualisation (Data Visualisation), as well as the interpretation of results. Find out more agencies. Furthermore, as with any model-based approach, the reliability of the results depends on the explanatory power power Indicator for evaluating formats in the context of a media plan. This is the prioritisation of formats according to their target audience. Find out more of the common variables over the specific variables.

Image
calage

In summary

Hybridisation Hybridisation In statistics, this is an approach that involves mixing two data sources which differ both in nature and in level in order to create a third, richer or more detailed one. Find out more of data data Digital data Find out more sources involves combining several complementary sources to produce a more accurate, richer, or more robust measure. This approach, building on the foundational work of survey survey A statistical method that aims to produce information about a population by interviewing part of the population (sample). This word is generally used to mean a survey conducted on a sample that have been interviewed using a questionnaire or else systematically observed. Find out more theory, employs various statistical methods to address the limitations of single-source measurements, particularly in the context of declining survey response rates and the growing fragmentation of data usage.