(2025) Enquête à haute fréquence Haïti 2025 : plan de sondage et pondération
Resume — Comment l'enquête à haute fréquence 2025 en Haïti a été échantillonnée et pondérée. L'échantillon a été tiré par composition aléatoire de numéros parmi toutes les lignes mobiles actives ; la note expose la stratification, les taux de réponse et les pondérations de post-stratification qui ramènent un échantillon téléphonique à la population. Sa portée dépasse cette enquête : l'insécurité limitant le travail de terrain en face-à-face, les enquêtes téléphoniques portent une grande part des données récentes sur les ménages haïtiens.
Constats Cles
- L'échantillon a été constitué par composition aléatoire de numéros couvrant tous les numéros de téléphone mobile actifs au moment de la sélection, en mai 2025.
- Les estimations représentent donc les ménages disposant d'au moins un téléphone mobile, et les individus de 18 ans et plus qui y vivent : c'est la limite de couverture que tout utilisateur des données doit garder à l'esprit.
- Dérive les probabilités de sélection des ménages et des individus des probabilités d'inclusion des numéros de mobile.
- A recueilli des données sur un enfant de 0 à 5 ans ou un adolescent de 6 à 17 ans tiré au hasard dans chaque ménage interrogé : les estimations au niveau de l'enfant reposent donc sur un tirage aléatoire intra-ménage.
Texte Integral du Document
Texte extrait du document original pour l'indexation.
Haiti High Frequency Survey (HFS) - 2025
Sampling Design and Weighting
June 23th 2025
Haiti High Frequency Survey (HFS) - 2025
Sampling Design and Weighting
The sample for the Haiti HFS-2025 was generated using a Random Digit Dialing (RDD) methodology,
which covered all active cell phone numbers at the time of sample selection (May 2025).
The survey estimates represent households with at least one cell phone, and individuals aged 18 or older
residing in those households.
1. Sampling design
The RDD methodology generates all possible phone numbers under the national numbering plan and
draws a random sample. This approach ensures full coverage of the population with a mobile phone.1
A large first-phase sample was drawn from the complete number frame. An automated screening process
identified active numbers. These active numbers were matched against business registries (e.g., yellow
pages and websites) to remove business numbers, which are ineligible for this survey.
A smaller second-phase sample2 was then selected from the list of active numbers identified in the firstphase sample and handed over to the field team for contact and interviews.3
2. Weighting
This survey includes two sample units: households and individuals. Sampling weights were computed for
each unit and should be used according to the estimate of interest. The weighting process follows five
steps:
1.
2.
3.
4.
5.
Estimation of cell phone inclusion probabilities.
Computation of design weights for households and individuals.
Adjustment for nonresponse.
Calibration using external population data (adjusted for national phone coverage).
Trimming and recalibration of weights.
1
Given that the survey used a sampling frame of telephone numbers, results represent the population with at least
one active phone and exclude the population with no phone.
2
Note that the selection of phone numbers involves two sampling phases, and not two sampling stages. The survey
involves only one sampling stage.
3
Furthermore, the second-phase sample was delivered in batches to the country teams during fieldwork. Delivering
large lists of numbers could have facilitated the “misuse” of the sample by easily replacing non-answering numbers,
raising nonresponse rates and potentially increasing nonresponse biases.
2
Step 1: Cell Phone Inclusion Probabilities
A first-phase sample was selected using simple random sampling without replacement. The selected
numbers were then screened and classified into active and inactive.
The first-phase inclusion probabilities of cell phone numbers are4
𝐶(1)𝑖 =
𝐶
𝑛(1)
𝐶
𝑁(1)
=
𝐶
𝐶
𝑛(1)𝐴
+ 𝑛(1)𝐼𝑁
𝐶
𝑁(1)
where
𝐶(1)𝑖 is the first-phase inclusion probability of the i-th cell phone number;
𝐶
𝐶
𝑛(1)
is the size of the first-phase sample of cell phones, composed of 𝑛(1)𝐴
active cell phones and
𝐶
𝑛(1)𝐼𝑁
inactive cell phones;
𝐶
𝑁(1)
Total number of possible cell numbers under the national plan.;
Next, a second-phase sample was selected systematically out of the first-phase samples of active cell
telephone numbers. The second-phase inclusion probabilities of cell phones are
𝐶(2)𝑖|(1)𝑖 =
𝐶
𝑛(2)𝐴
𝐶
𝑛(1)𝐴
where
𝐶(2)𝑖|(1)𝑖 is the second-phase inclusion probability of the i-th active cell phone number conditional
on being selected in the first phase;
𝐶
𝑛(2)𝐴
is the size of the second-phase sample of active cell phones;
The unconditional inclusion probabilities of the second-phase active cell phones are
𝐶𝑖 = 𝐶(1)𝑖 𝐶(2)𝑖|(1)𝑖 =
𝐶
𝐶
𝐶
𝑛(1)𝐴
+ 𝑛(1)𝐼𝑁
𝑛(2)𝐴
𝐶
𝑁(1)
𝐶
𝑛(1)𝐴
=
𝐶
𝐶
𝐶
𝑛(1)𝐴
+ 𝑛(1)𝐼𝑁
𝑛(2)𝐴
𝐶
𝑛(1)𝐴
𝐶
𝐶
𝑛(2)𝐴
𝑛(2)𝐴
= 𝐶
𝐶 = ̂𝐶
𝐶
𝑁(1)
𝑅𝐴(1) 𝑁(1)
𝐴̂(1)
4
Inclusion probabilities of cell phones do not show a stratum index since most cell phone samples were not stratified
for the reasons stated above.
3
where ̂
𝑅𝐴(1) is the rate of active phones estimated in the first phase.5 Hence, the unconditional inclusion
probabilities of the second-phase active numbers 𝐶𝑖 can be expressed as the ratio between the active
numbers selected in the second phase and an estimate of the total active numbers in the frame 𝐴̂(1) .
Step 2: Design weights for households and individuals
The selection probabilities for households and individuals aged 18 and over are derived from the inclusion
probabilities of the cell phone numbers through which they are reachable. Therefore, the computation of
household and individual weights should account for multiple chances of selection. This multiplicity
weighting adjusts estimates to eliminate the over-representation of households and individuals in the
sample that can be reached through more telephone numbers than other households and individuals,
thereby reducing bias associated with unequal probabilities of selection.
Multiplicity adjustment
There is multiplicity probability when a household has a larger selection probability because it can be
selected through different sample elements (telephone numbers). Households with more than one cell
phone number are over-represented in sample designs like this. As a result, their selection probabilities
need to be adjusted to account for this increased chance of selection. The multiplicity-adjusted household
selection probabilities are computed as
𝐶𝑚𝑗 = 𝑚𝑐𝑗 𝐶𝑖
where
𝐶𝑚𝑗 is selection probability of the j-th household when contacted through a cell phone, adjusted
for multiplicity of working cell phones in the household;
𝑚𝑐𝑗 is the number of working cell phones in the j-th household;
Therefore, if a household has mc cell phones, its chance of being selected through a cell phone is mc higher
than a household where there is only one cell phone. Since the number of cell phones in a household is
unknown at the time of the sample design, it needs to be asked during the interview in the questionnaire.
For this purpose, the survey collected information about the number of cell phones in the respondent
households through the following question:
How many working cell phones in total are owned by the persons in your household, including you?
The probability of an individual being selected through a cell phone equals the inclusion probability of his
or her cell phone number.
𝐶𝑘 = 𝐶𝑖
where
5 ̂
𝑅𝐴(1) estimates are highly precise due to the very large size of the first-phase samples.
4
𝐶𝑘 is the selection probability of the k-th individual when contacted through a cell phone.
Household and individual design weights, w0j and w0k respectively, are the inverse of the above selection
probabilities
𝑤0𝑗 = 𝜋𝑗−1
𝑤0𝑘 = 𝜋𝑘−1
Weighting of data on children & adolescents
HFS collected specific data about a randomly selected child (0 – 5) or adolescent 6 through 17 years of age
in each interviewed household. To implement this, the questionnaire first collected a roster of all children
and adolescents living in each respondent household and selected one at random.
The child/adolescent weight is based on his/her probability of selection within the household, conditional
on his/her household being selected in the sample. Hence
𝐶𝑛𝑗 = 𝐶𝑗 1/ ∑𝑗 𝑛
where
𝐶𝑛𝑗 is the selection probability of the n-th child/adolescent in the j-th household when the
household is contacted through a cell phone,
𝐶𝑗 is selection probability of the j-th household when contacted through a cell phone, adjusted for
multiplicity of working cell phones in the household;
∑𝑗 𝑛 is the number of eligible children and adolescents (6-17 years old) in the j-th household;
Children and adolescents’ design weight w0n is the inverse of the above selection probabilities
𝑤0𝑛 = 1/ 𝐶𝑛𝑗
Step 3: Nonresponse adjustment
When a phone number is called, it is not always possible to carry out an interview. Nonresponse occurs
because of a number of constraints. Most common are that nobody answers the call (no contact), the
respondent is unwilling to cooperate (refusal), or language barriers exist.
The design weights of responding households and individuals were adjusted for nonresponse. This
adjustment is based on the inverse of the weighted response rate estimate. This is the ratio of the sum of
the design weights of all units (respondents and nonrespondents) to the sum of the design weights of
respondents.
5
𝑎𝑗 =
∑𝑗,𝑅 𝑤0𝑗 + ∑𝑗,𝑁𝑅 𝑤0𝑗
∑𝑗,𝑅 𝑤0𝑗
;
𝑎𝑘 =
∑𝑘,𝑅 𝑤0𝑘 + ∑𝑘,𝑁𝑅 𝑤0𝑘
∑𝑘,𝑅 𝑤0𝑘
where aj is the nonresponse adjustment factor that should be applied to responding households and ak
is the nonresponse adjustment factor for responding individuals. R and NR indicate the responding and
nonresponding units, respectively.
Thus, the nonresponse adjusted weights for responding households and individuals are
𝑤′𝑗 = 𝑤0𝑗 𝑎𝑗 ;
𝑤′𝑘 = 𝑤0𝑘 𝑎𝑘
Step 4: Calibration
Finally, the weights for responding households and individuals were calibrated to align with the
distribution of the phone-owning population by sex, age, and region, based on external data from official
national sources.
Calibration was performed by minimizing a measure of the distance between the input weights
(nonresponse adjusted weights in this case) and the calibrated weights, under the constraint that the sum
of the calibrated weights equals the sum of the totals of the auxiliaries from the external source. Unlike
the nonresponse adjustment, weights calibration requires auxiliary variables for respondents only.
Among available calibration methods, this survey employed raking (iterative proportional fitting) using a
logit distance function, which ensures that calibrated weights remain within acceptable bounds while
satisfying marginal control totals.
The final weights for responding households and individuals can then be expressed as
𝑤𝑗 = 𝑤𝑗′ 𝑔𝑗 = 𝑤0𝑗 𝑎𝑗 𝑔𝑗
𝑤𝑘 = 𝑤𝑘′ 𝑔𝑘 = 𝑤0𝑘 𝑎𝑘 𝑔𝑘
where
𝑤0𝑗 is the design weight for the j-th household;
𝑎𝑗 is the nonresponse adjustment factor for households;
𝑔𝑗 is the calibration factor for the j-th household;
𝑤0𝑘 is the design weight for the k-th individual;
𝑎𝑘 is the nonresponse adjustment factor for individuals; and
𝑔𝑘 is the calibration factor for the k-th individual.
6