EN FR HT
Republic of Haiti
Document Library
2,507 documents 116,441 pages
Estimating and Forecasting Income Poverty and Inequality in Haiti using Satellite Imagery and Mobile Phone Data

Estimating and Forecasting Income Poverty and Inequality in Haiti using Satellite Imagery and Mobile Phone Data

World Bank 2020 44 pages
Summary — A study estimating and forecasting income poverty and inequality in Haiti from satellite imagery and mobile phone data rather than from household surveys.
Key Findings
Full Description
A study estimating and forecasting income poverty and inequality in Haiti from satellite imagery and mobile phone data rather than from household surveys. The method matters in a country whose last full household survey is old: it offers a way to produce poverty estimates between surveys, with the caveats that entails.
Topics
PovertyStatistics
Geography
National
Keywords
pauvreté, inégalité, imagerie satellitaire, téléphonie mobile, estimation
Entities
World Bank, Neeti Pokhriyal
Full Document Text

Extracted text from the original document for search indexing.

Estimating and Forecasting Income Poverty and Inequality in Haiti using Satellite Imagery and Mobile Phone data Neeti Pokhriyal | Omar Zambrano | Jennifer Linares | Hugo Hernández Cataloging-in-Publication data provided by the Inter-American Development Bank Felipe Herrera Library Estimating and forecasting income poverty and inequality in Haiti: using satellite imagery and mobile phone data / Neeti Pokhriyal, Omar Zambrano, Jennifer Linares, Hugo Hernández. p. cm. — (IDB Monograph ; 824) Includes bibliographic references. 1. Poverty-Haiti-Data processing. 2. Income distribution-Haiti-Data processing. 3. Geographic information systems-Haiti. 4. Machine learning-Haiti. 5. Haiti-Social conditions-Data processing. 6. Haiti-Economic conditions-Data processing. I. Pokhriyal, Neeti. II. Zambrano, Omar. III. Linares, Jennifer. IV. Hernández, Hugo. V. Inter-American Development Bank. Country Office in Haiti. VI. Series. IDB-MG-824 JEL Classification: I3, I32, I38, O3, O31, O35, O54 Keywords: Haiti, Caribbean, big data, satellite imagery, poverty maps, measurement and analysis of poverty, economic development, income inequality, social innovation, machine learning, geographic information systems Copyright © 2020 Inter-American Development Bank. This work is licensed under a Creative Commons IGO 3.0 Attribution- NonCommercial-NoDerivatives (CC-IGO BY-NC-ND 3.0 IGO) license (http://creativecommons.org/licenses/by-nc-nd/3.0/ igo/legalcode) and may be reproduced with attribution to the IDB and for any non-commercial purpose. No derivative work is allowed. Any dispute related to the use of the works of the IDB that cannot be settled amicably shall be submitted to arbitration pursuant to the UNCITRAL rules. The use of the IDB’s name for any purpose other than for attribution, and the use of IDB’s logo shall be subject to a separate written license agreement between the IDB and the user and is not authorized as part of this CC-IGO license. Note that link provided above includes additional terms and conditions of the license. The opinions expressed in this publication are those of the authors and do not necessarily reflect the views of the Inter-American Development Bank, its Board of Directors, or the countries they represent. 2 Estimating and Forecasting Income Poverty and Inequality in Haiti Estimating and Forecasting Income Poverty and Inequality in Haiti using Satellite Imagery and Mobile Phone data Authors: Neeti Pokhriyal, Omar Zambrano, Jennifer Linares, and Hugo Hernández Content Executive summary 4 1. Introduction 5 2. Literature Review 7 2.1 Estimating social indicators from census and household surveys 7 2.2 Estimating social indicators using auxiliary data 8 3. Snapshot of Haiti’s social indicators 10 4. Methodology and Results 12 5. Policies to tackle social deprivations in the COVID-19 era and the recovery period 33 6. Conclusion 35 Acknowledgements The authors would like to thank UN ECLAC, especially Fabiana Del Popolo and Alejandra Silva from CELADE, for access to Haiti’s valuable census microdata through the use of their Redatam software. The authors would also like to thank Lisseth Escalante, Ricardo Benzecry and Salvador Traettino for their excellent research assistance, and Jean Marie Cayemitte and Gihanne Ambroise from the Banque de la Republique d’Haïti, Evans Jadotte from the World Bank, Raulin Cadet from Quisqueya University, Marta Ruiz-Arranz, Juan José Barrios, Agustin Filippo, José Antonio Mejía-Guerra, Nicola Magri, and Boaz Anglade from the Inter-American Development Bank, Jonathan Hersh from Chapman University, and Michael Mann from George Washington University for their valuable feedback. Finally, the authors would like to thank Duare Pinto for the graphic design of the report. Estimating and Forecasting 3 Income Poverty and Inequality in Haiti Prologue Tracking poverty and shared prosperity is an Several trends were observed: First, that better- essential task to inform a country’s policymaking performing communes in Haiti were usually located process. This is especially the case at present times, in the Ouest department, where the capital of the given that the economic crisis associated with the country is located. Second, that some communes COVID-19 pandemic threatens to disproportionately located in the Sud departments became increasingly affect vulnerable households. Nonetheless, doing more deprived than the rest over the last five so requires the frequent collection of household years, likely due to Hurricane Matthew in late 2016. survey data, which can be a costly and labor- Third, that the Nord-Ouest communes increasingly intensive effort. In addition, the granularity of became poorer than the rest over the past years, a the collected data is not sufficient for the proper trend consistent with the deterioration of the food identification of pockets of poverty and of income emergency situation in this department reported inequality. In Haiti’s particular case, there has by other international organizations. Finally, our not been a household survey since 2012, and the machine learning framework identified that one out most recently available data is only representative of four of the Haitian poor lived in 10 communes at the department level, potentially diluting in the Artibonite, Ouest, Nord-Ouest and Nord important social deprivations present at the departments, and that only three out of the 140 more local levels. communes studied present low levels of income inequality. In view of these challenges, the IDB commissioned a machine learning framework that can estimate the Our resulting estimates, which were mapped for distribution of income poverty and inequality in 140 their convenient interpretation, provide evidence communes and 570 sections communales in Haiti that there is a need of incorporating a territorial using features extracted from anonymized mobile perspective to all growth strategies. We hope phone data and satellite imagery. These features that this innovative approach to measuring social have proven to correlate with income poverty indicators is used for COVID-19 response efforts and other social indicators in several developing and for the design of a more inclusive growth countries, including Brazil, Mexico, Belize, Cambodia, agenda in Haiti. We also hope that this framework and Bangladesh. While these techniques are not is adopted and further refined as new improvements intended to be a substitute of the valuable data to machine learning techniques are developed obtained through household surveys, they represent in order to complement future estimations an innovative and timely complement to the efforts of social indicators. of tracking social indicators. Yvon Mellinger IDB Country Representative in Haiti 4 Estimating and Forecasting Income Poverty and Inequality in Haiti 1 Haiti is not the only country that has struggled to carry out household surveys and censuses on a regular basis. In fact, Serajuddin et al. (2015) estimate that over a third of the world’s developing or middle-income countries had one or less poverty estimates between 2002 and 2011. This makes tracking poverty and shared prosperity a challenge in the developing world. One of the main reasons behind this challenge is the high cost of collecting ground data. According to Kilic et al. (2017), the average cost of a survey in a fragile state is Introduction about $186 per household2. If we were to take the sample of households surveyed in ECVMAS (about 5,000), the cost would be about $930,000 in real 2014 US dollars (excluding any capacity building Households surveys and censuses are the building costs). Further, household surveys may not be blocks of informed policymaking as they contain representative at the desired geographical level- relevant demographic and socioeconomic the 2012 ECVMAS, for instance, is representative at information of a country’s population. In order to the national and departmental levels (10 geographic monitor the potential progress of a country in the units for the latter), potentially masking variations achievement of development goals, both household at more granular levels of disaggregation-the surveys and census data must be collected on a commune level (144 geographic units) and the regular basis. In fact, the Enhanced General Data section communale level (571 geographic units). Dissemination System (e-GDDS) of the International Moreover, safety concerns may limit the capacity Monetary Fund encourages countries to collect of the government in carrying out these surveys population data (via population censuses) every in countries in conflict (Engstrom et al. 2017). ten years, and poverty data (via household surveys) These considerations make tracking development every three to five years. Haiti carried out its last goals such as the Sustainable Development Goals Population and Housing Census (Recensement (SDGs)-which, according to Kilic et al. (2017), rely General de la Population et de l’Habitat) in 2003 more on household survey data than the Millennium and its post-earthquake household living conditions Development Goals (MDGs)-a challenge for a survey (L’Enquête sur les Conditions de Vie des country like Haiti. Ménages Après Séisme or ECVMAS) in 2012. Therefore, at the time of publishing of this report, In view of this, researchers have recently turned Haiti had not carried out a census in seventeen to auxiliary sources of data such as mobile phone years and a household survey in eight years1. records, satellite imagery, and remote sensing 1 Prior to the 2012 ECVMAS, the household living conditions survey had been conducted in 2001. 2 The average was estimated using the cost per household for three fragile countries: Afghanistan ($109.25 per household), Iraq ($149.23) and Yemen ($298.42). Estimating and Forecasting 5 Income Poverty and Inequality in Haiti data to estimate social indicators such as income social indicators was done in two stages: a) First, poverty (Blumenstock et al. 2015; Jean et al. 2016; small-area estimates (SAE) of socioeconomic Engstrom et al. 2017) Many of these auxiliary data indicators were calculated using a combination sources have the advantage of being frequently of the 2003 census data and the 2012 ECVMAS, updated at a lower cost and of allowing for a b) auxiliary data was used to estimate the social further disaggregation of a country’s administrative indicators in 2014 and 2019 using a machine learning areas. Further, combining several auxiliary datasets model trained on the SAE from the first stage. The (Pokhriyal and Jacques, 2017 and Steele et al., 2017) 2014 commune level estimates were validated using like mobile phone and remote sensing data have the SAEs from the first stage. Our trained model is proven to result, in some cases, in more accurate also used to forecast the deprivations for 2019 both estimations than using each dataset separately. In at commune and sections communales. spite of the novelty and cost-efficiency of these techniques, it is important to note that auxiliary Three trends were observed from the resulting sources are meant to be a complement to ground- estimates: a) better-performing communes were truth data collection efforts— especially for the usually located in the Ouest department in both prolonged intervals between each household 2014 and 2019; b) Some communes located in survey— and should not be considered a substitution, the Sud departments became increasingly more as they provide a snapshot of the distribution of deprived than the rest between 2014 and 2019, likely poverty and other social deprivations within Haiti due to Hurricane Matthew in late 2016; and c) the based on the 2012 national poverty line (as we Nord-Ouest communes increasingly became poorer will explain in the methodology section). Updated and more deprived than the rest in 2014 relative to national poverty figures would require an update 2019, a trend consistent with the deterioration of of the national poverty line and the determination the food emergency situation in this department of the updated income and consumption levels of reported by the Food and Agriculture Organization the population, which can only be obtained from (FAO) and the World Food Programme (WFP). ground-truth data. In the context of the COVID-19 pandemic, it is The objective of this study is to build a computational essential to have a better idea of where the most framework that can estimate the distribution of vulnerable population in the country resides in income poverty, income inequality, and standards order to prioritize social interventions given the of living deprivation3 in 20144 and 2019 for 140 government’s limited fiscal space5. Therefore, the communes and 570 sections communales in Haiti present study lists and maps the names of the most using features extracted from anonymized mobile vulnerable communes and sections communales in phone data and satellite imagery (referred to as the 2014 and 2019. auxiliary data hereafter). The estimation of these 3 Risk of malnutrition and ownership of durable goods, based on the definition of Santos and Villatoro (2018) 4 The 2014 maps are dynamic, as they were built using 2014 satellite imagery and mobile phone metadata from March to May in years 2016-2018. 5 The Haitian Ministry of Economy and Finance estimates that Haiti’s non-financial public sector’s deficit will be 6.2%. 6 Estimating and Forecasting Income Poverty and Inequality in Haiti Literature Review 2 2 as household surveys often use samples that are too small to produce representative estimates beyond national or regional levels. Small area estimations (SAE) is a frequently used indirect estimation technique for the estimation of social indicators, notably poverty, at more disaggregate levels. According to Das and Haslett (2019), basic SAE methods can estimate small area linear parameters, including means and totals, yet they are not well Literature Review suited for the estimation of non-linear functions, such as distributions. Nonetheless, linear methods that incorporate unit records from a census or an administrative database can model non-linear Given the study’s two-stage approach for the functions with considerably increased accuracy. estimation of social indicators, the literature review section is divided into two subsections: the first One of such methods is the three-level Elbers, subsection summarizes works related to estimating Lanjow, and Lanjow (2002) or ELL method, which poverty from census and detailed household was the first one developed for these purposes. The surveys to create a representative sample at more ELL method consists of the following steps: (i) First, disaggregate levels, which act as regression targets similarities between census data and household for the poverty estimates generated from auxiliary surveys are identified in order to construct a data. The second subsection provides a summary of common vector of independent variables, (ii) the literature related to using auxiliary data sources second, a parametric model of household income for poverty measurement and the estimation of determinants is estimated based on the available other social dimensions. information at the highest level of representation of the household survey, and (iii) Finally, the empirical 2.1 Estimating social indicators from distribution obtained in (ii) is simulated over the census and household surveys census data to obtain small area estimations of greater geographical representation. The ELL Mapping socioeconomic conditions has historically method (and variations of this method) has been been of interest for governments, development successfully used for the estimation of social organizations, and academia. However, indicators6 in many developing countries, including representative data at granular administrative Belize (Hersh, et al. 2020), Brazil (Elbers, et al. levels is often unavailable in developing countries 2008), Mexico (Demombynes, et al. 2006), Thailand 6 Income poverty in most cases, yet some authors have also estimated indicators such as undernutrition. 7 According to Das and Haslett (2019), EBP is based on an area-specific two-level nested error regression model that includes small area level and household level effects, while ELL is based on cluster-specific two-level model using cluster and HH variability with contextual variables from the census. Estimating and Forecasting 7 Income Poverty and Inequality in Haiti (Healy et al., 2003), Cambodia (Fujii, 2004), South and 2001 poverty data in Bangladesh. The study Africa (Alderman et al., 2002), Brazil (Elbers et al., found that the ELL performed better than the EBP 2004), Bangladesh (BBS & UNWFP, 2004; Haslett and MQ in terms of relative bias and relative root et al., 2014), the Philippines (Haslett & Jones, 2005) mean squared errors when the majority of small and Nepal (Haslett & Jones, 2006). domains have single clusters in the sample and when there is minor between-area variation in the Another methodology for SAEs worth mentioning population. Given its relatively simple methodology is the one proposed by Molina and Rao (2010), and widespread success in the estimation of social which estimates non-linear small area population indicators in developing countries, a computational parameters using the empirical Bayes or best variant of the ELL method was our technique of prediction (EBP). According to Das and Haslett choice for the first part of our exercise, combining (2019), this methodology results in estimators with data from population census with data from minimum mean squared errors (MSEs) that are “best household surveys to generate a valid/robust predictors” through Monte Carlo approximation, estimations of per capita income at small levels of assuming that the transformed welfare variable territorial disaggregation8. follows a nested error regression model7. This method has been used mostly in European 2.2 Estimating social indicators Union countries. using auxiliary data According to Das and Haslett (2019), both of the aforementioned methods derive from standard Several studies have highlighted the relationship linear random effects models with potentially strong of poverty and information captured by satellites, distributional assumptions and formal specification including but not limited to, nightlights, weather, of the random portion. Moreover, they may not vegetation, and meteorological data. (Elvidge et al., be robust to outliers in the response variable, yet 2009; Pokhriyal and Jacques, 2017; Jean et al., 2016) outliers used in the survey-based model do not Moreover, the growth in the number of satellites necessarily imply outliers or non-robustness in the and the enhancement of their sensing capabilities aggregate SAEs. has enabled precise and timely reporting of various remote sensing measurements. In fact, many of An alternative to SAE is the MQ approach proposed the satellites have a revisit capability of 24 hours, by Chambers and Tzavidis (2006), which is based enabling almost real time scanning of earth’s on modelling quantile-like parameters of the surface. Newer sensors aboard these satellites conditional distribution of the target variable given also capture geo-spatial measurements like cars, the covariates. The MQ approach is distribution- shadows (which act as proxy for building heights), free and avoids the strong assumptions associated density of buildings, transportation information, with specification of random effects, allowing and roof types (which are useful in the mapping inter-area differences to be characterized by area- of slums). These measurements are captured at specific coefficients. Nonetheless, this method also very fine spatial resolution of 25 cm or 30 cm, with has a shortcoming: it does not capture cluster- varying spectral bands, which facilitate mapping specific random effects by failing to differentiate and other downstream tasks from these images. between small area-specific random effects and The basic idea behind using satellite imagery to cluster-specific effects. (Das and Haslett, 2019). determine deprivations within a country is to extract MQ has been used to estimate poverty at the local high-level patterns or signals which correlate with administrative level in several countries in Europe, socioeconomic deprivations of a geographical area. including Poland (Marchetti et al., 2018), Albania For example, satellite images of a given geographical (Tzavidis et al.,2008), and Italy (Giusti et al., 2011). area are scanned to find signs of urbanization, which are shown to correlate with economic development A simulation ran by Das and Haslett (2019) uses the of a country. three aforementioned methods to reconstruct 2000 8 In this approach, as in a traditional ELL method, at the first stage of the estimation a common support vector of variables merging household survey data and census data was determined. At the second stage, a predictive per-capita income model considering both cluster/locational effects, and idiosyncratic household effects, was predicted using a machine learning (ML) framework. The Machine Learning (ML) algorithm was trained using the smaller/representative data set (household survey, and then a probabilistic inference was projected over the out- of-sample larger dataset (census), in an analogue approach to a parametric SAE. The result was a complete set of predicted income or classification for all the households in the census sample, allowing the estimation of socioeconomic indicators at small territorial level. With no parametric structure to restrict the classification/imputation process, the machine learning approach allows for greater flexibility, focus on out-of-the-sample predictions, and systematic optimization of dimensionality problem to avoid overfitting. 8 Estimating and Forecasting Income Poverty and Inequality in Haiti Literature Review 2 Satellite imagery has been studied in the past to a specific given location as one of 62 target classes look for texture-based features. However, with (e.g. an airport, a flooded road, a place of worship, advances in deep learning algorithms, researchers a construction site, and so on) or as none of them increasingly focus in extracting automated features (false detections). FMOW images vary in quality from satellite imagery to learn signals of poverty and are distributed over more than 100,000 globe deprivations. (Jean et al., 2016 and Pandey et locations, which leads to high intraclass variations al., 2018) Researchers have tested the spatial and considerable interclass confusion. This model generalization capability of Jean et al., 2016 in was used as a feature extractor, as detailed in the developing countries (Head et al., 2017) such as methodology section of this report. This method is Haiti and Nepal, and reported that the method is also compared to a direct night-light-based poverty sensitive to hyper-parameter setting, and does estimation approach (Jean et al., 2016). not trivially generalize to other spatial locations, especially Haiti. They attribute it to the fact that In addition to the use of satellite imagery as Haiti has a high urban population density (55% of auxiliary data for poverty prediction, recent works total, with 74% of these living in slums9) and thus like Steele et al., 2017, Pokhriyal and Jacques, 2017 poses unique challenges in learning signals of have shown the efficacy of using a combination development from satellite imagery. of anonymized mobile phone metadata (which allows for extraction of features that correlate with Additionally, the need to study temporal socioeconomic measures, such as mobility and generalization and evolution of machine learning activity features), and geographical information models that predict poverty and wealth based data for poverty prediction at high resolution. on satellite imagery is highlighted in Engstrom Given that both data sources are generated at and Hersh (2017). Simultaneously, the need to different spatial resolutions, they complement have spatially high-resolution poverty metrics is each other well: the granularity of mobile data highlighted by Watmough et al. (2019). Tang et al. in urban areas compensates for the coarseness (2018) employ Normalized Difference Vegetation of geographical information data in these areas, Index (NDVI) data to predict the rate of change and allows for estimations at much disaggregate of poverty. However, given that it is based on a geographical areas, such as neighborhoods (Steele vegetation index, it is limited in its capability of et al., 2017). Using the combination of both data predicting broader deprivations of poverty and is sources provided improved predictive power and at coarser resolution of 250 m x 250 m. lower errors than using these separately, as shown in Pokhriyal and Jacques, 2017. Given that Haiti has One of the most recent initiatives feature extraction a high proportion of its population living in urban in this area is the Functional Map of the World areas, we decided to use this method for our 2014 (FMOW) challenge (Christie et al., 2018), which estimates to obtain more precise results. consists of creating automatic solutions to classify 9 According to the World Bank. Estimating and Forecasting 9 Income Poverty and Inequality in Haiti 3 Snapshot of Haiti’s There are significant differences in the provision social indicators of basic services among the urban and rural households. About 11 percent of the rural population has access to electricity, compared to 63 percent in urban areas. Likewise, only 16 percent of the According to the 2012 ECVMAS, 59 percent of population has access to improved sanitation the Haitian population lived under the national facilities, while 48 percent did in the cities. Therefore, poverty line (2.41 USD/day), while 24 percent lived while poverty is present in both urban and rural in extreme poverty. About 45 percent of the Haitian areas in Haiti, it tends to be a rural phenomenon. population lives in the rural areas, where, according Nonetheless, it is important to note that informal to (Ghayad et al., 2019), almost two-thirds of the urban settlements have increased over the population is considered chronically poor. Extreme last decade. poverty had declined from 31 percent in 2000 to 24 percent in 2012 at the national level, with urban In terms of income inequality, according to the World poverty halving, while poverty levels remained about Bank, the gini coefficient in Haiti was 0.61 in 2012, the same in the rural areas over the same period10. a level which had remained practically unchanged According to the World Food Programme’s (WFP) since 200111. The gini coefficient estimated by the 2020 Global Report on Food Crises, in Haitian rural Haitian Institute of Statistics (IHSI) in 2012 was areas, most vulnerable households lack agricultural even higher than the one estimated by the World work opportunities due to high labor costs and Bank, and stood at 0.6812. According to IHSI, there limited resources of farmers. Thus, vulnerable is heterogeneity within each of the ten Haitian households recur to alternative sources of income departments: In fact, 89 percent of household such as migration, petty trade, or selling of charcoal. income inequality in 2012 was intraregional. 10 According to Ghayad et al. (2019), rural poverty is common among developing countries, as opportunities for education, employment, and access to technology tend to be concentrated in the urban regions. 11 As reported in World Bank Group. 2014. “Poverty and Inclusion in Haiti: Social Gains at Timid Pace”. A similar trend is observed when using The Standardized World Income Inequality Database (Solt, 2019). 12 Note that these values significantly differ from the gini coefficients reported by the World Bank and the Standardized World Income Inequality Database in cross-country databases. In those databases, gini indices are standardized by welfare definition and adult-equivalence scale in order to facilitate comparability with other countries. 10 Estimating and Forecasting Income Poverty and Inequality in Haiti Snapshot of Haiti’s social indicators 3 While the gini coefficient is a widely-used measure between 2010 and 2018 (see figure 2), yet it has of income inequality, it is important to highlight been at a slower pace. According to the 2019 UNDP that it does not feature subgroup consistency: in Human Development Report, this was the case other words, if inequality declines in one subgroup due to an increase in inequality in terms of years (for instance, a region) and remains unchanged of schooling and life expectancy. Haiti is the fifth in the rest of the subgroups, the gini may not country in the world (out of 148) in terms of losses properly reflect this change. Therefore, the United in HDI when controlling for these inequalities13,14. Nations Development Program (UNDP) uses the Haiti had registered significant gains in both the Atkinson measure of inequality, which, in addition HDI and the IHDI over the past decade. However, to displaying subgroup consistency, is also sensitive in 2018, the IHDI registered a decrease. to inequality in the lower end of the distribution by putting more weight on the lower end. When Finally, food insecurity is widespread in Haiti. using this measure, the UNDP shows that income According to the WFP, one out of every three inequality has been increasing over the last decade, Haitian were acutely food insecure before COVID-19 with the most dramatic increase registered in 2018, became a pandemic. This is equivalent to 3.7 million when the sociopolitical crisis erupted (see figure 1). Haitians15. This was especially the case in the lower Northwestern region of the country and in the In terms of human development, Haiti registered urban commune of Cité Soleil. Disruptions in the important gains over the last decade, with the food supply chains due to the global pandemic, human development index (HDI) increasing from coupled with violent conflict, inflation and the 0.474 in 2010 to 0.503 in 2018. Nonetheless, the HDI currency depreciation registered over the past year is much lower when adjusted for income, education, could increase the severity of food insecurity and and health inequalities (the inequality adjusted HDI, malnutrition in the coming months. or IHDI). The IHDI still registered an improvement Figure 1: Income inequality (%) using the Atkinson Figure 2: Evolution of the Human Development Inequality Index. Source: UNDP-HDI 2019 Index Source: UNDP-HDI 2019 0.52 0.34 51 0.33 0.50 50 0.32 0.48 0.31 49 0.46 0.30 48 0.29 0.44 0.28 47 0.42 0.27 2010 2011 2012 2013 2014 2015 2016 2017 2018 46 2010 2011 2012 2013 2014 2015 2016 2017 2018 HDI Inequality-adjusted HDI (right axis) 13 The ranking was based on the average losses in HDI due to inequality between 2010 and 2018 for countries with 4 or less missing values. The top four countries with the biggest losses were: Comoros, Central African Republic, Guinea-Bissau, and Namibia. 14 Note that IHDI is based on the Atkinson index. IHDI is not association-sensitive, and therefore does not capture overlapping inequalities. In order to capture association sensitivities, all the data for each individual must be available from a single survey source, and this is currently not the case for Haiti. 15 Haiti is the tenth country with the highest number of people living in food insecurity crisis or worse after Yemen, DR of Congo, Afghanistan, Venezuela, Ethiopia, South Sudan, Syria, Sudan, and Northern Nigeria. Estimating and Forecasting 11 Income Poverty and Inequality in Haiti 4 Stage 1: Estimating Computational Small-Area Estimates (SAEs) from Census and Household Survey Data As discussed in the literature review, our method of choice for estimating SAEs is a computational Methodology variant of the ELL method, given its relatively simple methodology and widespread success in and Results the estimation of social indicators in developing countries. In the vein of the ELL method, our computational approach consists of the following The objective of this study is to estimate income steps: (i) first, similarities between census data poverty, income inequality (Gini), Foster Greer and household surveys are identified in order Thorbecke index (FGT) 16, and standards of to construct a common vector of independent living deprivation using two different sources of variables, (ii) second, a parametric model of auxiliary data: anonymized mobile phone data household income determinants is estimated based and satellite imagery17. The estimation of these on the available information at the highest level of social indicators was done in two stages: a) First, representation of the household survey (in Haiti’s small-area estimates (SAE) were calculated using a case, at the department level), and (iii) finally, the combination of the 2003 census data and the 2012 empirical distribution obtained in (ii) is projected ECVMAS, b) auxiliary data was used to estimate the over the census data using Machine Learning (ML) social indicators in 2014 and 2019 using a machine trained predictive modeling to obtain small area learning framework18. The 2014 estimates were estimations of greater geographical representation obtained by training and validating using the SAEs (in Haiti’s case, to obtain estimations at the from the first stage. commune level)19. 16 The Foster–Greer–Thorbecke (FGT) indices are a family of poverty metrics. FGT0 measures how widespread poverty is (headcount poverty), FGT 1 measures how poor the poor are (the expenditure discrepancy of poor people towards the poverty line), and FGT2 gives an indication of how severe poverty is (as it puts higher weight on the poverty of the poorest individuals, making it a combined measure of poverty and income inequality). 17 Satellite imagery denotes both aerial imagery for year 2014 as well as the publicly available satellite imagery for 2019. 18 More precisely, both aerial and mobile data were used for 2014, while only satellite data was used for 2019, due to the unavailability of mobile phone records. 19 The basic problem structure of a typical SAE, which is a smaller data set (household survey) that contains information that has to be imputed over a larger data set (census), replicates the typical problem structure of predictive modelling from the Machine Learning framework. Predicting modeling algorithms exploit the value of information of smaller/representative data sets, where the algorithm learns the optimal parameters to solve estimation/classification problems (training sets), to then project a probabilistic inference over an out-of-sample larger dataset (scoring set). Given this shared structure, it is plausible to use a machine learning approach to solve the problem of SAE of socioeconomic indicators to produce poverty maps. 12 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Two data sources were used for the construction the effect of γc, which is the cluster/locational effect, of Haiti’s SAE estimates: the 2012 household survey and εc,h which is the idiosyncratic household effect. (ECVMAS) and the 2003 population census (the most recent one done in Haiti). The 2012 household Our predictive modeling process for generation of data was obtained from the IHSI website, while the poverty indicators in this study included continuous, census data was obtained from the United Nations binary and multinomial specifications. A variety Economic Commission for Latin America and the of Machine Learning algorithms were evaluated Caribbean (ECLAC) Population and Development including decision trees, including Random Forest Division (CELADE), through their Redatam software. (RF) and Gradient Boosting Machines (GBM), linear models (GLS with Lasso), deep learning models Just like other household surveys, ECVMAS collects (Neural Networks), and a combination of all of them data on primary and secondary sources of income (Stacked Learning). ML models were selected over and transfers for the households, but it is only the performance criteria of the cross-validation representative at the department level. On the phase of the training set20. Our preferred ML model, other hand, the 2003 Population Census contains the linear GBM, minimized the sum of the squared detailed information at the individual level but does errors between the estimates and actual headcount not report income data. Nonetheless, both data poverty rates at the statistically representative sources have 129 household and individual variables territorial portions of the ECVMAS 2012, thus in common (see table A.1 in the appendix), making accurately reproducing national and departmental it possible to construct a common support vector headcount poverty rates (FGT0) for the base year. of poverty-related independent variables. This is the first step of the ELL framework. Table 1 reports the fittest models for each model specification, its estimated headcount poverty rate For the second step, a per-capita income model was at the national level and a measure of each of their estimated for each household in every department, sum of squared errors with respect to the ground using the following structure: truth ECVMAS data. As previously mentioned, the selected model closely replicates the national (1) Y h,c=βXc,h+μc,h poverty headcount rate and has the lowest sum of squared errors at the department level. In our case, (2) μc,h=γc+ εc,h this model is the GBM linear specification adjusted by a constant. As shown in figure 3, this GBM model specification closely replicates each department’s Where in (1) and (2), Yh,c used household per capita headcount poverty rate, with a R2 of 0.685. income as a dependent variable with a constant (k=2,510 Haitian gourdes) linear adjustment and μc,h is the uncorrelated random error term composed by Table 1: Estimated models for commune poverty rate in Haiti using ECVMAS data Resulting Sum of Squared Model national Errors poverty rate (Department level) Survey (ground truth) 0.597 0 Linear GBM specification adjusted by a constant 0.571 8.857 Linear specification adjusted by a constant for each income percentile 0.616 9.927 Dichotomous specification adjusted to the national poverty rate 0.631 20.475 Linear (natural logarithm) specification 0.759 48.716 20 A Confusion Matrix for each evaluated model allowed to compare the models’ out-of sample predictive capabilities in terms of accuracy, recall and precision metrics. Further, ROC-AUC Curves and Predictive Deciles Plots supported the metrics for the selection of learning models. Our preferred machine learning routine was Gradient Boosting Machine, a technique that combines predictions from multiple decision trees to generate a model that minimizes out-of-sample prediction error. Estimating and Forecasting 13 Income Poverty and Inequality in Haiti Results Figure 3: Poverty headcount ratio (using the 2012 national poverty line) R2 = 0.685 Income Poverty 0.78 Nord Sud Our model yields a national poverty rate of 57.1%, which is equivalent to approximately 5.8 million Sud-Est individuals. Figure 4 shows the resulting poverty Artibonite Nord-Est rates at the commune level, which range from 0.65 Center 12.3% (Delmas in the Ouest Department) to Household Survey Nord-Ouest 89.8% (Île à Vache in the Sud Department). Note Ouest Nippes that there is a concentration of communes with very high levels of poverty in the eastern side of Artibonite, in the southernmost communes 0.53 in the Sud and Sud-Est departments, and in the Centre department. However, notice also that Grande’Anse the mean and median poverty levels among all communes are very high, at 72.7% and 75.5%, 0.40 respectively. Among the 10 communes with the 0.40 0.53 0.65 0.78 highest poverty rate in 2012, six were located in the northern region (which included the Nord-Est, Model Nord, and Nord-Ouest). Figure 4: 10 Poorest communes in 2012 (income poverty in USD) Mean Median St. deviation Minimum Maximum 72.7% 75.5% 12.2% 12.3% 89.8% Commune Department % Île à Vache Sud 89.8% Terre Neuve Artibonite 89.6% Capotille Nord-Est 87.5% Borgne Nord 87.4% Baie-de-Henne Nord-Ouest 87.1% Cornillon / Ouest 86.8% Grand Bois Sainte Suzanne Nord-Est 86.6% Bahon Nord 86.3% Bas Limbe Nord 86.2% Chantal Sud 86.2% 14 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 An alternative way to report the results is the to note that the fact that a commune registered a number of people living in poverty. Among the 10 lower gini coefficient (like in the case of Cité Soleil, communes with the highest number of people living which has a poverty level of 77.8%) is not indicative in poverty in 2012 (Figure 5), most of them were of the level of wealth of its households, rather, a concentrated in two departments, Artibonite (in comparison of income between households within Gonaïves, Saint-Marc, Dessalines, Petite-Rivière and this particular commune. Thus, a low gini coefficient Saint-Michel de l’Attalaye) and Ouest (Cité Soleil, in Cité Soleil indicates that most households within Port-au-Prince, and Carrefour). In addition, one out this commune have similar levels of (low) income. of four of the Haitian poor lived in the 10 communes listed in figure 5. The FGT2 index, which shows how severe poverty is22, shows that the communes with the most intense In terms of income inequality, the communes with poverty levels in 2012 were located in the Sud and the highest levels of inequality were located in the Artibonite departments. The communes with less Sud Department and Artibonite. Nonetheless, notice poverty intensity were those located in the Ouest that that the average commune in Haiti has a gini department (Port-au-Prince, Pétion-Ville, Carrefour, index of about 0.7021. These high levels of income Delmas) and Cap-Haïtien (Nord). Notice in figure inequality are reflected in figure 6, which shows 7 that there is a concentration of communes with that communes in most departments have similar severe poverty in Centre, the eastern region of levels of income inequality. In fact, only three out of the Artibonite department, and the southernmost the 140 communes examined in this exercise had a communes of the Sud department. gini coefficient of less than 0.50. These were Cité Soleil, Delmas, and Tabarre in Ouest. It is important Figure 5: 10 Communes with the largest number of people living in income poverty in 2012 Commune Department Gonaïves Artibonite 196,080 Cité Soleil Ouest 180,171 Port-au-Prince Ouest 171,501 Saint-Marc Artibonite 145,331 Cap-Haïtien Nord 134,924 Carrefour Ouest 133,418 Dessalines Artibonite 126,156 Petite Rivière Artibonite 121,573 Saint-Michel Artibonite 117,312 de l’Attalaye Port-de-Paix Nord-Ouest 112,005 21 The gini index or coefficient is a measure of income inequality, where 0 represents perfect equality and 1 represents perfect inequality. Communes with gini coefficients approaching 1 feature the highest levels of income inequality. 22 This measure gives greater weight to those that fall far below the poverty line than those that are closer to it. Estimating and Forecasting 15 Income Poverty and Inequality in Haiti Figure 6: 10 communes with the highest levels of income inequality in 2012 (Gini Index) Mean Median St. deviation Minimum Maximum 70.1% 70.6% 6.7% 37.7% 81.5% Commune Department % Côteaux Sud 81.5% Port- à -Piment Sud 81.4% Chantal Sud 80.6% Île à Vache Sud 80.3% Bahon Nord 80.0% La Chapelle Artibonite 80.0% Terre Neuve Artibonite 79.7% Sainte Suzanne Nord -Est 79.5% Marmelade Artibonite 79.4% Roche-à- Sud 77.8% Bateau Figure 7: 10 Communes with highest FGT2 poverty (in USD) Mean Median St. deviation Minimum Maximum 51.4% 51.9% 11.1% 17.8% 71.5% Commune Department % Chantal Sud 71.5% Bahon Nord 70.8% Île à Vache Sud 69.4% Terre Neuve Artibonite 69.1% La Chapelle Artibonite 68.6% Port- à -Piment Sud 68.5% Sainte Suzanne Nord -Est 68.5% Cornillon / Ouest 67.5% Grand Bois Marmelade Artibonite 67.5% Les Anglais Sud 67.1% 16 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 In terms of standards of living deprivation23, most Most of these are located in the Ouest department communes registered high levels of deprivation (except for Cap-Haïtien, which is located in the Nord (over 90%, on average). The communes with the department). These communes are highly urban. least deprivations were Delmas, Port-au-Prince, Carrefour, Pétion-Ville, Cap-Haïtien, and Tabarre. Figure 8: 10 Communes most deprived of standards of living Mean Median St. deviation Minimum Maximum 91.5% 96.1% 12.1% 26.2% 99.7% Commune Department % Terre Neuve Artibonite 99.7% Cornillon / Grand Ouest 99.6% Bois Belle Anse Sud-Est 99.5% Boucan Carre Centre 99.4% Plaisance du Sud Nippes 99.3% Savanette Centre 99.3% Arnaud Nippes 99.2% Île à Vache Sud 99.2% Maissade Centre 99.1% Saint Jean du Sud Sud 99.1% 23 The standards of living deprivation is a subcomponent of the multidimensional poverty index according to the Santos and Villatoro (2018) definition. The results discussed in this report, however, only account for the incidence (H) of the standards of living deprivation, and not the intensity (A). Estimating and Forecasting 17 Income Poverty and Inequality in Haiti Stage 2: Estimating Social Pagerank is an important centrality measure and has been used in context of quantifying information Deprivations using Mobile Phone accessibility from mobile phone data to study Data and Satellite Imagery socioeconomic deprivations. (Pokhriyal and Dong, Mobile phone data description, 2015; Soto et al., 2011). feature extractions, and data Since we have access to very limited mobile aggregating procedures features, aggregating them to communes and sections communales depends on the specific The results obtained from the previous stage are feature. We calculated three features, one each from used to validate more contemporary estimates of population estimates, migration and road usage poverty and other social deprivations using mobile data, respectively. The first resulting feature is the phone metadata and satellite imagery. This section log of population density for each microregion. The provides the data description for each source, second feature is the network’centric pagerank along with feature extraction and building targets feature, as discussed earlier. The third feature is for deprivations. Please note, that since our study the road usage statistics per microregion and is was performed both at the spatial granularity of calculated from road usage data. commune and sections communales. We will refer to these collectively as “microregions” hereafter, Imagery data description, feature unless otherwise stated. extractions, and data aggregating The anonymized mobile phone data was obtained procedures from FlowMinder24 and it includes data from March In the case of the imagery, we used 3,246 aerial 1 to May 31 for 2016, 2017, and 2018. In addition images from 2014. Each of the images was 12,000 x to allowing for population estimates, the dataset 12,000 pixels and each pixel represents 0.25 x 0.25 allowed for the estimation of road usage by sq. m. on the ground. Thus, each image sums up to 3 producing a set of trips from mobile data, and then x 3 sq. km on the ground. Given the high resolution mapping these trips. From this data, we get the of satellite images, several pre-processing steps estimated road usage per micro region, by finding are needed that involve coarsening to a required the road segments that lie within microregional resolution and dividing the image into units that can boundaries and aggregating the road usage for be taken as input by the feature extractor model. those segments. For feature extraction, we use two state-of-the-art methods described as follows: Moreover, the mobile phone dataset allowed for the estimation of the number of subscribers in a given 1. Functional features-based extractor: It iis based location across different time points, thus producing on the Functional Map of the World (FMOW), an estimate of population-wide movement for the which is a dataset consisting of 1 million satellite years 2016, 2017 and 2018 available at microregional images spanning over 200 countries to train levels (for more information, please refer to the machine learning methods that can predict the appendix box A.1). functional purpose of the buildings/land in the images. The idea is to classify each image as We conceptualize the migration data as a directed belonging to one of the 63 functional classes weighted graph G, with each microregion as a node, (like airport, construction site, shopping mall, or V, and an edge exists between two nodes Va and residential unit). We use a state-of-the-art deep Vb if there is a population flow from Va to Vb. The neural network model for imageanalysis called amount of population flow determines the weight DenseNet (Huang et al., 2016), which is trained of the edge. The graph G encodes the migration on the FMOW dataset images (for details of structure across the micro regions. To ascertain a the model, see Christie et al.,2018) To extract metric with each node, we run centrality measures, features from the satellite imagery, we divide namely the Pagerank algorithm that calculates the the imagery into smaller units, each covering influence or importance of a node, recursively, an area of 112 X 112 sq. m. The model takes in using the importance of the nodes connected to it. 24 Inter-American Development Bank mimeo: Flowminder final consulting report: Caracol Industrial Park and National Road Network. An analysis of commuting and migration patterns. 18 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 satellite imagery and runs it through its network. to automatically reduce the dimensions that retain We intercept the model at the 6th penultimate 90% of the variance in the data. The benefit of layer and use the output at that stage as our dimensionality reduction is that it considerably helps feature vector for that image. reduce computational cost. Also, since we have limited spatial locations (communes and sectional 2. Nightlight-based extractor: As a second communes), reducing the dimensions helps prevent feature extractor model, we use another state- overfitting the model to the data. of-the-art model, based in Convolutional Neural Network (CNN) that is trained to learn It is important to note that the satellite imagery the relationships between daytime satellite is at a higher resolution than the policy planning imagery and its corresponding nighttime light units in Haiti for which poverty and other social imagery in the context of predicting poverty deprivations is ideally reported. Therefore, as a next and socio-economic deprivations. It is trained step, the features from the satellite images were for over 3,000 DHS (Demographic and Health aggregated to the desired microregions using the Surveys) clusters in Africa (For details, see Jean following approach. et al., 2016). This model takes as inputs satellite imagery at resolution of roughly a 1-sq. km area Satellite images are in the form of rectangular grids, for the satellite imagery of years 2014 and 2019 whereas the microregions are available as shapefiles and extracts features corresponding to it. containing their irregular polygon boundaries. Numerous rectangular grids fall within a polygon During our experiments, we found that the night- (see Figure 9). All grids that fully lie within the light (NL) based feature extractor performed better boundaries of a microregion are assigned to that than the FMOW. It is what we had expected, as the microregion. Grids that lie within the intersection NL feature extractor is tuned for predicting poverty, of multiple microregions are assigned to the while the other extractor is mostly tuned on high microregion where it has the most intersections. resolution images of the developed world, and is To get the image features from a microregion, built to look for functional forms or uses of land we estimate an average of all the image features (Christie et al., 2018). Hence, for the subsequent corresponding to satellite images that are assigned analysis, we used the features from the NL to that microregion. This gives an aggregated feature extractor. feature vector summarizing the satellite imagery for a microregion. Since both feature extractor models produce a long feature vector (order of 2,000 -4,000-dimensional length), we use principal component analysis (PCA) Estimating and Forecasting 19 Income Poverty and Inequality in Haiti BOX 1: Important Data Considerations and Challenges There are several challenges that result from working with auxiliary data sources, namely satellite imagery and mobile phone data. These datasets, usually, exist at different spatial and temporal granularity. Mobile phone data are available for each subscriber, while environmental data have mixed spatial resolution, from very accurate vector data to low-resolution satellite imagery. The specific challenge in dealing with data heterogeneity is to identify the optimal spatial resolution for the satellite imagery that can capture the signal for development and can be exploited to predict poverty. Since the satellite imagery exists in high-resolution grids, one must decide the resolution which will act as the unit of analysis. We believe it is dependent on the task and the country in question, as all spatial locations within a country will not be homogeneous in containing the signal for development. An example is the highly dense urban areas versus the mostly agriculture-based rural areas. Another consideration is the computational expertise and resources needed to analyze the auxiliary data. If state of the art deep neural network based methods are to be explored, then it is important to keep in mind that training a deep model is a computationally expensive task, both in terms of time, money, and computational expertise and resources availability of computational resources and expertise. While satellite imagery at a certain resolution is publicly available, mobile phone data is always proprietary and is thus not openly available to researchers. It also has privacy concerns, though all analyses in this work have used anonymized and aggregated metadata. Another major challenge is how to validate the model estimates at finer spatial granularity of sections communales. This is owing to the unavailability of ground truth data that can be extracted from census/surveys at that spatial granularity and/or period. A note about bias in the datasets: It is important to note that there are biases in the auxiliary datasets, which might impact the inferences that are drawn. Mobile phone data suffers from selection bias, as it represents only subscribers of a particular telecom provider. Thus, important demographics like children and the ultra-poor (who may not have access to a mobile phone) may not be properly represented in the data. Also, results might be biased towards urban areas, rather than rural ones, due to satellite imagery mostly picking up urbanization signals as well as the access to electricity for mobile phone service. Given that Haiti has a higher urban density than other low-income countries, this should not be as big of a challenge as in other countries. We did see improved performance accuracy when different auxiliary data sources were combined, especially for targets such as average income and FGT indicators. Thus, combining satellite imagery and mobile phone data might assist in mitigating some of the biases existing in these datasets, but further longitudinal studies need to be undertaken to get a conclusive evidence. Finally, our model estimates and uncertainties are also affected by the poor data quality owing to temporal and spatial resolution in satellite imagery. However, better quality of input data should benefit model estimates. 20 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Figure 9: Left - Satellite Images for Haiti for 2014. Right - Road Map of Haiti for 2016 calculated from the mobile phone metadata. Methodology Our research question is expressed as a regression, Each of our individual regression model is of the with the independent variables/covariates being the form: features extracted from the data sources and the dependent variable being the target that we want (3) yi = βT xi + f( xi )+ϵ to estimate: Poverty (in USD), FGT indices, gini, and standards of living deprivations. The targets are where yi is the target value and xi is a vector of calculated using the 2012 ECVMAS and the 2003 independent variables derived from the particular census. data source for the ith microregion. The first term is a linear combination of the independent variables. Our idea is to build individual regression models The function f(xi) models the non-linear relationship on each source of data, and then combine those between yi and xi. The residual term, ϵ, models the predictions in a Bayesian weighted manner. A remaining unexplained noise, and is modeled as a naïve way to combine different data sources is to zero-mean Gaussian random variable, i.e., ϵ~N(0,σn2). concatenate the different feature spaces together, Without the non-linear term, f(xi) in the equation, but such concatenation with few data points leads the model is equivalent to an ordinary linear to overfitting (Xu et al.,2013). We use Gaussian regression. However, instead of assuming a fixed Process regression, which provides uncertainty parametric form for f(), we adopt a non-parametric with its predictions, which are then used to get approach, by assuming a Gaussian Process (GP) the combined predictions. The advantage of this prior on f(). For details on the generative process approach lies in its modularity, as more data sources and the methodology, please refer to the appendix. can be easily added (since we run regression on individual data and their predictions combined). Combining source-specific models Additionally, each data source remains private to its ecosystem, and only the output predictions need To predict poverty metrics for a microregion, we to be shared. This is an important concern when used the model specified in the previous equation. If working with datasets that lies within different only one data source (mobile data or imagery) was private and public entities and has privacy concerns available for this region, then only the model specific (like mobile phone metadata). Estimating and Forecasting 21 Income Poverty and Inequality in Haiti to that data source can be employed. For D data 9 of the Haitian departments and tested on the sources, D independent models are employed. Now, remaining department. Our CV procedure is run 10 for each data source, denoted as d (where d ϵ D), our times, to ensure that all communes are tested. model produces a posterior Gaussian distribution, denoted by yid ~N((yid ) ̅,σid2). The combined poverty It is important to note that, since we do not have estimate, yi, is assumed to be a mixture distribution access to SAEs at the section communale level, we consisting of d Gaussians, defined above. The mixing cannot validate these predictions with the ground weights for a given data source, d (where d ϵ D) and data. This lack of validation of the extrapolatory microregion i, are defined as: power of the poverty estimation models is a 1 challenge as discussed by Rodriguez Castelan et σid2 al., 2019. However, these estimates can be validated (4) wid= when newer surveys are done. ∑v=1 D 1 σ2 iv We use two metrics of evaluation- Mean Absolute Error (MAE) and Root Mean Squared Error The weights assign greater importance to the (RMSE). MAE is the average of the errors for a set source that provides a smaller predictive variance, of predictions. RMSE is square root of the mean signifying higher confidence in the prediction for the squared error. Both MAE and RMSE are in the same particular microregion. The mean and the variance units as the target variable, and lower values of for the combined poverty estimate for a microregion them indicate a better fit. Both are good measures are given as: of how accurately the model predicts the target. However, since in RMSE the errors are squared, its values can be higher if some errors are higher. (5) E[yi ] = ∑v=1 D wiv yiv var[yi ] = The methodology detailed above is summarized in figure 10 below. ∑v=1 D w σ 2 + ∑D ∑D w w iv iv v=1 v'=1 iv iv' Our model produces estimates of poverty and its associated deprivations for 2014 at commune and (yiv ̅ yiv')2 sections communales, as well as forecasts for 2019. The results and maps that are named as 2014, are built using 2014 satellite imagery and mobile phone Out-of-sample generalization using metadata spanning few months for 2016/17/18, and are thus dynamic in nature. We understand that it spatial validation would be better for the model to have concurrent To measure the extrapolation capacity of the mobile phone and satellite imagery, but owing to model on out-of-sample data, we used spatial problems with data acquisition, we decided to cross- validation (CV) techniques, which are more make the best use of data at hand. Also, it should robust (Deville et al., 2014; Bahn and McGill, 2013). be noted that poverty and social deprivations tend In spatial CV, the training and evaluation/test sets to evolve slowly over time, and thus one needs to are taken from geographically distinct regions. be mindful that the input data should not capture Such validation is more robust that standard cross- short-lived variations. The forecasted maps for validation techniques (Blondel et al., 2012; Bahn and 2019 are produced using only satellite imagery McGill, 2013), and thus proves that our model is for 2019, due to the unavailability of recent mobile generalizing well to out-of-sample instances. For phone data. Haiti, during each CV run, our model is trained on 22 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Figure 10: Summary of the methodology for the estimation of social indicators using auxiliary data. Tables 2 and 3 summarize the cross-validated results 2. Linear Model: Linear Regression model. at commune level for each of the targets in 2014, related to poverty and its deprivations. To assess 3. Mobile: It is our GP model, and the independent the efficacy of our model, we experimented with variables are the features extracted from mobile the following settings: phone metadata. 1. Multiview Gaussian Process (GP) Model: This 4. Satellite Imagery: It is our GP model, and the is our GP model, with two auxiliary sources of independent variables are features extracted data- mobile and satellite imagery. using the NL model from the satellite images. Table 2: Mean Absolute Error (MAE) values from spatial cross-validation at commune level. Standards of Model Statistics Average Income Poverty Extreme Poverty Poverty (in USD) Living Deprivation Min 1375.20 0.11 0.06 0.10 0.26 Max 10012.20 0.88 0.79 0.87 1.00 Multiview GP MAE 620.73 (215.72) 0.07 (0.02) 0.10 (0.03) 0.08 (0.03) 0.04 (0.02) Linear MAE 622.57 (160.59) 0.07 (0.02) 0.09 (0.03) 0.08 (0.02) 0.04 (0.01) Image Only MAE 778.47 (430.08) 0.09 (0.04) 0.11 (0.04) 0.09 (0.04) 0.08 (0.07) Mobile Only MAE 641.57 (206.80) 0.08 (0.03 0.10 (0.03) 0.08 (0.03) 0.03 (0.01) Model Statistics FGT1 HTG FGT2 HTG Gini FGT1 USD FGT2 USD Min 14.73 217.11 0.32 0.35 0.12 Max 87.83 7713.61 0.82 2.09 4.38 Multiview GP MAE 6.67 (2.42) 916.10 (323.98) 0.06 (0.03) 0.16 (0.06) 0.52 (0.18) Linear MAE 7.85 (3.14) 1043.55 (369.26) 0.07 (0.03) 0.19 (0.08) 0.59 (0.21) Image Only MAE 6.92 (2.47) 968.12 (318.92) 0.06 (0.02) 0.16 (0.06) 0.55 (0.18) Mobile Only MAE 7.30 (2.95) 1006.47 (399.38) 0.06 (0.03) 0.17 (0.07) 0.57 (0.22) Note: Mean values and the corresponding standard deviation (in parenthesis) over 10 cross-validation runs are reported. Estimating and Forecasting 23 Income Poverty and Inequality in Haiti Table 3: Root Mean Squared Error (RMSE) values from spatial cross-validation at commune level. Standards of Model Statistics Average Income Poverty Extreme Poverty Poverty (in USD) Living Deprivation Min 1375.20 0.11 0.06 0.10 0.26 Max 10012.20 0.88 0.79 0.87 1.00 Multiview GP RMSE 785.97 (269.47) 0.09 (0.03) 0.13 (0.05) 0.10 (0.03) 0.07 (0.07) Linear RMSE 815.12 (199.84) 0.10 (0.03) 0.13 (0.05) 0.10 (0.03) 0.04 (0.02) Image Only RMSE 1159.38 (859.27) 0.13 (0.07) 0.16 (0.09) 0.13 (0.08) 0.13 (0.15) Mobile Only RMSE 742.21 (205.17) 0.09 (0.02) 0.12 (0.03) 0.09 (0.02) 0.04 (0.02) Model Statistics FGT1 HTG FGT2 HTG Gini FGT1 USD FGT2 USD Min 14.73 217.11 0.32 0.35 0.12 Max 87.83 7713.61 0.82 2.09 4.38 Multiview GP RMSE 7.98 (2.94) 1082.58 (350.57) 0.07 (0.03) 0.19 (0.07) 0.62 (0.21) Linear RMSE 9.55 (4.65) 1295.78 (559.63) 0.09 (0.04) 0.23 (0.11) 0.74 (0.32) Image Only RMSE 7.85 (3.06) 1109.18 (392.08) 0.08 (0.03) 0.19 (0.07) 0.63 (0.22) Mobile Only RMSE 9.21 (3.50) 1267.92 (422.18) 0.08 (0.03) 0.22 (0.08) 0.72 (0.24) Note: Mean values and the corresponding standard deviation (in parenthesis) over 10 cross-validation runs are reported. Our model does consistently better than the linear than using either of the datasets separately (mobile model, across all the targets. Figure 11 highlights or imagery). For some of the targets, like income the performance of our model, which maps the poverty, combining the different sources of data is non-linear relationships between the auxiliary data beneficial. sources and the targets. Our model also does better Figure 11: From Left: a) Denotes the comparison of actual and predicted Gini index values for all communes for 2014, b) Denotes the comparison of actual and predicted poverty (USD) values for all communes for 2014 Model fit for predicting Gini index Model fit for predicting poverty (USD) 1.0 1.0 Quartile 1 Quartile 3 High Quartile 1 Quartile 3 High Quartile 2 Quartile 4 Low Quartile 2 Quartile 4 Low 0.8 0.8 Predicted using Auxiliary Data Predicted using Auxiliary Data 0.6 0.6 0.4 0.4 0.2 0.2 0.0 0.0 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 Estimated from Census Estimated from Census 24 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Figures 12, 13, 14, and 15 display the cross-validated Figure 12 (left pane) shows the predicted income values of different social indicators of interest poverty at the commune level. The average commune compared with the ground data (SAEs estimated has a poverty level exceeding 66%. The commune from the census and ECVMAS). There are regions with the lowest level of predicted income poverty (like those around Port-au-Prince) where our was Port-au-Prince in the Ouest department (with model predicts values close to the ones derived 11.8% of income poverty), while the commune with from census/surveys (the SAE, or the ground- the highest predicted level of poverty was Cornillon/ truth). However, there are also regions that feature Grand Bois in the Ouest department (80.9%). Only marked differences between the predictions and eight out of the 13826 examined communes for the ground truth, which calls for further research25. the income poverty prediction has poverty levels Nonetheless, on average, our predicted commune below 50%. High poverty is concentrated in regions estimates had a different of only 1 percentage point within departments of Nord-Est and Centre in 2014. difference with the ground-truth data for the income However, the section communale maps (which will poverty and standards of living deprivation, and a be later discussed, starting on figure 16) confirm 0.5 percentage point difference with the ground- that there is heterogeneity in the spatial dispersion truth, in the case of the income inequality estimates. of poverty and associated deprivations27. Figure 12: 2014 Dynamic Maps of income poverty (in USD) at the commune level. The maps display the quantiles of the predicted (left) vs. the ground-truth data (right) Mean Median St. deviation Minimum Maximum 66.4% 68.1% 8.7% 11.8% 80.9% Note: Please notice that the values for each color range slightly vary between the predicted and ground-data maps. 25 Tables A.2, A.3, and A.4 in the annex list the communes with more than 10 percentage point difference between the ground truth and the predicted values. 26 While the model was estimated for 140 communes, it failed to estimate the income poverty values for Delmas and Tabarre. 27 In fact, within the least poor department, the Ouest department, there are several microregions registering income poverty exceeding 50%. Estimating and Forecasting 25 Income Poverty and Inequality in Haiti Figure 13 shows the predicted and ground truth Ouest department. The commune with the lowest gini coefficients for Haitian communes. The average predicted gini coefficient is Delmas, while the commune has a predicted gini coefficient of 0.639. one with the highest predicted gini coefficient is Similar to the case of the SAE, our predicted Cornillon/Grand Bois, both of these in the Ouest data shows that only six communes have a gini department. coefficient of under 0.50, most of these in the Figure 13: 2014 Dynamic Maps of Gini Index at commune level. The maps display the quantiles of the predicted (left) vs. the ground-truth data (right). Mean Median St. deviation Minimum Maximum 0.639 0.648 0.066 0.338 0.763 Figure 14: 2014 Dynamic Maps of FGT2 Index (in USD) at commune level. The maps display the quantiles of the predicted (left) vs. the ground-truth data (right). 26 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Figure 15 shows the predicted and ground-truth to have a 91.3% deprivation. The least deprived values for the Standards of Living deprivation at commune is Port-au-Prince28 (18.9%), while the the commune level. Just as with the SAE, most most deprived one is Vallières in the Nord-Est communes register high levels of standards of living department (100%). deprivation. The average commune is predicted Figure 15: 2014 Dynamic Maps of Standards of Living Deprivation at commune level. The maps display the quantiles of the predicted (left) vs. the ground-truth data (right). Mean Median St. deviation Minimum Maximum 91.3% 93.7% 10.5% 18.9% 100% Maps at the section Figure 16: 2014 Dynamic Map of Poverty communale level (in USD) at the section communale level Figures 16-19 show the maps for the same four social indicators at the section communale level. These maps highlight the heterogeneity of socioeconomic deprivations that get masked when taking spatially aggregated statistics. However, as mentioned before, given that maps at this level of disaggregation have not been formally validated, these should be interpreted with caution29 and are only presented to provide evidence of the heterogeneity within most communes. 28 Notice however that Port-au-Prince is among the communes with the highest difference between the ground truth and estimates. The second least deprived commune is Pétionville (46.1%), followed by Croix-des-Bouquets (50%), both in Ouest. 29 For this reason, no interpretation of the values or identification of the most affected sections communales will be provided Estimating and Forecasting 27 Income Poverty and Inequality in Haiti Figure 17: 2014 Dynamic Map of Gini Index at the section communale level Figure 18: 2014 Dynamic Map of FGT2 index (in USD) at the section communale level Figure 19: 2014 Dynamic Map of Standard of Living Deprivations at the section communale level 28 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 A note on dataset uncertainty Figure 20 show the uncertainty associated with power, which is also reported in Pokhriyal and each dataset, which can assist policymakers when Jacques (2017). We also noticed a strong spatial deciding which dataset is better in making poverty correlation, which is understandable as poverty and predictions for a given microregion. We see that its associated deprivations are spatially correlated. mobile phone metadata holds better predictive Figure 20: Uncertainty associated with each data source evidenced by the most accurate one for income poverty (in USD) predictions (left) and standards of living deprivations (right) for 2014 (commune level). (a) (b) Forecasting 2019 deprivations with 2019 satellite imagery Since both the satellite and mobile phone metadata and aggregated them to commune level, as well as are regularly available, the resulting model could be the section communale level. updated several times in-between the household surveys. To test the forecasting ability, we check how We used our model that is trained on 2014 satellite a model that is trained on learning the relationship imagery at commune level. To predict the poverty between geospatial covariates and poverty for a and associated deprivations at the commune specific point in time (say 2014) performs when level in 2019, we provide our model with the 2019 newer data (say 2019) is inputted. Given that we satellite features aggregated at commune level did not have access to more recent mobile data, the during testing time. Similarly, to predict poverty and 2019 maps were created exclusively with satellite associated deprivations at the section communale imagery (specifically, using Google API satellite level, we test/evaluate our model with aggregated images for Haiti from 2019) (Gorelick et al., 2017). section communale data for 2019. Figures 21-24 Each of the images was available at zoom level 16, show the resulting maps at both microregional which is equivalent to approximately 1 x 1 sq. km levels. Once again, notice the heterogeneity within on the ground. Just like with the 2014 images, we every commune. used the night-light-based feature extractor to get features corresponding to each satellite imagery Estimating and Forecasting 29 Income Poverty and Inequality in Haiti Figure 21: Income poverty (in USD) at the commune level (left) and section communale level (right) in 2019. Note: notice that the colors and quantiles vary between maps. Figure 22: Gini index at the commune level (left) and section communale level (right) in 2019. Note: notice that the colors and ranges vary between maps. 30 Estimating and Forecasting Income Poverty and Inequality in Haiti Methodology and Results 4 Figure 23: FGT2 (in USD) at the commune level (left) and section communale level (right) in 2019. Note: notice that the colors and ranges vary between maps. Figure 24: Standards of living deprivations at the commune level (left) and section communale level (right) in 2019. Note: notice that the colors and ranges vary between maps. Estimating and Forecasting 31 Income Poverty and Inequality in Haiti It is important to note that the forecasted maps 2. Nord-Ouest communes increasingly became from 2019 must be interpreted with caution, as poorer and more deprived in 2019 relative to these have not been validated by SAEs and were 2014, when the worst-performing communes only elaborated with satellite imagery (not the were mainly located in Nord-Est and Centre. mobile data). As a reminder, the combination This is consistent with the latest findings of of both datasets seems to perform best than the World Food Programme’s 2020 Global images alone. Report on Food Crisis, which shows that the population in the Nord-Ouest department is the Nonetheless, despite the lack of validation for these one undergoing the most urgent food insecurity last maps, several general trends can be observed levels (IPC Phase 4). when comparing both groups of maps: 3. Some communes in the Sud Department 1. Better-performing communes were usually became more deprived in 2019 relative 2014. located in the Ouest Department in both 2014 This may be due to the impact that Hurricane and 2019. The following communes consistently Matthew had on these communes in late 2016. perform better than the rest in terms of poverty, inequality and social deprivations: Pétion-Ville, Port-au-Prince, Delmas, Cap-Haitïen and Tabarre. 32 Estimating and Forecasting Income Poverty and Inequality in Haiti Policies to tackle social deprivations in the COVID-19 era and the recovery period 5 5 be 10.9% of GDP30, and the Ministry of Economy and Finance expects it to be in the order of 11.6% of GDP in FY2020. Haiti’s tax system also tends to be regressive, relying heavily on indirect taxes. In FY2019, income taxes amounted to only 26% of total tax revenues, while the TCA (sales tax) and custom taxes were equivalent to 28% and 26.3%, respectively31. Policies to tackle Further, Haiti’s health and education spending are relatively low: according to Ghayad et al., 2019, social deprivations Haiti’s spending in these two categories summed up to less than 4% of GDP in FY2018, well below the regional average (over 8% of GDP). With the arrival in the COVID-19 of COVID-19, the Haitian government increased its health spending allocation by a factor of four, era and the bringing the budget allocation of the Ministry of Public Health and Population (MSPP) to 10.9% of recovery period the total FY2020 budget. In terms of social safety nets, several efforts toward creating social programs have been rolled out The preceding sections have clearly demonstrated over the past five years, yet the social protection that a large portion of the Haitian population was system has remained notably fragmented, and have in a vulnerable position prior to the arrival of the consisted mainly of tuition assistance and nutritional COVID-19 pandemic. The crisis associated with this support programs funded by various development pandemic will further deepen social disparities in and humanitarian organization (Ghayad et al., 2019). Haiti if no comprehensive social protection plan In 2014, an Action Plan for Reduction of Poverty is introduced. According to the IMF (Ghayad et (PAARP) was launched, which introduced a set of al., 2019), to promote inclusive growth, a country social programs under a social assistance umbrella has three tools at its disposal: progressive income called Ede Pep32, and included a national fee waiver taxation, health and education spending, and a program for basic education (PSUGO). In 2016, a social safety net. Nonetheless, Haiti’s tax collection Sectorial Table of Social Protection was introduced is low to begin with: in FY2019, it was estimated to by the Ministry of Social Affairs and Labor (MAST), 30 This figure has been falling from its value in the previous years, as it was, on average, 13.4% of GDP between FY2015-2018. 31 One could argue that this could be related to the socioeconomic crisis in FY2019, yet this regressivity has also been present over (at least) the ten preceding years. 32 In total, EDE PEP includes 11 social programs to help mothers of schoolchildren and university students, vulnerable and food insecure households, and farmers and households affected by natural disasters. Estimating and Forecasting 33 Income Poverty and Inequality in Haiti which consisted of consultations with stakeholders wallet. These challenges must be tackled in order nationwide in order to draft a national social to ensure the effective distribution of the PNPPS protection policy. In addition, the National Strategic cash transfers. Development Plan (PDSH) — which intended to transform Haiti into an emerging country by 2030 — Further, the following actions should be has a social refoundation pillar which contemplates considered when preparing to roll out the PNPPS: modern health and education networks accessible (i) the expansion and update of the Information to all Haitians, as well as policies of social inclusion System (SIMAST) of the MAST so that there is a and gender equality. comprehensive source of information that can be consulted to locate beneficiaries. Our poverty More recently, according to the 2020 Article maps could be used as a first step to try to locate IV Staff Report of the IMF with Haiti, a National beneficiaries at a more disaggregate level (ii) Plan for Social Protection and Progress (PNPPS) the identification of clear and reliable means of was being finalized. In this staff report, the distribution or payments depending on the type IMF recommends that the PNPPS reduces the of benefit, beneficiaries and their location (e.g. fragmentation and overlap of existing programs, mobile payments might work in the city, but in the and that a limited number of unconditional, quasi remote areas, for instance, it would be necessary to universal cash transfer programs for vulnerable distribute cash through local credit unions); and (iii) groups is established. Cash transfer programs that a clear articulation of the role of the various actors are simple in design have proven to be effective for each step of the intervention, from the definition in other low-income countries and have the of the benefits to the deployment of these, as Haiti potential to be introduced in a quicker fashion than currently relies on a variety of actors, including conditional cash transfers. Further, Banerjee et al., financial operators and NGOs. 2019 demonstrated that more refined targeting programs tend to be less effective in countries Finally, it will be important to consider making unable to identify beneficiary or lacking capacity to complementary investments such as transport implement such targeted programs, thus providing infrastructure, in order to reduce the high cost evidence in favor of simplified, unconditional cash of mobility for Haitians living in remote areas. transfers in countries like Haiti. This will be key in places such as the Nord-Ouest department, which, as shown in figure 9 above, With the arrival of COVID-19, the launching of is practically disconnected from the rest of the a set of unconditioned cash transfer programs country. According to Hausmann (2015), there are under the PNPPS umbrella is more essential than enormous productivity differences among regions ever to provide quick relief to the most vulnerable within several countries. These differences are segments of the population. As part of the measures explained by the fact that some regions within to mitigate the impact of COVID-19 on the most a country are better endowed with inputs for vulnerable, the Haitian government distributed productivity to take place (such as electricity, unconditional cash transfers33 (via a mobile payment more readily available forms of transportation, and system) and food kits to 1.5 million vulnerable the convergence of different talents). Therefore, households. Nonetheless, the distribution of the the poor are potentially being excluded from the cash transfers faced several challenges, including higher productivity areas of the country. Investing the lack of a comprehensive and accurate list of the on quality infrastructure will allow for Haitians to beneficiaries’ phone numbers, the fact that many connect to the rest of the productive sectors, will target beneficiaries’ telecom provider was not linked result in productivity gains, and could ultimately be to the mobile payment system34, and the fact that associated with a more inclusive growth. a number of beneficiaries had no active mobile 33 Equivalent to about US$30. 34 Only one of the two mobile carriers (Digicel) is connected to the mobile payment system. 34 Estimating and Forecasting Income Poverty and Inequality in Haiti 6 Conclusion Given the high costs associated with collecting and the section communale levels. The results from this process are consistent with the SAEs, although only the 2014 dynamic maps at the commune level were validated with the ground-truth data. Similar household survey data, reporting poverty and to the SAEs, the 2014 validated estimates show other social indicators on a regular basis is a that only eight out of the communes for which challenge in several developing countries. The predictions were available had poverty levels below present report adds to the poverty literature on 50%, and that only six out of the 140 communes Haiti by disaggregating 2012 estimates for income for which predictions were available had gini poverty, income inequality, and standard of living coefficients under 0.50. Further, high standards of deprivations at the commune level (previously, living deprivations continued to be a trend present and to the best of our knowledge, these were only in the 2014 dynamic predictions. available at the departmental level). These estimates where then used to validate the prediction of these The section communale maps, which should be values for 2014, which were estimated using features interpreted with caution due to the lack of validation, extracted from aerial imagery and anonymized call show high heterogeneity of income poverty, income detailed records straddling 2016, 2017, and 2018 inequality, and standard of living deprivations within (thus making the 2014 values dynamic in nature). each commune. Further, when comparing the 2014 dynamic maps with the 2019 maps (the latter which The 2012 disaggregate estimates (called Small Area also lack validation), the following general trends Estimates or SAEs throughout the report) were were observed: a) Better-performing communes obtained using a model based using the Gradient were usually located in the Ouest Department in Boosting Machine technique. The national poverty both 2014 and 2019. These communes were: Pétion- rate estimated from this model is 57.1%, which is very Ville, Port-au-Prince, Delmas, Cap-Haitïen and close to the 2012 official national poverty rate using Tabarre ; b) Nord-Ouest communes increasingly department level data (58.5%). The SAEs allowed became poorer and more deprived than the rest in for us to determine that one out of four Haitian 2019 relative to 2014, when the worst-performing living in poverty resided in ten communes (out of communes were mainly located in Nord-Est the 140 examined communes): Gonaïves, Cité Soleil, and Centre; and c) some communes in the Sud Port-au-Prince, Saint-Marc, Cap-Haïtien, Carrefour, Department became more deprived than the rest Dessalines, Petite Riviere, Saint-Michel de l’Attalaye in 2019 relative 2014. This may be due to the impact and Port-de-Paix. Further, there is a concentration that Hurricane Matthew had on these communes of communes with very high levels of poverty in in late 2016. the eastern side of Artibonite, the southernmost communes in the Sud and Sud-Est departments, Our maps provide evidence that Haiti needs a and in the Centre department. Moreover, Among comprehensive plan for inclusive growth, but this the 10 communes with the highest poverty rate in will require that more resources be mobilized 2012, six were located in the northern region. The toward social spending and that the Haitian tax communes with the most intense poverty levels in system is made more progressive. More importantly, 2012 (measured by the FGT2 index) were located in in the context of COVID-19, the rolling out of a set the Sud and Artibonite departments. The communes of simple, unconditioned cash transfers (via the with less poverty intensity were those located in PNPPS) is more urgent than ever in order to mitigate the Ouest department (Port-au-Prince, Pétion-Ville, the disproportionate effect that the pandemic will Carrefour, Delmas) and Cap-Haïtien (Nord). have on the most vulnerable. Nonetheless, several logistical challenges must be tackled first to Income inequality is rampant in Haiti, with only guarantee the timely delivery of these transfers. The three out of the 140 communes examined in this poverty and inequality maps presented in this report exercise with a gini coefficient of less than 0.50. represent a powerful tool for the identification Furthermore, in terms of standards of living of the most vulnerable at a more disaggregate deprivations, most communes registered high level and provide the framework for further levels of deprivation (over 90%, on average). The update with more recent data and for other social communes with the least deprivations were mostly protection applications. Finally, it will be important located in the Ouest department (Delmas, Port-au- to consider important infrastructure investments in Prince, Carrefour, Pétion-Ville, and Tabarre). order to better connect some of the most remote areas in Haiti to more productive regions of the The 2014 dynamic maps and the 2019 predictions country. All else constant, such actions will result were estimated using a model based on uncertainty- in productivity gains and thus, in a more inclusive weighted multi-view Gaussian process regressions. growth in Haiti. These predictions were presented at the commune Estimating and Forecasting 35 Income Poverty and Inequality in Haiti Bibliography Alderman, H., Babita, M., Demombynes, G., Makhatha, N. & Özler, B. (2002). How low can you go? Combining census and survey data for mapping poverty in South Africa. J. Afr. Econom., 11(2), 169– 200. Bahn, V. and B. J. McGill. Testing the predictive performance of distribution models. Oikos, 122(3): 321–331, 2013. Banerjee, Abhijit, Paul Niehaus, and Tavneet Suri. 2019. Universal basic income in the developing world. Annual Review of Economics 11. BBS & UNWFP. Local Estimation of Poverty and Malnutrition in Bangladesh. Dhaka: The Bangladesh Bureau of Statistics and The United Nations World Food Programme. 2004. Blondel, V.D., M. Esch, C. Chan, F. Cl´erot, P. Deville, E. Huens, F. Morlot, Z. Smoreda, and C. Ziemlicki. Data for development: the d4d challenge on mobile phone data. arXiv preprint arXiv:1210.0137, 2012. Blumenstock, J., G. Cadamuro, and R. On. Predicting poverty and wealth from mobile phone metadata. Science, 350(6264):1073{1076, 2015. Chambers, R. & Tzavidis, N. (2006). M-quantile models for small area estimation. Biometrika, 93(2), 255– 268. Christie, G., N. Fendley, J. Wilson, and R. Mukherjee. Functional map of the world. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. Cressie, N. The origins of kriging. Mathematical Geology, 22(3), 1990. Das, S. and Haslett, S., 2019. A Comparison of Methods for Poverty Estimation in Developing Countries. International Statistical Review, 87(2), pp.368-392. Demombynes, G., Lanjouw, J.O., Lanjouw, P. and Elbers (2006) “How Good a Map? Putting Small Area Estimation to the Test”, Policy Research Working Paper No. 4155, The World Bank. Deville, P., C. Linard, S. Martin, M. Gilbert, F. R. Stevens, A. E. Gaughan, V. D. Blondel, and A. J. Tatem. Dynamic population mapping using mobile phone data. Proceedings of the National Academy of Sciences, 111(45):15888–15893, 2014. Elbers, C. & Lanjouw, J. O. & Lanjouw, P., (2002). "Micro-level estimation of welfare," Policy Research Working Paper Series 2911, The World Bank. Elbers, C., Lanjouw, J.O., Lanjouw, P. & Leite, P.G. (2004). Poverty and inequality in Brazil: new estimates from combined PPV-PNAD. In Inequality and Economic Development in Brazil, pp. 81–104. Washington, D.C.: World Bank. Elbers, Chris, Peter F. Lanjouw, and Phillippe G. Leite. "Brazil within Brazil: Testing the poverty map methodology in Minas Gerais." World Bank Policy Research Working Paper Series, Vol (2008). 36 Estimating and Forecasting Income Poverty and Inequality in Haiti Bibliography Elvidge, C. D., Sutton, P. C., Ghosh, T., Tuttle, B. T., Baugh, K. E., Bhaduri, B., & Bright, E. (2009). A global poverty map derived from satellite data. Computers & Geosciences, 35(8), 1652-1660. Engstrom, R.; J. Hersh, and D. Newhouse. Poverty from space: using high-resolution satellite imagery for estimating economic well-being. Policy Research working paper, World Bank, 2017. Fernandez-Taranco, O., D. Keita, P. Hiebra, J. Le Nay, A. Ambroise, P. Rouzier,C. Cadet. Haiti and Governance for Human Development. 2013. United Nations Development Programme. Flowminder final consulting report: Caracol Industrial Park and National Road Network. An analysis of commuting and migration patterns (Mimeo) Fujii, T. (2004). Commune-Level Estimation of Poverty Measures and its Application in Cambodia, Vol. 2004/48 Helsinki: UNU-WIDER, United Nations University (UNU). Giusti, C., Marchetti, S., Pratesi, M. & Tzavidis, N. (2011). Small area methodologies for poverty estimation: an application to Italian data. In International Statistical Institute: Proceedings of the 58th World Statistical Congress, Dublin, pp. 6234– 6237.Haslett, S. & Jones, G. Estimation of Local Poverty in the Philippines. Philippines National Statistics Co-ordination Board / World Bank, November 2005. Ghayad, R., F. Lambert, M. Rousset, and M. Bellon (2019). Haiti Selected Issues: IMF Country Report No. 20/122; December 23, 2019 Gorelick, N.et al. Google earth engine: Planetary-scale geospatial analysis for everyone. in Remote Sensing of Environment (Elsevier, 2017). Haslett, S. & Jones, G. Small Area Estimation of Poverty, Caloric Intake and Malnutrition in Nepal. Kathmandu: Nepal Central Bureau of Statistics/World Food Programme, United Nations/World Bank. 2006. Haslett, S., Jones, G. & Isidro, M. (2014). Small-area Estimation of Child Undernutrition in Bangladesh. Dhaka: Bangladesh Bureau of Statistics, United Nations World Food Programme and International Fund for Agricultural Development. ISBN 978-984-33-9085-1. Hausmann, R., 2015. What Should We Do About Inequality?. Copy at http://www.tinyurl.com/y6byadd3 Head, A., M. Manguin, N. Tran, and J. E. Blumenstock. Can human development be measured with satellite imagery? In Proceedings of the Ninth International Conference on Information and Communication Technologies and Development, ICTD ’17, 2017. Healy, A.J., Hitsuchon, S. and Vajaragupta, Y., 2003. Spatially disaggregated estimates of poverty and inequality in Thailand. Massachusetts Institute of Technology and Thailand Development Research Institute. Hersh, J., R. Engstrom, M. Mann, A. Mejia and L. Martin. Mapping Income Poverty in Belize Using Satellite Features and Machine Learning. Inter-American Development Bank. 2020. Huang, G., Z. Liu, and K. Q. Weinberger. Densely connected convolutional networks. CoRR, 2016. Jean, N., M. Burke, M. Xie, W. M. Davis, D. B. Lobell, and S. Ermon. Combining satellite imagery and machine learning to predict poverty. Science, 353(6301):790-794, 2016. Kilic, Talip; Serajuddin, Umar; Uematsu, Hiroki; Yoshida, Nobuo. 2017. Costing household surveys for monitoring progress toward ending extreme poverty and boosting shared prosperity (English). Policy Research working paper; no. WPS 7951; LSMS. Washington, D.C.: World Bank Group. http://documents. worldbank.org/curated/en/260501485264312208/Costing-household-surveys-for-monitoring-progress- toward-ending-extreme-poverty-and-boosting-shared-prosperity Estimating and Forecasting 37 Income Poverty and Inequality in Haiti Bibliography Marchetti, S., Beresewicz, M., Salvati, N.S. & Wawrowski, L. (2018). The use of a three-level M-quantile model to map poverty at local administrative unit 1 in Poland. J. R. Stat. Soc.: Ser. A (Stat. Soc.), 181(3), 1077– 1104. Molina, I. & Rao, J.N. (2010). Small area estimation of poverty indicators. Can. J. Stat., 38(3), 369– 385. Pandey, S. M., T. Agarwal, and N. C. Krishnan. Multi-task deep learning for predicting poverty from satellite images. In IAAI, 2018. Pokhriyal, N. and W. Dong. Virtual network and poverty analysis in Senegal. D4D Challenge Senegal Scientific Papers, Netmob, 2015. Pokhriyal N., and D.C. Jacques. Combining disparate data sources for improved poverty prediction and mapping. Proceedings of the National Academy of Sciences, 2017. Rasmussen, C.E. and C. K. I. Williams. Gaussian Processes for Machine Learning. The MIT Press, 2006. Rodríguez Castelán, C., Ingmar Weber, Damien Jacques, and Trevor Monroe. “Making a better poverty map”. World Bank. 2019. https://blogs.worldbank.org/opendata/making-better-poverty-map. Santos, M. & Villatoro, P. (2018): A Multidimensional Poverty Index for Latin America. Review of Income and Wealth, Vol. 64, Issue 1, pp. 52-82, 2018 Solt, Frederick. 2019. “Measuring Income Inequality Across Countries and Over Time: The Standardized World Income Inequality Database.” SWIID Version 8.2, November 2019. Soto, V., V. Frias-Martinez, J. Virseda, and E. Frias-Martinez. Prediction of socioeconomic levels using cell phone records. In Proceedings of the 19th International Conference on User Modeling, Adaption and Personalization, pages 377–388. Springer, 2011. Steele, J.E., P. R. Sundsøy, C. Pezzulo, V. A. Alegana, T. J. Bird, J. Blumenstock, J. Bjelland, K. Engø-Monsen, Y.-A. de Montjoye, A. M. Iqbal, K. N. Hadiuzzaman, X. Lu, E. Wetter, A. J. Tatem, and L. Bengtsson. Mapping poverty using mobile phone and satellite data. Journal of The Royal Society Interface, 14(127), 2017. Tang, B. V., Y. Sun, Y. Liu, and D. S. Matteson. Dynamic poverty prediction with vegetation index. Workshop on Modeling and Decision-Making in the Spatiotemporal Domain, 32nd Conference on Neural Information Processing Systems. 2018. Tzavidis, N., Salvati, N., Pratesi, M. & Chambers, R. (2008). M-quantile models with application to poverty mapping. Stat. Methods Appl., 17(3), 393– 411. Watmough, G. R., C. L. J. Marcinko, C. Sullivan, K. Tschirhart, P. K. Mutuo, C. A. Palm, and J.-C. Svenning. Socioecologically informed use of remote sensing data to predict rural household poverty. Proceedings of the National Academy of Sciences, 2019. World Bank Group. Investing in people to fight poverty in Haiti. 2014. World Food Programme. 2020 - Global Report on Food Crises. https://www.wfp.org/publications/2020- global-report-food-crises United Nations Development Programme. 2019. “Global Multidimensional Poverty Index 2019: Illuminating Inequalities”. http://hdr.undp.org/en/2019-MPI Xu, C., D. Tao, and C. Xu. A survey on multi-view learning. arXiv, abs/1304.5634, 2013 38 Estimating and Forecasting Income Poverty and Inequality in Haiti Appendix Table A.1 List of Common Variables Home ID Access to public water service network Haitian Department (categorical variable ranging from 1 Access to treated water to 10) Total household income Water is drawn from wells Final official weight Water is obtained from rain collection Family income equivalent home Water is obtained from other sources Total income of the household with imputations Water is obtained from river Natural logarithm of the total income of the household Dependency ratio 1. Total people of non-working age as with imputations a percent of the total people of working age Family income equivalent with imputations Dependency ratio 2. People who do not work as a percent of the people of working age who are currently working Natural logarithm of household income equivalent Number of women at home charges Persons per household Number of children at home Overcrowding. Persons per household on all rooms of a Number of people working in the public sector dwelling Thatched roof Number of people working Cement roof Number of people working in a cooperative Plastic roof Number of people working in an NGO Tin Roof Number of people working in other jobs Tile roof Number of teenagers at home Other type of roof Number of children at home Wooden walls Number of people with elementary school education Mud walls Number of people who are employers at their jobs Concrete walls Number of people who are self-employed Plate walls Number of people who are home helpers Cardboard walls Number of people who are interns Brick walls Number of people who are on a payroll Glisse walls Number of people with high school level education Other type of walls Number of people without education Wooden floor Number of people with higher education Estimating and Forecasting 39 Income Poverty and Inequality in Haiti Earthen floor Number of people who are literate Concrete floor Number of people who are studying tiled floor Number of people over 65 who do not work Ceramic floor Number of household members over 15 who do not work Other type of floor Average age of home occupants Kay house The head of the household in women slum (house built with materials from other buildings) There is a spouse at home Ajoupa (hut) Head of the household has primary education low-rise house Head of the household has tertiary education One-story house Head of the household has secondary education Apartment Spouse has primary education Villa house Spouse has tertiary education Other type of dwelling Spouse has secondary school education Access to public waste disposal collection services Age of the head of household Waste disposal through a private service Head of the household works as a civil servant Waste disposal in vacant lot Head of the household works as a domestic worker Waste disposal in ravines Head of the household works in a cooperative Waste disposal into sewers Head of the household works in an NGO Waste disposal on public roads Head of the household works in another place Waste disposal into the sea Spouse works as a civil servant Waste incineration Spouse works as a domestic worker Cooks with wood Spouse works in a cooperative Cooks with propane Spouse works in an NGO Cooks with electric stove Spouse works in another place Cooks with kerosene Head of the household is a salaried worker Cooks with charcoal Head of the household is an employer Cooks with solar oven Head of household is self-employed Cooks with other methods Head of the household is a family assistant Home owner Head of the household is intern Tenant Spouse is a salaried worker Home is a farm Spouse is an employer Inhabits a home for free Spouse is self-employed De-facto occupant Spouse is assistant and self-employed Other type of house occupant Spouse is an intern Number of rooms per dwelling Type of locality Unknown ID Weight Individual weighting per home. Under 12 worth .5 and 13 Post stratification to 18 are worth .75 The water comes from private source 40 Estimating and Forecasting Income Poverty and Inequality in Haiti Appendix BOX A.1: An explanation of Call Detailed Records (CDR) The data used for the estimation of our maps was provided by Flowminder, a non-profit organization that collects, aggregates, integrates and analyzes anonymous mobile operator data, satellite and household survey data across several countries. This organization was hired to elaborate a report for the IDB on commuting and migration patterns of the Haitian population using Digicel call detailed records (CDR). According to Flowminder, CDR data is generated by telecommunications providers and are the basis for the customer billing process. An operator’s network consists of a group of base stations which route signals from an initiating mobile device to a receiving mobile device. A single base station can consist of multiple cells (receivers) operating at different mobile technology generations (For instance: 3G and 4G). A base station’s geographic location is usually fixed, but on some occasions, specially designed, mobile base stations – also known as cells on wheels - can be used to provide extra capacity at particular locations. A base station can only receive signals from mobile devices in a constrained region around its physical location. The spatial density of base stations can vary widely. In Port-Au-Prince, this density can reach up to six base stations per square kilometer, while in rural areas, this density can be significantly lower with a single base station covering many square kilometers. A mobile device is uniquely identified by its Mobile Station International Subscriber Directory Number (MSISDN) and International Mobile Subscriber Identifier (IMSI). In order to bill its customers, a telecommunication operator maintains a database of CDR. Every voice call results in the following information stored in the database: � Start time of the call � Duration of the call � MSISDN of initiating party � Location ID of cell used by initiating party � MSISDN of receiving party � Location ID of cell used by receiving party. Sometimes, additional information such as the handset type is included. Due to the highly private nature of CDR data, Flowminder employed a series of anonymization techniques to reduce the risk of individual or small group disclosure. One of these measures was to “hash” unique identifiers (such as MSISDN). The hashing process was undertaken by telecom staff before Flowminder had access to the data. The infrastructure used for analysis was securely maintained behind telecommunication operator firewalls to reduce the risk of unauthorized access to the raw data. Estimating and Forecasting 41 Income Poverty and Inequality in Haiti Table A.2: Communes with more than 10 percentage point difference between the ground truth and the estimates (income poverty) Deparment Commune Difference Sud Île à Vache 0.37 Artibonite Grande Saline 0.27 Nord-Ouest Baie de Henne 0.19 Artibonite L'Estère 0.15 Artibonite Terre Neuve 0.14 Sud Torbeck 0.14 Artibonite La Chapelle 0.14 Sud Saint louis du Sud 0.14 Sud Chantal 0.13 Centre Saut d'Eau 0.12 Nord Plaisance 0.12 Sud Saint Jean du Sud 0.12 Ouest Cité Soleil 0.12 Artibonite Saint-Michel de l'Attalaye 0.12 Sud Cavaillon 0.11 Nord Pignon 0.11 Nord Bahon 0.11 Nord-Est Capotille 0.11 Table A.3: Communes with more than 10 Table A.4: Communes with more than 10 percentage point difference between the percentage point difference between the ground truth and the estimates ground truth and the estimates (gini coefficient) (standard of living deprivation) Deparment Commune Difference Deparment Commune Difference Sud Île à Vache 0.35 Ouest Delmas 0.26 Artibonite Grande Saline 0.20 Ouest Port-au-Prince 0.22 Sud Torbeck 0.14 Artibonite Grande Saline 0.19 Sud Chantal 0.13 Ouest Croix-Des-Bouquets 0.18 Nord Bahon 0.13 Ouest Cité Soleil 0.12 Artibonite La Chapelle 0.13 Sud Île à Vache 0.11 Nord Plaisance 0.13 Sud Saint louis du Sud 0.12 Nord Pignon 0.11 Sud Port-Salut 0.11 Sud Saint Jean du Sud 0.11 42 Estimating and Forecasting Income Poverty and Inequality in Haiti Appendix Details of the Methodology The data generative process for our Gaussian Process based regression given as: f(x) ∼ GP(m(x),k(x,x’)) yi ∼ N(βTxi + f(xi),σn2),∀i A GP is a stochastic process, indexed by x ∈ Rd. Any finite sample generated from it is jointly multivariate normal (Rasmussen and Williams, 2006). m(x) is the mean of f(x) and k(x,x’) is a kernel function that defines the covariance between any two evaluations of f(x), i.e., m(x) = E[f(x)], and k(x,x’) = E[(f(x)− m(x)) (f(x’)−m(x’))]. For model simplicity, we assume that m(x) = 0, which is a standard practice in GP based methods. Given a training set of examples, , the GP prior on f(), and other terms in the equation, the posterior distribution of y∗ (for an unseen input vector, x∗), is a Gaussian distribution, with the following mean and variance: y̅ * =E[y*]=βT x+kT (K + σn2 I)-1 y σ∗2 := var[y∗] = k∗ −kT(K + σn2I)−1k + σn2 Here, y = [y1,y2,...]T, and K is a matrix which contains the kernel function evaluation on each pair of training inputs, i.e., K[i,j] = k(xi,xj), k is a vector of the kernel computation between each training input and the test input, i.e., k[i] = k(x∗,xi), k∗ = k(x∗,x∗), and I is an identity matrix. Choice of kernel function The role of the kernel function is to specify how the function values, f(x) and f(x’), vary as the function of their corresponding inputs, x and x’. We use the following kernel function: ‖x-x'‖2 ‖xs-xs'‖2 k(x,x')= σf2 exp (-‖x-x'‖^) exp(-‖x-x'‖^ ) 2l 2 2l2s where xs and x’s are the spatial coordinates (latitude, longitude) of the commune centers corresponding to x and x’, respectively. The first exponent term captures non-linear dependencies in the feature space. The second exponent term plays the same role, but in the geographic space, and it models the spatial autocorrelation as a continuous function, which is the same as Kriging, a widely used methods in geostatistics (Cressie, 1990). The parameter σf 2 is the variance of the stochastic process f, l is the process length-scale for the feature space part, and ls is the process length-scale for the spatial part. The quantities β,l,ls,σn2, and σf2 are estimated by maximizing the marginalized log-likelihood of the training data. Estimating and Forecasting 43 Income Poverty and Inequality in Haiti Estimating and Forecasting Income Poverty and Inequality in Haiti using Satellite Imagery and Mobile Phone data June 2020 Neeti Pokhriyal | Omar Zambrano | Jennifer Linares | Hugo Hernández