Data pre‐processing steps - datasciencecampus/WorldPop-Malawi-fork GitHub Wiki
AI was used to generate this content.
1. Raster_Mosaicking_Buildings
Builds a buffered Malawi boundary (or loads it if already present), aligns it to a reference building raster, then mosaics 2018 or 2023 building rasters from Malawi + neighboring countries into a single Malawi-focused raster per variable(1), cropping and masking to the buffered boundary before saving. The idea is to take the Malawi building file to create a "Malawi" boundary + taking files for all the neighboring countries so that the Malawi boundary can be extended by 10 km into each neighboring country. This is to stop any edge issues.
(1)The “variables” are derived from the raster name suffix after the first three country letters, e.g.
- buildings_count_2018_glv2_5_t0_5_C_100m_v1.tif
- buildings_count_BCB_gl_100m_v1_1.tif
- buildings_count_BCB_ms_100m_v1_1.tif
Zambia only has the first variable, while the other three appear in Malawi/Mozambique/Tanzania. The script will still attempt to mosaic all three variables based on the union of file names.
From WP website covatiate information: https://hub.worldpop.org/project/categories?id=14:
-
BCB_gl: Gridded maps of building patterns - Google (BCB) Building metrics are calculated based on the pixel in which their respective centroids are located (Building Centroid Based – BCB method), or the pixel(s) that their geometries intersect (Pixel Intersected Based – PIB method). For buildings that straddle multiple grid cells, the grid cell in which the buildings' centroids are located will be allocated the area value. Note: For the BCB method, total building area may exceed the area of a grid cells if the centroid of a large building falls within the grid cell.
-
BCB_ms: Gridded maps of building patterns - Microsoft (BCB) The value of each grid cell represents the count of buildings within respective grid cells divided by the grid cell's area, based on the two methods (PIB/BCB). Building metrics are calculated based on the pixel in which their respective centroids are located (Building Centroid Based – BCB method), or the pixel(s) that their geometries intersect (Pixel Intersected Based – PIB method).
-
glv2: From Ortis:
mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tifis a new file I created by mosaickingbuildings_count_2018_glv2_5_t0_5_C_100m_v1.tiffiles from Malawi, Tanzania, Mozambique and Zambia. You will not find it on the WorldPop server. Running this script01_Raster_Mosaicking_Buildings_2018create this file.
Potential issue: This script takes all the raster files for all the countries and uses the first file's CRS and reproject all rasters to the CRS of the first raster. If order changes and that file has a different CRS?
01_Raster_Mosaicking_Buildings_2018.R
Inputs
- Boundary shapefile:
data/Shapefiles/Country_Shapefile_Buffer_10km.shp(loaded if present; otherwise generated viagenerate_buffered_country_boundary()in utils.R) made from2018_MPHC_EAs_Final_for_Use.shp(provided by MNSO) - Reference raster for CRS:
data/Malawi_Covs/2018_Buildings/mwi_buildings_count_2018_glv2_5_t0_5_C_100m_v1.tif. - All
.tifrasters in:data/Malawi_Covs/2018_Buildingsdata/Tanzania_Covs/2018_Buildingsdata/Mozambique_Covs/2018_Buildingsdata/Zambia_Covs/2018_Buildings
Outputs
- Mosaicked, cropped, masked rasters written to
data/Mosaic_Buildings_2018, named with prefixMOS_MLW+ the shared suffix of the input raster names (after stripping the first 3 characters).
01_Raster_Mosaicking_Buildings_2024.R
Inputs
- Boundary shapefile same as in the above section
- Reference raster for CRS: data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif.
- All
.tifrasters in:- data/Malawi_Covs/2024_Buildings
- data/Mozambique_Covs/2024_Buildings
- data/Tanzania_Covs/2024_Buildings
- data/Zambia_Covs/2024_Buildings
Outputs
- Mosaicked, cropped, masked rasters written to data/Mosaic_Buildings_2024, named
MOS_MLW+ the shared suffix of the input raster names (after stripping the first 3 characters).
Raster_Mosaicking_Workflow
These scripts are doing a similar thing to the Raster_Mosaicking_Buildings but for the other covariate datasets. Essentially, it finds each unique raster name (after trimming the first 3 characters of filenames), it mosaics the matching rasters across countries (Malawi, Mozambique, Tanzania, Zambia), crops and masks to the Malawi boundary, and writes the result to the 2018 or 2024 mosaic output folder.
Assumptions
- Filenames across country folders share the same suffix after removing the first 3 characters; this is used to match rasters for mosaicking.
- Output filenames start with 7 characters that can be trimmed for comparison; the script expects this to align with files_trim logic.
- All relevant covariate rasters are
.tifand live directly in the specified folders (no deeper subfolders). - The first raster found for a given name has the desired CRS; all others can be safely reprojected to it without loss.
- The reference building raster exists and is readable; its CRS is valid for reprojecting the boundary.
- Mosaic order with
fun="first"is acceptable; overlapping pixels keep the first raster’s value. - The output directory can be created and overwritten without issues.
01_Raster_Mosaicking_Workflow_2018.R
Inputs
Boundary shapefile (read or generated): data/Shapefiles/Country_Shapefile_Buffer_10km.shp
Reference raster for CRS (if ordered the same): data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif
Source covariate rasters (all .tif):
- data/Malawi_Covs/2018_Covariates
- data/Mozambique_Covs/2018_Covariates
- data/Tanzania_Covs/2018_Covariates
- data/Zambia_Covs/2018_Covs
Outputs
Mosaicked, cropped, masked rasters written to Mosaic_Covariates_2018 with names like MOS_MLW<raster_name>.
01_Raster_Mosaicking_Workflow_2024.R
Inputs
Boundary shapefile (read or generated): data/Shapefiles/Country_Shapefile_Buffer_10km.shp
Reference raster for CRS: data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif
Source covariate rasters (all .tif):
- data/Malawi_Covs/2024_Covariates
- data/Mozambique_Covs/2024_Covariates
- data/Tanzania_Covs/2024_Covariates
- data/Zambia_Covs/2024_Covs
Outputs
Mosaicked, cropped, masked rasters written to Mosaic_Covariates_2024 with names like MOS_MLW<raster_name>.
Notes When running this script this error occurred at the stage of "# Loop through each unique raster name and process":
_tiffWriteProc: No space left on device.
==========Error: [project] cannot write values (err: 3)
This was cause by the temporary directory being full. To fix it run in R:
tempdir()
This will tell you where the temp directory is in your file system - delete any unnecessary (all) files and re-run the script which will take into account any files that were processed already.
00_Data_processing
There are two scripts for data processing: 00_Data_Processing2.R is intended to run after 00_Data_Processing.R and uses the same core inputs, but adds spatial assignment logic, DHS enrichment, and optional district datasets. Both scripts write to the same key outputs: data/Output_Data/summarized_survey_data.csv and an EA-level household size spatial file (shp vs gpkg). Running the second script will effectively update/replace the summary CSV with the more spatially informed version.
00_Data_Processing.R
Overall: Ingests multiple survey/census datasets, summarizes population and household metrics at EA level, derives building counts from structure points, combines sources with a priority rule for observed household counts, and exports summary tables and EA-level shapefile outputs.
Inputs
- Census/survey
.dtafiles inMNSO-Data:
- mphc2018Data_AllRegions.dta
- mphc2018Data_structures.dta
- ICT Listing WorldPop.dta
- IHS6 Listing WorldPop.dta
- Naca Listing WorldPop.dta
- MDHS_2024_NoDZLK_anonymized.dta
- FINAL MDHS LISTING DATA_Annon.dta
- EA polygons:
- 2018_MPHC_EAs_Final_for_Use.shp
Outputs
EA-level summary CSV: data/Output_Data/summarized_survey_data.csv
EA-level household size shapefile: data/Output_Data/hh_size_data.shp and sidecar files
Structures points GPKG (if not already present): data/Output_Data/mphc_structures_points.gpkg
Assumptions
- Each census row represents one person (uses hh_count = 1 for population summing).
- EA_CODE is correctly formed by concatenating district + padded TA + padded EA.
- Nearest EA polygon assignment for structure points is valid enough for building counts.
- The priority rule for
observed_hh_count(IHS6 > Naca > ICT) is acceptable for analysis. - Inputs contain valid GPS columns with expected field names.
00_Data_Processing2.R
Overall: Extends/replicates EA-level summarization using spatial location for GPS points (nearest EA polygon), handles GPS accuracy thresholds, incorporates DHS listing/survey data (when available), and optionally Zomba/Malemia datasets. Exports a more enriched EA summary and EA-level household size GPKG.
Inputs
- Same core
.dtafiles as above inMNSO-Data - EA polygons:
2018_MPHC_EAs_Final_for_Use.shp - Precomputed GPS EA geopackage (created here if missing):
mphc_2018_sf_ea.gpkg - Optional files (if present):
- data/MNSO-Data/DHS_Segmented_File.csv
- data/Output_Data/zomba_rbind_data.csv
- data/MNSO-Data/malemia_hh_without_IDs.csv
Outputs
EA-level summary CSV (overwrites/updates): data/Output_Data/summarized_survey_data.csv
EA-level household size GPKG (appends): data/Output_Data/hh_size_data.gpkg
GPS-assigned census geopackage (if missing): mphc_2018_sf_ea.gpkg
Assumptions
- GPS accuracy thresholds: points with accuracy > 5m are summarized by original EA code; < 5m by nearest EA polygon.
- DHS centroids more than 5km from EA polygons are excluded.
- Optional Zomba/Malemia files might be missing; the script inserts NA placeholders if absent.
- EA polygons provide a reliable nearest feature for all GPS points.
02_Covariates_Extraction
Overall, it extracts EA-level covariate summaries from raster mosaics for 2024 and 2018 (building counts, non-residential counts, and continuous covariates), adds EA centroid coordinates, joins EA-level survey summaries, and exports two modeling tables plus variable-name maps.
Inputs
- EA polygons:
2018_MPHC_EAs_Final_for_Use.shp - EA survey summaries:
data/Output_Data/summarized_survey_data.csv - Building mosaics:
- data/Input_Data/Mosaic_Buildings_2024
- data/Input_Data/Mosaic_Buildings_2018
- Covariate mosaics:
- data/Input_Data/Mosaic_Covariates_2024
- data/Input_Data/Mosaic_Covariates_2018
- Non-residential raster:
data/Output_Data/non_res_raster.tif
Outputs
- EA-level covariate tables:
- data/Output_Data/Malawi_2024_data.csv
- data/Output_Data/Malawi_2018_data.csv
- Variable name maps:
- data/Output_Data/var_names_2024.csv
- data/Output_Data/var_names_2018.csv
Assumptions
- EA polygons are valid and can be reprojected to the building raster CRS.
- Covariate rasters align spatially and can be summarized by exact_extract with mean (continuous surfaces).
- Building and non-residential rasters are appropriate for sum aggregation within EAs.
- Covariate stacks contain 64 layers (renamed to x1–x64), and that count is stable across years.
- EA_CODE in the survey summary matches the shapefile’s EA_CODE after type conversion.
Covariate information
esalc - Land Cover covariates
esalc — ESA CCI Land Cover. The ESA Climate Change Initiative (CCI) generated global land cover maps at approximately 300m resolution, with a legend of 22 land cover classes defined using the UN FAO Land Cover Classification System (LCCS). CEOS WorldPop uses these as covariate inputs.
The full LCCS class codes used by WorldPop include:
Code Land Cover Class 10 Cropland, rainfed 20 Cropland, irrigated 30 Mosaic cropland / natural vegetation 50 Tree cover, broadleaved, evergreen 60 Tree cover, broadleaved, deciduous 70 Tree cover, needleleaved, evergreen 80 Tree cover, needleleaved, deciduous 90 Tree cover, mixed leaf type 100 Mosaic tree and shrub / herbaceous 110 Mosaic herbaceous / tree and shrub 120 Shrubland 130 Grassland 140 Lichens and mosses 150 Sparse vegetation 160 Tree cover, flooded, fresh or brackish water 170 Tree cover, flooded, saline water 180 Shrub/herbaceous, flooded 190 Urban / built-up 200 Bare areas 210 Water bodies 220 Permanent snow and ice
dst — Distance to. WorldPop computes "distance to edges of reclassified ESA-CCI-LC classes" as annual covariates. WorldPop .
2022 — The year the underlying ESA CCI land cover map refers to.
100m — The spatial resolution: approximately 100 metres (3 arc-seconds).
v1 — Version 1 of the WorldPop covariate dataset. WorldPop systematically collated 73 annual, spatio-temporally harmonised datasets at ~100m global resolution for use in population modelling. WorldPop