Data pre‐processing steps - datasciencecampus/WorldPop-Malawi-fork GitHub Wiki

AI was used to generate this content.

1. Raster_Mosaicking_Buildings


Builds a buffered Malawi boundary (or loads it if already present), aligns it to a reference building raster, then mosaics 2018 or 2023 building rasters from Malawi + neighboring countries into a single Malawi-focused raster per variable(1), cropping and masking to the buffered boundary before saving. The idea is to take the Malawi building file to create a "Malawi" boundary + taking files for all the neighboring countries so that the Malawi boundary can be extended by 10 km into each neighboring country. This is to stop any edge issues.

(1)The “variables” are derived from the raster name suffix after the first three country letters, e.g.

  • buildings_count_2018_glv2_5_t0_5_C_100m_v1.tif
  • buildings_count_BCB_gl_100m_v1_1.tif
  • buildings_count_BCB_ms_100m_v1_1.tif

Zambia only has the first variable, while the other three appear in Malawi/Mozambique/Tanzania. The script will still attempt to mosaic all three variables based on the union of file names.

From WP website covatiate information: https://hub.worldpop.org/project/categories?id=14:

  • BCB_gl: Gridded maps of building patterns - Google (BCB) Building metrics are calculated based on the pixel in which their respective centroids are located (Building Centroid Based – BCB method), or the pixel(s) that their geometries intersect (Pixel Intersected Based – PIB method). For buildings that straddle multiple grid cells, the grid cell in which the buildings' centroids are located will be allocated the area value. Note: For the BCB method, total building area may exceed the area of a grid cells if the centroid of a large building falls within the grid cell.

  • BCB_ms: Gridded maps of building patterns - Microsoft (BCB) The value of each grid cell represents the count of buildings within respective grid cells divided by the grid cell's area, based on the two methods (PIB/BCB). Building metrics are calculated based on the pixel in which their respective centroids are located (Building Centroid Based – BCB method), or the pixel(s) that their geometries intersect (Pixel Intersected Based – PIB method).

  • glv2: From Ortis: mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif is a new file I created by mosaicking buildings_count_2018_glv2_5_t0_5_C_100m_v1.tif files from Malawi, Tanzania, Mozambique and Zambia. You will not find it on the WorldPop server. Running this script 01_Raster_Mosaicking_Buildings_2018 create this file.

Potential issue: This script takes all the raster files for all the countries and uses the first file's CRS and reproject all rasters to the CRS of the first raster. If order changes and that file has a different CRS?

01_Raster_Mosaicking_Buildings_2018.R

Inputs

  • Boundary shapefile: data/Shapefiles/Country_Shapefile_Buffer_10km.shp (loaded if present; otherwise generated via generate_buffered_country_boundary() in utils.R) made from 2018_MPHC_EAs_Final_for_Use.shp (provided by MNSO)
  • Reference raster for CRS: data/Malawi_Covs/2018_Buildings/mwi_buildings_count_2018_glv2_5_t0_5_C_100m_v1.tif.
  • All .tif rasters in:
    • data/Malawi_Covs/2018_Buildings
    • data/Tanzania_Covs/2018_Buildings
    • data/Mozambique_Covs/2018_Buildings
    • data/Zambia_Covs/2018_Buildings

Outputs

  • Mosaicked, cropped, masked rasters written to data/Mosaic_Buildings_2018, named with prefix MOS_MLW + the shared suffix of the input raster names (after stripping the first 3 characters).

01_Raster_Mosaicking_Buildings_2024.R

Inputs

  • Boundary shapefile same as in the above section
  • Reference raster for CRS: data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif.
  • All .tif rasters in:
    • data/Malawi_Covs/2024_Buildings
    • data/Mozambique_Covs/2024_Buildings
    • data/Tanzania_Covs/2024_Buildings
    • data/Zambia_Covs/2024_Buildings

Outputs

  • Mosaicked, cropped, masked rasters written to data/Mosaic_Buildings_2024, named MOS_MLW + the shared suffix of the input raster names (after stripping the first 3 characters).

Raster_Mosaicking_Workflow


These scripts are doing a similar thing to the Raster_Mosaicking_Buildings but for the other covariate datasets. Essentially, it finds each unique raster name (after trimming the first 3 characters of filenames), it mosaics the matching rasters across countries (Malawi, Mozambique, Tanzania, Zambia), crops and masks to the Malawi boundary, and writes the result to the 2018 or 2024 mosaic output folder.

Assumptions

  1. Filenames across country folders share the same suffix after removing the first 3 characters; this is used to match rasters for mosaicking.
  2. Output filenames start with 7 characters that can be trimmed for comparison; the script expects this to align with files_trim logic.
  3. All relevant covariate rasters are .tif and live directly in the specified folders (no deeper subfolders).
  4. The first raster found for a given name has the desired CRS; all others can be safely reprojected to it without loss.
  5. The reference building raster exists and is readable; its CRS is valid for reprojecting the boundary.
  6. Mosaic order with fun="first" is acceptable; overlapping pixels keep the first raster’s value.
  7. The output directory can be created and overwritten without issues.

01_Raster_Mosaicking_Workflow_2018.R

Inputs Boundary shapefile (read or generated): data/Shapefiles/Country_Shapefile_Buffer_10km.shp Reference raster for CRS (if ordered the same): data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif Source covariate rasters (all .tif):

  • data/Malawi_Covs/2018_Covariates
  • data/Mozambique_Covs/2018_Covariates
  • data/Tanzania_Covs/2018_Covariates
  • data/Zambia_Covs/2018_Covs

Outputs Mosaicked, cropped, masked rasters written to Mosaic_Covariates_2018 with names like MOS_MLW<raster_name>.

01_Raster_Mosaicking_Workflow_2024.R

Inputs Boundary shapefile (read or generated): data/Shapefiles/Country_Shapefile_Buffer_10km.shp Reference raster for CRS: data/Malawi_Covs/2024_Buildings/mwi_buildings_count_2023_glv2_5_t0_5_C_100m_v1.tif Source covariate rasters (all .tif):

  • data/Malawi_Covs/2024_Covariates
  • data/Mozambique_Covs/2024_Covariates
  • data/Tanzania_Covs/2024_Covariates
  • data/Zambia_Covs/2024_Covs

Outputs Mosaicked, cropped, masked rasters written to Mosaic_Covariates_2024 with names like MOS_MLW<raster_name>.

Notes When running this script this error occurred at the stage of "# Loop through each unique raster name and process":

_tiffWriteProc: No space left on device.
==========Error: [project] cannot write values (err: 3)

This was cause by the temporary directory being full. To fix it run in R:

tempdir()

This will tell you where the temp directory is in your file system - delete any unnecessary (all) files and re-run the script which will take into account any files that were processed already.

00_Data_processing


There are two scripts for data processing: 00_Data_Processing2.R is intended to run after 00_Data_Processing.R and uses the same core inputs, but adds spatial assignment logic, DHS enrichment, and optional district datasets. Both scripts write to the same key outputs: data/Output_Data/summarized_survey_data.csv and an EA-level household size spatial file (shp vs gpkg). Running the second script will effectively update/replace the summary CSV with the more spatially informed version.

00_Data_Processing.R

Overall: Ingests multiple survey/census datasets, summarizes population and household metrics at EA level, derives building counts from structure points, combines sources with a priority rule for observed household counts, and exports summary tables and EA-level shapefile outputs.

Inputs

  1. Census/survey .dta files in MNSO-Data:
  • mphc2018Data_AllRegions.dta
  • mphc2018Data_structures.dta
  • ICT Listing WorldPop.dta
  • IHS6 Listing WorldPop.dta
  • Naca Listing WorldPop.dta
  • MDHS_2024_NoDZLK_anonymized.dta
  • FINAL MDHS LISTING DATA_Annon.dta
  1. EA polygons:
  • 2018_MPHC_EAs_Final_for_Use.shp

Outputs EA-level summary CSV: data/Output_Data/summarized_survey_data.csv EA-level household size shapefile: data/Output_Data/hh_size_data.shp and sidecar files Structures points GPKG (if not already present): data/Output_Data/mphc_structures_points.gpkg

Assumptions

  • Each census row represents one person (uses hh_count = 1 for population summing).
  • EA_CODE is correctly formed by concatenating district + padded TA + padded EA.
  • Nearest EA polygon assignment for structure points is valid enough for building counts.
  • The priority rule for observed_hh_count (IHS6 > Naca > ICT) is acceptable for analysis.
  • Inputs contain valid GPS columns with expected field names.

00_Data_Processing2.R

Overall: Extends/replicates EA-level summarization using spatial location for GPS points (nearest EA polygon), handles GPS accuracy thresholds, incorporates DHS listing/survey data (when available), and optionally Zomba/Malemia datasets. Exports a more enriched EA summary and EA-level household size GPKG.

Inputs

  1. Same core .dta files as above in MNSO-Data
  2. EA polygons: 2018_MPHC_EAs_Final_for_Use.shp
  3. Precomputed GPS EA geopackage (created here if missing): mphc_2018_sf_ea.gpkg
  4. Optional files (if present):
  • data/MNSO-Data/DHS_Segmented_File.csv
  • data/Output_Data/zomba_rbind_data.csv
  • data/MNSO-Data/malemia_hh_without_IDs.csv

Outputs

EA-level summary CSV (overwrites/updates): data/Output_Data/summarized_survey_data.csv EA-level household size GPKG (appends): data/Output_Data/hh_size_data.gpkg GPS-assigned census geopackage (if missing): mphc_2018_sf_ea.gpkg

Assumptions

  • GPS accuracy thresholds: points with accuracy > 5m are summarized by original EA code; < 5m by nearest EA polygon.
  • DHS centroids more than 5km from EA polygons are excluded.
  • Optional Zomba/Malemia files might be missing; the script inserts NA placeholders if absent.
  • EA polygons provide a reliable nearest feature for all GPS points.

02_Covariates_Extraction


Overall, it extracts EA-level covariate summaries from raster mosaics for 2024 and 2018 (building counts, non-residential counts, and continuous covariates), adds EA centroid coordinates, joins EA-level survey summaries, and exports two modeling tables plus variable-name maps.

Inputs

  1. EA polygons: 2018_MPHC_EAs_Final_for_Use.shp
  2. EA survey summaries: data/Output_Data/summarized_survey_data.csv
  3. Building mosaics:
  • data/Input_Data/Mosaic_Buildings_2024
  • data/Input_Data/Mosaic_Buildings_2018
  1. Covariate mosaics:
  • data/Input_Data/Mosaic_Covariates_2024
  • data/Input_Data/Mosaic_Covariates_2018
  1. Non-residential raster: data/Output_Data/non_res_raster.tif

Outputs

  1. EA-level covariate tables:
  • data/Output_Data/Malawi_2024_data.csv
  • data/Output_Data/Malawi_2018_data.csv
  1. Variable name maps:
  • data/Output_Data/var_names_2024.csv
  • data/Output_Data/var_names_2018.csv

Assumptions

  • EA polygons are valid and can be reprojected to the building raster CRS.
  • Covariate rasters align spatially and can be summarized by exact_extract with mean (continuous surfaces).
  • Building and non-residential rasters are appropriate for sum aggregation within EAs.
  • Covariate stacks contain 64 layers (renamed to x1–x64), and that count is stable across years.
  • EA_CODE in the survey summary matches the shapefile’s EA_CODE after type conversion.

Covariate information

esalc - Land Cover covariates

esalc — ESA CCI Land Cover. The ESA Climate Change Initiative (CCI) generated global land cover maps at approximately 300m resolution, with a legend of 22 land cover classes defined using the UN FAO Land Cover Classification System (LCCS). CEOS WorldPop uses these as covariate inputs.

The full LCCS class codes used by WorldPop include:

Code Land Cover Class 10 Cropland, rainfed 20 Cropland, irrigated 30 Mosaic cropland / natural vegetation 50 Tree cover, broadleaved, evergreen 60 Tree cover, broadleaved, deciduous 70 Tree cover, needleleaved, evergreen 80 Tree cover, needleleaved, deciduous 90 Tree cover, mixed leaf type 100 Mosaic tree and shrub / herbaceous 110 Mosaic herbaceous / tree and shrub 120 Shrubland 130 Grassland 140 Lichens and mosses 150 Sparse vegetation 160 Tree cover, flooded, fresh or brackish water 170 Tree cover, flooded, saline water 180 Shrub/herbaceous, flooded 190 Urban / built-up 200 Bare areas 210 Water bodies 220 Permanent snow and ice

dst — Distance to. WorldPop computes "distance to edges of reclassified ESA-CCI-LC classes" as annual covariates. WorldPop .

2022 — The year the underlying ESA CCI land cover map refers to.

100m — The spatial resolution: approximately 100 metres (3 arc-seconds).

v1 — Version 1 of the WorldPop covariate dataset. WorldPop systematically collated 73 annual, spatio-temporally harmonised datasets at ~100m global resolution for use in population modelling. WorldPop