Ratio change metadata - datasciencecampus/WorldPop-Malawi-fork GitHub Wiki

Location of code and data

Code

  • malawi_ratio_change.R on ons-compatibility-ratiochange branch

Data

Summary of ratio change method

The ratio change method is a simple method to update census population estimates for Enumeration Areas (EA) in Malawi (persons, households, etc) by using the ratio change of a separate data source as a proxy for change in the population

For our initial implementation of the ratio change method, we take the ratio of buildings estimated in each Malawi EA from Google footprint data from 2018 and 2023. We then apply this ratio to the 2018 Census estimate of households for each EA, to give an updated estimated number of households for 2023.

For example, a Malawi EA may have:

  • Google footprint data - count of buildings in 2018 = 100, count of buildings in 2023 = 150
  • Ratio of Google footprint data from 2023-2018 = 1.5
  • The estimate number of households from the 2018 Census = 300
  • Ratio change of Census households = 300 * 1.5 = 450

Quality of estimates is measured by comparing the updated household estimate to the number of households estimated from the most recent survey data, the absolute error is simply:

  • number of households from the survey - number of households estimated from ratio change
  • Three surveys are available to us with household estimates for EAs; the ICT (2023), the IHS (December 2024) and the NACA (August 2024). Most EAs only have an estimate from one source , however where there are estimates from multiple surveys, the most recent one is taken.
  • Are we certain on the quality of the survey estimates? Is one survey better than another?

We can also get the absolute error of the WorldPop modelled estimate number of households. This means we can compare the WorldPop model to the ratio change method by summarising the absolute error between the two methods. We calculate:

  • Mean Absolute Error (MAE) - average of absolute error values across all Malawi EAs
  • Variance of Absolute Error - standard deviation of absolute error values across all Malawi EAs
  • Max Absolute Error - maximum absolute error value from all Malawi EAs

We can also summarise performance from the ratio change vs WorldPop method by summarising the estimated number of households for each Malawi EA. For example:

  • Maximum EA household estimate
  • Minimum EA household estimate
  • Total number of EAs over a given threshold (300 households for Urban EAs, 200 households for Rural EAs), which we also express as a percentage
  • Total number of EAs under a given threshold (100 households for both Urban and Rural EAs)

Performance of Ratio Change method vs initial WorldPop method

Summary of absolute error - urban

<style> </style>
Method max_abs_error sd_abs_error mean_abs_error
WorldPop model 1705 102.6641497 77.43739635
Ratio change 776.6196441 64.94645663 65.45033271

Summary of absolute error - rural

<style> </style>
Method max_abs_error sd_abs_error mean_abs_error
WorldPop model 4151 265.7733351 215.039267
Ratio change 540.7403809 65.69975957 66.62777277

Summary of estimated households - urban

<style> </style>
Metric Max Min Sum_over_threshold Sum_less_100hh Total Too_big_perc Not_too_big_count
WorldPop model 4763 34 2187 7 2770 78.95306859 583
Ratio change 1592.963902 0.987518261 684 74 2767 24.71991326 2083

Summary of estimated households - rural

<style> </style>
Metric Max Min Sum_over_threshold Sum_less_100hh Total Too_big_perc Not_too_big_count
WorldPop model 4114 0 8369 1749 15961 52.43405802 7592
Ratio change 1552.81559 0 7792 1694 15944 48.87104867 8152

Compare performance of ratio change method - using google temporal satellite data

Alternate satellite data sources are available to use instead of the original Google Open Buildings dataset (https://sites.research.google/gr/open-buildings) where we are concerned that the temporal aspect of this data (i.e., comparing metrics from 2018 to 2023) may be blurred due to multiple years worth of data being combined for the '2023' estimates.

Instead, Google Temporal Buildings dataset (https://sites.research.google/gr/open-buildings/temporal) provides estimates of buildings from satellite imagery from 2016 through to 2023. This dataset has 3 measures; 1) fractional building count, 2) building height, and 3) building presence, all of which are provided for 4m pixels across Malawi.

Fractional building count provides an estimate as to how much of a fraction of a building is within a given pixel. Building height is the height in metres from the ground, and is capped at 100m. Building presence is effectively a confidence score, ranging from 0-1, to determine how confident the model is that a given pixel contains a part of a building. Screenshots are provided at the end to show the different data layered onto part of Malawi satellite image.

From these measures, we calculate two important metrics; 1) building count estimate, and 2) total building area (building volume can also be created, this is to be added).

Building count estimate is taken by summing together all of the fractional building counts across all pixels within EA boundaries which should, in theory, give the total number of buildings within an EA (note that the method applies a scaling factor correction, so the raw fractional counts are scaled, I need to understand this methodology more and will add detail here when ready).

Total building area starts with the building presence measure, takes only pixels that have a value above 0.5 (i.e., the model is more confident than not that the given pixel has a building part in it), and then multiplies the total number of remaining pixels within an EA by 16 (16 metre squared, because each pixel is 4m x 4m). This gives total building area of an EA.

Because the data is available from 2016 to 2023, we are able to take the years 2018 and 2023 to feed into the same ratio change methodology, substituting in these two sets of estimates (which I call 'Building Count Temporal' and 'Total Building Area Temporal'). By calculating the absolute error and number of buildings estimated as above, we can compare the ratio change method using these new variables to previous results:

<style> </style>
Urban      
Method max_abs_error sd_abs_error mean_abs_error
WorldPop model 4151 265.77334 215.03927
Building Count Original 540.7404 65.69976 66.62777
Building Count Temporal 588.317 64.84683 71.81328
Total Building Area Temporal 642.2436 75.5704 72.24901
       
Rural      
Method max_abs_error sd_abs_error mean_abs_error
WorldPop model 1705 102.66415 77.4374
Building Count Original 776.6196 64.94646 65.45033
Building Count Temporal 800.3693 60.86491 60.67796
Total Building Area Temporal 952.0508 69.94239 64.87916
<style> </style>
Urban                
Metric Max Median Min Sum_over_threshold Sum_less_100hh Total Too_big_perc Not_too_big_count
WorldPop model 4763 402 34 2187 7 2770 78.95306859 583
Building Count Original 1592.963902 241.76 0.987518261 684 74 2767 24.71991326 2083
Building Count Temporal 913.69 255.95 0.99 845 67 2767 30.54 1922
Total Building Area Temporal 1393.84 247.43 0.97 736 70 2767 26.6 2031
                 
                 
Rural                
Metric Max Median Min Sum_over_threshold Sum_less_100hh Total Too_big_perc Not_too_big_count
WorldPop model 4114 207 0 8369 1749 15961 52.43 7592
Building Count Original 1552.82 197.33 0 7792 1694 15944 48.87 8152
Building Count Temporal 1467.77 194.03 0 7589 1727 15969 47.52 8380
Total Building Area Temporal 1807.81 195.21 0 7664 1696 15960 48.02 8296

Consequently, we can say that using either of the Google satellite data for building counts produces similar outputs, leaving to ratio change estimates that perform better than when using Google satellite total area, which in turn is better than the current WorldPop model. However, there are a few things to consider/explore with the Google temporal data:

  • The confidence metric for pixels - ChatGPT assumes a pixel over 0.5 is a valid enough threshold. The academic paper and demonstration code accompanying Earth Engine instead suggested a threshold of 0.34, but it is said "This however might need tweaking for a given location and time since the model confidence varies based on a number of factors such as cloud cover, imagery alignment, etc.". Therefore, I think there is more work to do to explore this temporal dataset further, fine-tuning the confidence metric to ensure that we are selecting valid data for our context

  • How consistent is the confidence metric across areas in Malawi - e.g. urban vs rural areas. Are there likely to be more pixels with higher confidence estimates in urban areas because the likelihood of there being more buildings in urban areas, compared to rural areas?

  • Validate outputs - I need to check that the estimates roughly correspond to the number of buildings visible on satellite images. Though we can gain confidence from the fact that the different Google satellite sources are similar (though would we expect the temporal data to differ, given our concerns about the original Google data?). Validation will be tricky though, because Earth Engine only provides one detailed satellite image, so I can't access detailed images from 2018 and 2023 for example to verify both years.

Further comparisons using alternate satellite data sources

To start with, some datasets have been mentioned but brief reasons outlined here as to why these haven't been considered:

  • Global Human Settlement Layer (GHSL) data is a composition of worldwide estimates from the EU Copernicus satellite data. Various datasets are provided including population estimates, urban/rural classifications, building heights, surface area and volume, built-up surface area, further split down to identify non-residential surface area (so theoretically you could derive a residential surface area, though I assume there are limitations to this, otherwise GHSL would provide this?). The main challenge with using GHSL data is the temporal inconsistency. Data is available from 1975 to 2030 (projections), every 5 years, at 100m grid cell resolution. However, there are three years provided in addition, including 2018, provided at 10m resolution. Consequently, the relevance to this work is that we have no data available for 2023, and the data available for 2018 compared to other years is at a different spatial scale (which may also raise differences in quality of estimates at different scales for different time points).

  • OpenStreetMap (OSM) is a volunteered geographic information (VGI) product where a community of users and volunteers provide information about buildings, using uploaded satellite imagery and their local knowledge of the area. As this is dependent on volunteers, the coverage of information is very patchy, with some academic work suggesting that buildings counts using only OSM are substantially lower than other products. There may also be concerns when using as a measure of change, what is the coverage like from 2018 compared to 2023/2026? I would hazard a guess that coverage would be better in later years for several reasons - better access to technology, greater demand/need for geospatial information.

Next I present comparisons of outputs from the ratio change method, with several options considered in using measure of change of buildings. These include the options already presented:

  • Google Open Buildings - count
  • Google Open Temporal Buildings - count
  • Google Open Temporal Buildings - total area

Plus the following latest additions:

  • Google Temporal Buildings smoothed - count
  • Google Temporal Buildings smoothed - total area
  • Google Temporal Buildings smoothed - volume

The Google Temporal Buildings smoothed is a product from academics (Zhang et al. A temporally harmonised gridded building-stock dataset for the Global South, 2016-2023) who take the raw Google Open Temporal Buildings data, and try to address inconsistencies in the time series of this data. For example, satellite imagery over time may have changes in observation conditions (cloud cover, haze, atmospheric conditions, seasonal vegetation, viewing geometry). There may also be changes in satellite data processing/modelling. This led to some cases of temporal instability and "zig-zag" amplitudes, where changes of building metrics for a given area would change drastically from one year to the next. Instead, it may be expected that any change would be gradual. Consequently, a methodology was developed to smooth the time series of building metrics (count, area and volume) for the global south.

The raw Google pixel data were aggregated to 100m pixels for this analysis, which the smoothed 100m grid estimates made publicly available. I took this data (for counts, area and volume), aggregated that data to Malawi EAs for 2018 and 2023, and passed this data through the ratio change pipeline.

Below I present the same metrics from before (absolute error, and summary of estimated number of households within EAs, separately for urban and rural EAs).

Summary of absolute error - urban

<style> </style>
Method max_abs_error sd_abs_error mean_abs_error
WorldPop 4151 265.7733351 215.039267
Ratio Change - original - count 540.7403809 65.69975957 66.62777277
Ratio Change - temporal - count 588.3169847 64.84683133 71.81328449
Ratio Change - temporal - area 642.2435746 75.57039655 72.24901358
Ratio Change - smoothed - count 466.5951258 67.16950375 71.80745337
Ratio Change - smoothed - area 743.4253243 78.09700279 76.15261418
Ratio Change - smoothed - volume 928.4899419 86.16713446 73.99321049

So using smooth temporal count data leads to better performance in terms of reduced max error, but mean error is best using original Google Open Buildings data. Area and volume tend to perform worse.

Summary of absolute error - rural

<style> </style>
Method max_abs_error sd_abs_error mean_abs_error
WorldPop 1705 102.6641497 77.43739635
Ratio Change - original - count 776.6196441 64.94645663 65.45033271
Ratio Change - temporal - count 800.3693031 60.8649133 60.67796163
Ratio Change - temporal - area 952.0508114 69.94239442 64.87916437
Ratio Change - smoothed - count 736.5958989 65.74520155 63.77889113
Ratio Change - smoothed - area 1139.847608 73.31376983 64.63212926
Ratio Change - smoothed - volume 1310.543831 78.46916417 66.58149336

Best in terms of max error is smoothed count data or original Google data. Mean error very similar. Area and volume much worse in terms of max error.

Summary of estimated household size - urban

<style> </style>
  Min HH count Median HH count Max HH count % EAs over threshold % EAs under 100 Total EAs
WorldPop 34 402 4763 78.95 0.25 2770
Ratio Change - original - count 0.9875183 241.7554 1592.9639 24.72 2.67 2767
Ratio Change - temporal - count 0.9699394 247.4297 1393.8428 26.60 2.53 2767
Ratio Change - temporal - area 0.9956518 255.953 913.6911 30.54 2.42 2767
Ratio Change - smoothed - count 1.0146817 256.1538 1577.1393 30.32 2.31 2767
Ratio Change - smoothed - area 0.9924106 257.9221 1556.3914 31.77 2.28 2767
Ratio Change - smoothed - volume 0.9700229 236.791 2872.534 24.50 2.57 2767

Original Google data is best performing overall, but smoothed volume data is also comparable here.

Summary of estimated household size - rural

<style> </style>
  Min HH count Median HH count Max HH count % EAs over threshold % EAs under 100 Total EAs
WorldPop 0 207 4114 52.43 10.96 15961
Ratio Change - original - count 0 197.3256 1552.816 48.87 10.62 15944
Ratio Change - temporal - count 0 195.2137 1807.811 48.02 10.63 15960
Ratio Change - temporal - area 0.01211886 194.0269 1467.77 47.52 10.81 15969
Ratio Change - smoothed - count 0 211.2961 1658.41 54.52 8.05 15956
Ratio Change - smoothed - area 0 210.6304 1896.408 54.09 7.60 15967
Ratio Change - smoothed - volume 0 200.7425 2294.24 50.36 8.42 15967

Some interesting patterns. WorldPop model similar to ratio change. Over threshold, the original and temporal data perform slightly better. The smoothed data performs better looking at number EAs below threshold.

Next steps

  • Census data - treating unoccupied households, households of multiple occupation (both of which we can work out from the structures/all regions datasets)
  • Measuring uncertainty - adopt WP growth model? Adapt ratio change to incorporate some kind of simulation to produce confidence intervals?
  • Further exploration of satellite imagery - including conversations with WP/Soton experts

Google Earth Engine screenshots

Base satellite image

Screenshot 2026-07-21 at 09 53 59

Fractional estimate

Screenshot 2026-07-21 at 09 52 17

At a glance, this doesn't look like it covers the entirety of a building?

Fractional estimate with building confidence (green = less confident, red = more confident)

Screenshot 2026-07-21 at 09 53 03

Building confidence - threshold at 0.34 (recommended on Earth Engine demo)

Screenshot 2026-07-21 at 09 53 20

Building confidence - threshold at 0.5

Screenshot 2026-07-21 at 09 53 36

Building confidence - threshold at 0.75

Screenshot 2026-07-21 at 09 53 46

So confidence threshold of 0.5 is arguably too stringent? Perhaps next step to consider the threshold at 0.34.

Other satellite data options

Global Human Settlement Layer (GHSL) data is a composition of worldwide estimates from the EU Copernicus satellite data. Various datasets are provided including population estimates, urban/rural classifications, building heights, surface area and volume, built-up surface area, further split down to identify non-residential surface area (so theoretically you could derive a residential surface area, though I assume there are limitations to this, otherwise GHSL would provide this?).

The main challenge with using GHSL data is the temporal inconsistency. Data is available from 1975 to 2030 (projections), every 5 years, at 100m grid cell resolution. However, there are three years provided in addition, including 2018, provided at 10m resolution. Consequently, the relevance to this work is that we have no data available for 2023, and the data available for 2018 compared to other years is at a different spatial scale (which may also raise differences in quality of estimates at different scales for different time points).

⚠️ **GitHub.com Fallback** ⚠️