Ratio change metadata - datasciencecampus/WorldPop-Malawi-fork GitHub Wiki
- malawi_ratio_change.R on ons-compatibility-ratiochange branch
The ratio change method is a simple method to update census population estimates for Enumeration Areas (EA) in Malawi (persons, households, etc) by using the ratio change of a separate data source as a proxy for change in the population
For our initial implementation of the ratio change method, we take the ratio of buildings estimated in each Malawi EA from Google footprint data from 2018 and 2023. We then apply this ratio to the 2018 Census estimate of households for each EA, to give an updated estimated number of households for 2023.
For example, a Malawi EA may have:
- Google footprint data - count of buildings in 2018 = 100, count of buildings in 2023 = 150
- Ratio of Google footprint data from 2023-2018 = 1.5
- The estimate number of households from the 2018 Census = 300
- Ratio change of Census households = 300 * 1.5 = 450
Quality of estimates is measured by comparing the updated household estimate to the number of households estimated from the most recent survey data, the absolute error is simply:
- number of households from the survey - number of households estimated from ratio change
- Three surveys are available to us with household estimates for EAs; the ICT (2023), the IHS (December 2024) and the NACA (August 2024). Most EAs only have an estimate from one source , however where there are estimates from multiple surveys, the most recent one is taken.
- Are we certain on the quality of the survey estimates? Is one survey better than another?
We can also get the absolute error of the WorldPop modelled estimate number of households. This means we can compare the WorldPop model to the ratio change method by summarising the absolute error between the two methods. We calculate:
- Mean Absolute Error (MAE) - average of absolute error values across all Malawi EAs
- Variance of Absolute Error - standard deviation of absolute error values across all Malawi EAs
- Max Absolute Error - maximum absolute error value from all Malawi EAs
We can also summarise performance from the ratio change vs WorldPop method by summarising the estimated number of households for each Malawi EA. For example:
- Maximum EA household estimate
- Minimum EA household estimate
- Total number of EAs over a given threshold (300 households for Urban EAs, 200 households for Rural EAs), which we also express as a percentage
- Total number of EAs under a given threshold (100 households for both Urban and Rural EAs)
| Method | max_abs_error | sd_abs_error | mean_abs_error |
|---|---|---|---|
| WorldPop model | 1705 | 102.6641497 | 77.43739635 |
| Ratio change | 776.6196441 | 64.94645663 | 65.45033271 |
| Method | max_abs_error | sd_abs_error | mean_abs_error |
|---|---|---|---|
| WorldPop model | 4151 | 265.7733351 | 215.039267 |
| Ratio change | 540.7403809 | 65.69975957 | 66.62777277 |
| Metric | Max | Min | Sum_over_threshold | Sum_less_100hh | Total | Too_big_perc | Not_too_big_count |
|---|---|---|---|---|---|---|---|
| WorldPop model | 4763 | 34 | 2187 | 7 | 2770 | 78.95306859 | 583 |
| Ratio change | 1592.963902 | 0.987518261 | 684 | 74 | 2767 | 24.71991326 | 2083 |
| Metric | Max | Min | Sum_over_threshold | Sum_less_100hh | Total | Too_big_perc | Not_too_big_count |
|---|---|---|---|---|---|---|---|
| WorldPop model | 4114 | 0 | 8369 | 1749 | 15961 | 52.43405802 | 7592 |
| Ratio change | 1552.81559 | 0 | 7792 | 1694 | 15944 | 48.87104867 | 8152 |
Alternate satellite data sources are available to use instead of the original Google Open Buildings dataset (https://sites.research.google/gr/open-buildings) where we are concerned that the temporal aspect of this data (i.e., comparing metrics from 2018 to 2023) may be blurred due to multiple years worth of data being combined for the '2023' estimates.
Instead, Google Temporal Buildings dataset (https://sites.research.google/gr/open-buildings/temporal) provides estimates of buildings from satellite imagery from 2016 through to 2023. This dataset has 3 measures; 1) fractional building count, 2) building height, and 3) building presence, all of which are provided for 4m pixels across Malawi.
Fractional building count provides an estimate as to how much of a fraction of a building is within a given pixel. Building height is the height in metres from the ground, and is capped at 100m. Building presence is effectively a confidence score, ranging from 0-1, to determine how confident the model is that a given pixel contains a part of a building. Screenshots are provided at the end to show the different data layered onto part of Malawi satellite image.
From these measures, we calculate two important metrics; 1) building count estimate, and 2) total building area (building volume can also be created, this is to be added).
Building count estimate is taken by summing together all of the fractional building counts across all pixels within EA boundaries which should, in theory, give the total number of buildings within an EA (note that the method applies a scaling factor correction, so the raw fractional counts are scaled, I need to understand this methodology more and will add detail here when ready).
Total building area starts with the building presence measure, takes only pixels that have a value above 0.5 (i.e., the model is more confident than not that the given pixel has a building part in it), and then multiplies the total number of remaining pixels within an EA by 16 (16 metre squared, because each pixel is 4m x 4m). This gives total building area of an EA.
Because the data is available from 2016 to 2023, we are able to take the years 2018 and 2023 to feed into the same ratio change methodology, substituting in these two sets of estimates (which I call 'Building Count Temporal' and 'Total Building Area Temporal'). By calculating the absolute error and number of buildings estimated as above, we can compare the ratio change method using these new variables to previous results:
<style> </style>| Urban | |||
|---|---|---|---|
| Method | max_abs_error | sd_abs_error | mean_abs_error |
| WorldPop model | 4151 | 265.77334 | 215.03927 |
| Building Count Original | 540.7404 | 65.69976 | 66.62777 |
| Building Count Temporal | 588.317 | 64.84683 | 71.81328 |
| Total Building Area Temporal | 642.2436 | 75.5704 | 72.24901 |
| Rural | |||
| Method | max_abs_error | sd_abs_error | mean_abs_error |
| WorldPop model | 1705 | 102.66415 | 77.4374 |
| Building Count Original | 776.6196 | 64.94646 | 65.45033 |
| Building Count Temporal | 800.3693 | 60.86491 | 60.67796 |
| Total Building Area Temporal | 952.0508 | 69.94239 | 64.87916 |
| Urban | ||||||||
|---|---|---|---|---|---|---|---|---|
| Metric | Max | Median | Min | Sum_over_threshold | Sum_less_100hh | Total | Too_big_perc | Not_too_big_count |
| WorldPop model | 4763 | 402 | 34 | 2187 | 7 | 2770 | 78.95306859 | 583 |
| Building Count Original | 1592.963902 | 241.76 | 0.987518261 | 684 | 74 | 2767 | 24.71991326 | 2083 |
| Building Count Temporal | 913.69 | 255.95 | 0.99 | 845 | 67 | 2767 | 30.54 | 1922 |
| Total Building Area Temporal | 1393.84 | 247.43 | 0.97 | 736 | 70 | 2767 | 26.6 | 2031 |
| Rural | ||||||||
| Metric | Max | Median | Min | Sum_over_threshold | Sum_less_100hh | Total | Too_big_perc | Not_too_big_count |
| WorldPop model | 4114 | 207 | 0 | 8369 | 1749 | 15961 | 52.43 | 7592 |
| Building Count Original | 1552.82 | 197.33 | 0 | 7792 | 1694 | 15944 | 48.87 | 8152 |
| Building Count Temporal | 1467.77 | 194.03 | 0 | 7589 | 1727 | 15969 | 47.52 | 8380 |
| Total Building Area Temporal | 1807.81 | 195.21 | 0 | 7664 | 1696 | 15960 | 48.02 | 8296 |
Consequently, we can say that using either of the Google satellite data for building counts produces similar outputs, leaving to ratio change estimates that perform better than when using Google satellite total area, which in turn is better than the current WorldPop model. However, there are a few things to consider/explore with the Google temporal data:
-
The confidence metric for pixels - ChatGPT assumes a pixel over 0.5 is a valid enough threshold. The academic paper and demonstration code accompanying Earth Engine instead suggested a threshold of 0.34, but it is said "This however might need tweaking for a given location and time since the model confidence varies based on a number of factors such as cloud cover, imagery alignment, etc.". Therefore, I think there is more work to do to explore this temporal dataset further, fine-tuning the confidence metric to ensure that we are selecting valid data for our context
-
How consistent is the confidence metric across areas in Malawi - e.g. urban vs rural areas. Are there likely to be more pixels with higher confidence estimates in urban areas because the likelihood of there being more buildings in urban areas, compared to rural areas?
-
Validate outputs - I need to check that the estimates roughly correspond to the number of buildings visible on satellite images. Though we can gain confidence from the fact that the different Google satellite sources are similar (though would we expect the temporal data to differ, given our concerns about the original Google data?). Validation will be tricky though, because Earth Engine only provides one detailed satellite image, so I can't access detailed images from 2018 and 2023 for example to verify both years.
To start with, some datasets have been mentioned but brief reasons outlined here as to why these haven't been considered:
-
Global Human Settlement Layer (GHSL) data is a composition of worldwide estimates from the EU Copernicus satellite data. Various datasets are provided including population estimates, urban/rural classifications, building heights, surface area and volume, built-up surface area, further split down to identify non-residential surface area (so theoretically you could derive a residential surface area, though I assume there are limitations to this, otherwise GHSL would provide this?). The main challenge with using GHSL data is the temporal inconsistency. Data is available from 1975 to 2030 (projections), every 5 years, at 100m grid cell resolution. However, there are three years provided in addition, including 2018, provided at 10m resolution. Consequently, the relevance to this work is that we have no data available for 2023, and the data available for 2018 compared to other years is at a different spatial scale (which may also raise differences in quality of estimates at different scales for different time points).
-
OpenStreetMap (OSM) is a volunteered geographic information (VGI) product where a community of users and volunteers provide information about buildings, using uploaded satellite imagery and their local knowledge of the area. As this is dependent on volunteers, the coverage of information is very patchy, with some academic work suggesting that buildings counts using only OSM are substantially lower than other products. There may also be concerns when using as a measure of change, what is the coverage like from 2018 compared to 2023/2026? I would hazard a guess that coverage would be better in later years for several reasons - better access to technology, greater demand/need for geospatial information.
Next I present comparisons of outputs from the ratio change method, with several options considered in using measure of change of buildings. These include the options already presented:
- Google Open Buildings - count
- Google Open Temporal Buildings - count
- Google Open Temporal Buildings - total area
Plus the following latest additions:
- Google Temporal Buildings smoothed - count
- Google Temporal Buildings smoothed - total area
- Google Temporal Buildings smoothed - volume
The Google Temporal Buildings smoothed is a product from academics (Zhang et al. A temporally harmonised gridded building-stock dataset for the Global South, 2016-2023) who take the raw Google Open Temporal Buildings data, and try to address inconsistencies in the time series of this data. For example, satellite imagery over time may have changes in observation conditions (cloud cover, haze, atmospheric conditions, seasonal vegetation, viewing geometry). There may also be changes in satellite data processing/modelling. This led to some cases of temporal instability and "zig-zag" amplitudes, where changes of building metrics for a given area would change drastically from one year to the next. Instead, it may be expected that any change would be gradual. Consequently, a methodology was developed to smooth the time series of building metrics (count, area and volume) for the global south.
The raw Google pixel data were aggregated to 100m pixels for this analysis, which the smoothed 100m grid estimates made publicly available. I took this data (for counts, area and volume), aggregated that data to Malawi EAs for 2018 and 2023, and passed this data through the ratio change pipeline.
Below I present the same metrics from before (absolute error, and summary of estimated number of households within EAs, separately for urban and rural EAs).
<style> </style>| Method | max_abs_error | sd_abs_error | mean_abs_error |
|---|---|---|---|
| WorldPop | 4151 | 265.7733351 | 215.039267 |
| Ratio Change - original - count | 540.7403809 | 65.69975957 | 66.62777277 |
| Ratio Change - temporal - count | 588.3169847 | 64.84683133 | 71.81328449 |
| Ratio Change - temporal - area | 642.2435746 | 75.57039655 | 72.24901358 |
| Ratio Change - smoothed - count | 466.5951258 | 67.16950375 | 71.80745337 |
| Ratio Change - smoothed - area | 743.4253243 | 78.09700279 | 76.15261418 |
| Ratio Change - smoothed - volume | 928.4899419 | 86.16713446 | 73.99321049 |
So using smooth temporal count data leads to better performance in terms of reduced max error, but mean error is best using original Google Open Buildings data. Area and volume tend to perform worse.
<style> </style>| Method | max_abs_error | sd_abs_error | mean_abs_error |
|---|---|---|---|
| WorldPop | 1705 | 102.6641497 | 77.43739635 |
| Ratio Change - original - count | 776.6196441 | 64.94645663 | 65.45033271 |
| Ratio Change - temporal - count | 800.3693031 | 60.8649133 | 60.67796163 |
| Ratio Change - temporal - area | 952.0508114 | 69.94239442 | 64.87916437 |
| Ratio Change - smoothed - count | 736.5958989 | 65.74520155 | 63.77889113 |
| Ratio Change - smoothed - area | 1139.847608 | 73.31376983 | 64.63212926 |
| Ratio Change - smoothed - volume | 1310.543831 | 78.46916417 | 66.58149336 |
Best in terms of max error is smoothed count data or original Google data. Mean error very similar. Area and volume much worse in terms of max error.
<style> </style>| Min HH count | Median HH count | Max HH count | % EAs over threshold | % EAs under 100 | Total EAs | |
|---|---|---|---|---|---|---|
| WorldPop | 34 | 402 | 4763 | 78.95 | 0.25 | 2770 |
| Ratio Change - original - count | 0.9875183 | 241.7554 | 1592.9639 | 24.72 | 2.67 | 2767 |
| Ratio Change - temporal - count | 0.9699394 | 247.4297 | 1393.8428 | 26.60 | 2.53 | 2767 |
| Ratio Change - temporal - area | 0.9956518 | 255.953 | 913.6911 | 30.54 | 2.42 | 2767 |
| Ratio Change - smoothed - count | 1.0146817 | 256.1538 | 1577.1393 | 30.32 | 2.31 | 2767 |
| Ratio Change - smoothed - area | 0.9924106 | 257.9221 | 1556.3914 | 31.77 | 2.28 | 2767 |
| Ratio Change - smoothed - volume | 0.9700229 | 236.791 | 2872.534 | 24.50 | 2.57 | 2767 |
Original Google data is best performing overall, but smoothed volume data is also comparable here.
<style> </style>| Min HH count | Median HH count | Max HH count | % EAs over threshold | % EAs under 100 | Total EAs | |
|---|---|---|---|---|---|---|
| WorldPop | 0 | 207 | 4114 | 52.43 | 10.96 | 15961 |
| Ratio Change - original - count | 0 | 197.3256 | 1552.816 | 48.87 | 10.62 | 15944 |
| Ratio Change - temporal - count | 0 | 195.2137 | 1807.811 | 48.02 | 10.63 | 15960 |
| Ratio Change - temporal - area | 0.01211886 | 194.0269 | 1467.77 | 47.52 | 10.81 | 15969 |
| Ratio Change - smoothed - count | 0 | 211.2961 | 1658.41 | 54.52 | 8.05 | 15956 |
| Ratio Change - smoothed - area | 0 | 210.6304 | 1896.408 | 54.09 | 7.60 | 15967 |
| Ratio Change - smoothed - volume | 0 | 200.7425 | 2294.24 | 50.36 | 8.42 | 15967 |
Some interesting patterns. WorldPop model similar to ratio change. Over threshold, the original and temporal data perform slightly better. The smoothed data performs better looking at number EAs below threshold.
- Census data - treating unoccupied households, households of multiple occupation (both of which we can work out from the structures/all regions datasets)
- Measuring uncertainty - adopt WP growth model? Adapt ratio change to incorporate some kind of simulation to produce confidence intervals?
- Further exploration of satellite imagery - including conversations with WP/Soton experts
At a glance, this doesn't look like it covers the entirety of a building?
So confidence threshold of 0.5 is arguably too stringent? Perhaps next step to consider the threshold at 0.34.
Global Human Settlement Layer (GHSL) data is a composition of worldwide estimates from the EU Copernicus satellite data. Various datasets are provided including population estimates, urban/rural classifications, building heights, surface area and volume, built-up surface area, further split down to identify non-residential surface area (so theoretically you could derive a residential surface area, though I assume there are limitations to this, otherwise GHSL would provide this?).
The main challenge with using GHSL data is the temporal inconsistency. Data is available from 1975 to 2030 (projections), every 5 years, at 100m grid cell resolution. However, there are three years provided in addition, including 2018, provided at 10m resolution. Consequently, the relevance to this work is that we have no data available for 2023, and the data available for 2018 compared to other years is at a different spatial scale (which may also raise differences in quality of estimates at different scales for different time points).