AirGap

Method

How these 20 sites were chosen

Five layers, from estimating air quality across Indonesia to presenting it on this page.

Data period:29 June 20231 June 2026· 54 representative dates spread across that span

Terms used on this page

LOCO
Leave-One-Cell-Out. One 5 km grid cell is withheld entirely from training, then estimated as if it had never existed.
Spatial blocking
A whole cluster of neighbouring cells (within 50 km) is withheld together, rather than one cell at a time.
Sigma
How much the ten models disagree about a cell. The larger it is, the more the model is guessing.
Mahalanobis
How unfamiliar a location's environmental conditions are compared with everywhere ever measured.
Regime
A group of environmental condition types. There are 15, formed from elevation, density, land use, and emissions.
Variogram
A measure of how far one sensor still speaks for its surroundings. For our data that distance is 104.4 km.
Submodular greedy
Picking sites one at a time from the highest score, damping the scores of nearby areas after each pick.
1

PM2.5 Estimation

LightGBM model, 18 features, square-root target

Tujuan
Estimate PM2.5 everywhere that has no sensor.
Cara
Indonesia is divided into 77,226 five-kilometre cells. A LightGBM model learns from 19,455 daily observations using 18 features: weather, emissions, elevation, land use, population density, satellite imagery, and seasonal markers. The target is square-root transformed and squared back afterwards.
Hasil
PM2.5 estimates for all 77,226 cells, averaged across 54 representative dates.
Dampak
A national pollution map covering the whole landmass, not just the area around sensors.

Why 158 labelled cells shrink to 52 used for training: 158 cells carry a PM2.5 label, 157 of them fall inside the land grid, and only 52 have labels dated inside the satellite feature window so they can be paired up.

2

Uncertainty Quantification

Ensemble sigma + Mahalanobis + regime clustering

Tujuan
Find out where the model is guessing, not merely what it guesses.
Cara
Ten models are trained separately with resampling at the spatial-block level rather than per row, so environmental diversity between models is preserved. Every cell is also measured for Mahalanobis distance against conditions already observed, then grouped into 15 regimes.
Hasil
A sigma map, an unfamiliarity threshold of 6.13, and 15 regimes — 8 of which hold no sensor at all, covering 41,058 cells or 53.2% of the country.
Dampak
Ignorance becomes measurable and mappable, not merely admitted.
3

Placement Optimisation

Submodular greedy + spatial updating

Tujuan
Choose the 20 sites that add the most knowledge, not the busiest ones.
Cara
Of 77,226 cells, 50,813 pass the eligibility filter: at least 50 residents and no higher than 3,000 metres. Each candidate receives a weighted score — unfamiliarity 0.30, uncertainty 0.20, population 0.15, pollution 0.15, minus overlap 0.20. After each pick, nearby cells have their scores damped by distance.
Hasil
20 sites reaching 10 of 15 area types with a closest pair of 94.6 km, covering 561,912 people.
Dampak
The sensors are spread out by design, not by luck.
4

Validation

Leave-K-blocks-out + head-to-head against two strategies

Tujuan
Test whether this selection is genuinely better, not merely tidy-looking.
Cara
Two separate tests. First, whether adding sensors at our sites lowers model error — repeated 8 times with different block splits. Second, whether the spread is more efficient than random placement and city-based placement.
Hasil
Error falls 2.1% (±5.4), and 18 of 20 sensors stand alone.
Dampak
The claim of superiority carries supporting numbers, including its confidence limits.
5

Interface

The page you are reading

Tujuan
Let anyone check the results, not only the team.
Cara
Every result is computed once during preparation, then stored as map images and compact files. This page calls no server when it opens.
Hasil
A map clickable down to the 5 km cell, in two languages, with no backend.
Dampak
It stays light over constrained connections — an important consideration for use outside major cities.

The main finding: how you test decides what you find

A model can look excellent if you test it the wrong way. We tested with three progressively stricter protocols, and quote the strictest as our official figure.

ProtocolRMSEMAEBias
CV acak10 seeds

The same cell may appear in both training and test data. This is the easiest exam.

0.641± 0.00213.215± 0.0357.959± 0.017-1.298± 0.022
LOCO10 seeds

One cell is withheld entirely. The model is tested somewhere it has never seen.

0.269± 0.04118.843± 0.52512.527± 0.4440.181± 0.932
Blokir spasial10 seeds

A whole cluster of neighbouring cells is withheld together, so the model cannot copy from next door.

0.150± 0.01020.330± 0.11514.026± 0.108-1.242± 0.200

Values are mean ± standard deviation across seeds. R²: higher is better (maximum 1). RMSE and MAE: lower is better. Bias: closer to zero is better.

The 0.491 gap is a finding in itself

R² falls from 0.641 on the easiest exam to 0.150 on the strictest. Conventional validation inflates apparent performance almost fourfold on spatially clustered data. That gap is itself our methodological contribution.

See why on the map: switch on the spatial-blocks layer and note that those 52 training cells are really only 16 independent locations.

Open the map

Does every feature actually earn its place

The two derived features give a slight but real improvement. We report it as it stands, including the two variants that made things worse.

Feature variantRMSE
A 16 fitur dasar0.155620.26
B + rh, aod_per_blh0.162320.18
C B tanpa mutu_sensor0.122920.65
D B + bobot blok0.011421.92

What drives the model's estimates

Each feature's contribution to the model's decisions, out of 18 used. The top four are highlighted.

Does unfamiliarity predict model error?

One dot per training cell: Mahalanobis distance on the horizontal axis, model error on the vertical.

01223351.02.33.54.86.02.08 → 13.21.17 → 19.42.10 → 7.62.38 → 9.21.69 → 6.71.56 → 9.31.59 → 6.71.71 → 12.72.31 → 6.61.80 → 7.32.12 → 10.02.14 → 11.42.99 → 10.81.18 → 16.51.94 → 19.41.91 → 24.93.66 → 15.33.35 → 13.32.32 → 12.33.70 → 12.72.01 → 16.73.60 → 15.72.18 → 10.91.88 → 15.52.36 → 21.62.47 → 18.72.24 → 19.72.53 → 29.81.68 → 18.31.33 → 14.71.37 → 18.71.62 → 30.81.75 → 9.61.51 → 25.11.83 → 25.01.38 → 8.01.49 → 22.91.49 → 7.31.44 → 8.61.60 → 8.11.60 → 8.12.09 → 7.51.78 → 7.51.44 → 13.31.89 → 5.71.40 → 12.41.97 → 15.31.24 → 12.11.59 → 12.41.35 → 11.81.95 → 10.65.45 → 6.8MahalanobisModel error (MAE)

The shape is a U, not a rise. The correlation is -0.004 with p=0.976 across 52 cells — not significant. We did not succeed in showing that Mahalanobis predicts error, and we display that as it stands. It is still used because its independence test against sigma passes (-0.146), so the two measure different sources of ignorance.

How far one sensor speaks for its surroundings

104.4 km · 1,326 location pairs · out-of-fold (Lapis 2.1, 160 model)

This 104.4 km figure is what serves as the overlap threshold during validation.

A note on data sources

Sensor data comes from three sources of differing quality: OpenAQ as the best reference, ISPU KLHK, and PurpleAir, which are low-cost community sensors. The source is recorded as a flag in the model. When computing the national map, every cell is treated as if measured by the best instrument so results are not biased downward.

Limitations We Acknowledge

The following six points are reported without softening. Hiding them would make the rest of this report fair to doubt.

  1. 1

    R² on the strictest protocol is only 0.150

    Far below our initial target of 0.6, which came from a comparison study in a single city. This model is useful for revealing patterns and prioritising sites, not for replacing measurement.

  2. 2

    The Mahalanobis validation is not significant

    Its correlation with model error is -0.004 (p=0.976) across 52 cells, and the relationship is U-shaped rather than rising. We still use it because the independence test against sigma passes, but the claim that it predicts error is unproven.

  3. 3

    Satellite AOD contributes almost nothing

    Ranked 15th of 18 features at roughly 1% contribution. The cause is tropical cloud cover, which leaves daily availability at only 5–31%.

  4. 4

    Boundary layer height behaves against theory

    Its correlation with PM2.5 is positive (+0.114) where theory predicts negative. Our suspicion: daily aggregation erases the diurnal cycle, while the physical mechanism operates in the night and morning boundary layer.

  5. 5

    Fire hotspot features are unavailable

    Blocked by compute quota, despite being physically likely to explain the extreme PM2.5 spikes in Sumatra and Kalimantan.

  6. 6

    Regression to the mean remains strong

    National predictions top out at 45.9 µg/m³ while training data reaches 428.4. The model consistently under-estimates heavily polluted locations.