Methodological and data platform for the collection, expansion, and curation of training and validation samples for the new 10 m pasture mapping of the Cerrado biome (Product 1 · Project GR000153-73077-CARGILL-BRAZIL). Integrating 701 Embrapa Cenargen field reference polygons (~1 ha; 608 classified + 93 unlabeled), 62,009 candidate points (50,000 persistent Sentinel-2 + 5,453 Col. 4 S2 + 6,556 Landsat validation), monthly Sentinel-2 series (2020–2025), 64-dimensional AlphaEarth Foundations embeddings, SOM topological filtering, and ~2 m CBERS-4A quality control.
The high spatial, temporal, and ecological heterogeneity of pastures in the Cerrado biome demands a traceable strategy to expand sample sets from high-quality field references, filter pseudo-labeling errors in complementary representation spaces, and rigorously document uncertainty.
Structured candidate universe combining 5,453 Sentinel-2 Col. 4 points, 6,556 Landsat validation points, and 50,000 persistent pasture points (2020–2025; ~35% low, ~44% medium, ~21% high vigor). Sensitivity analysis compares the inclusive universe (62,009) against the Collection 11 agreement scenario (56,952 points).
AlphaEarth Foundations annual embeddings (64 latent dimensions at 10 m with L2-renormalization) combined with monthly Sentinel-2 spectral series (NDVI, EVI, NDII, NDWI, PSRI, CAI) extracted via a 60-day moving window using the joint median of valid pixel-date pairs.
Two-stage filtering via L2 cosine similarity against 608 classified field reference polygons, Top-1/Top-3 md3 concordance, 12×12 SOM clustering with Bayesian smoothing, and stratified ~2 m CBERS-4A visual inspection before Random Forest training.
Detailed technical overview of the sampling plan, representation spaces, two-stage filtering, quality control protocols, and machine learning roadmap under the WRI Brasil · FUNAPE · LAPIG/UFG agreement (Report Product 1 — August 17, 2026).
The project operates under the technical agreement between WRI Brasil and FUNAPE, with technical execution by LAPIG/UFG (Project GR000153-73077-CARGILL-BRAZIL), under the scientific coordination of Prof. Dr. Laerte G. Ferreira and statistical consulting by Prof. Luis Rodrigo Fernandes Baumann (IME/UFG).
Field Reference Base (701 Polygons): 701 in situ polygons (~1 ha). Currently, 608 polygons have consolidated labels across pasture condition categories, while 93 remain unclassified (retained for unsupervised tasks).
Candidate Universe (62,009 Points) & Independence Caveats:
Spatial Support & Scenarios: Candidate extraction uses a 3×3 pixel window (900 m² / 0.09 ha), carrying spectral mixing risks at field boundaries. The Agreement Scenario (corroborated by Col. 11) totals 56,952 candidates, while the Inclusive Scenario comprises 62,009 candidates.
Monthly Sentinel-2 Time Series (2020–2025): Extracted on Google Earth Engine over a centered moving window of ~60 days, excluding observations contaminated by clouds and shadows (cloud/shadow masking audit pending). The monthly value is calculated as the joint median of all valid pixel-date pairs:
Covering NDVI, EVI, NDII, NDWI, PSRI, and CAI. A final audit of valid observation counts, edge months, and the resampling of native 20 m bands to 10 m is currently in progress.
Stage 1 (Embedding Similarity — Preliminary): Cosine similarity against 608 classified field references, assigning Top-1 provisional labels, Top-3 nearest neighbors, and concordance index md3 ∈ {1, 2, 3} across five sensitivity scenarios (Scenarios A–E). All thresholds remain exploratory and subject to final calibration.
Stage 2 (Self-Organizing Maps — Under Development): 12×12 Kohonen grid (144 neurons) organizing multivariate time series:
High-Resolution Verification: High-resolution visual inspection of pre-filtered candidates and borderline cases using ~2 m nominal resolution pan-sharpened CBERS-4A imagery (WPM sensor).
Sample Sizing with Finite Population Correction: Designed using the finite population correction formula based on an illustrative filtered universe of N = 40,000 candidate points:
Statistical Qualification: The N = 40,000 figure is strictly illustrative, and the 3% margin of error at 95% confidence applies solely to an overall binary proportion (e.g., valid pasture vs. non-pasture / correct vs. incorrect pseudo-label), not to estimating per-category accuracy across the 7 individual condition classes.
Blind Multi-Interpreter Protocol: Candidate pseudo-labels are concealed from interpreters. Inter-interpreter agreement will be quantitatively measured, with systematic adjudication for disagreements.
Machine Learning Pasture Models: Assembly of the final curated training database with documented provenance and confidence flags. Models will employ Random Forest (10 m) partitioned by spatial block cross-validation (spatial k-fold) to prevent spatial autocorrelation leakage.
In Situ Field Campaign & Precision Limits: Planned operational pilot of ~100 field localities across the Cerrado, designed in collaboration with IME/UFG. Logistical priority is allocated to Goiás, with complementary transects across biome gradients.
Independent Validation: Post-classification statistical evaluation combining high-resolution imagery and in situ field observations to estimate overall map accuracy.
To guarantee scientific integrity and reproducibility, the project integrates proactive controls across seven core methodological risk areas (Report Product 1, Table 5):
Preserve candidate provenance; evaluate model scenarios with and without Collection 11 constraints; base final validation on independent high-resolution imagery and field observations.
Track per-category retention rates; calibrate category-specific similarity thresholds; prevent artificial absorption of minority categories.
Compute homogeneity and edge metrics for the 3×3 candidate windows (900 m²) to detect and flag mixed boundary pixels.
Anchor reference labels strictly to their field collection year prior to multi-year propagation across 2020–2025 series.
Explicitly compute and unit-normalize L2 norms on spatial mean embeddings to eliminate non-unit dot product distortion.
Segregate field reference support from candidate pseudo-label frequencies in the SOM to prevent circular reinforcement loops.
Pair the Goiás-centered field pilot with biome-wide stratified satellite inspection (CBERS-4A ~2 m) to ensure balanced statistical inference across all Cerrado gradients.
In alignment with the concluding roadmap of Product 1, the immediate operational phase requires the resolution of eight technical milestones:
Consolidate a single versioned dictionary for the 7 condition categories, formalizing operational criteria and reconciling the 608 classified field samples.
Finalize verification of cloud/shadow masks, minimum valid observations, edge month behavior, and resampling of native 20 m bands to 10 m (NDII, PSRI, CAI).
Complete the recomputation of similarity matrices using verified L2-renormalized embedding means.
Finalize empirical thresholds for Top-1, Top-3, md3, and separation margins (s₁ − s₂) across Scenarios A through E.
Define final SOM grid parameters, evaluate input dimensionality (12×J vs 72×J), and calibrate Dirichlet-Multinomial vs Normal-Normal smoothing.
Draw the stratified quality control sample (n ≈ 1,040 points) and execute the blind multi-interpreter inspection protocol.
Construct the approved training dataset and evaluate spectral-temporal separability among condition categories.
Train Random Forest 10 m pasture models with spatial block partitioning and execute independent image- and field-based validation.
Open data, processing routines, technical reports, and interactive spatial analytics.
701 in situ field reference areas (~1 ha polygons) surveyed across the Cerrado biome (608 classified samples and 93 unclassified units retained for unsupervised steps). Dictionary Reconciliation & Integrity Rule: An initial 5-class proposal is being reconciled with the 7 implemented categories to construct a single versioned dictionary; category distributions remain provisional until this formal reconciliation is completed.
Tracking of methodological components, current execution status, and next technical deliverables.