LAPIG/UFG · WRI Brasil

Improve Pasture Mapping in the Cerrado Biome

Methodological and data platform for the collection, expansion, and curation of training and validation samples for the new 10 m pasture mapping of the Cerrado biome (Product 1 · Project GR000153-73077-CARGILL-BRAZIL). Integrating 701 Embrapa Cenargen field reference polygons (~1 ha; 608 classified + 93 unlabeled), 62,009 candidate points (50,000 persistent Sentinel-2 + 5,453 Col. 4 S2 + 6,556 Landsat validation), monthly Sentinel-2 series (2020–2025), 64-dimensional AlphaEarth Foundations embeddings, SOM topological filtering, and ~2 m CBERS-4A quality control.

Historical Series 2020 – 2025 (6 Years) Cerrado Biome · 62,009 Candidates · 701 References Sentinel-2 & AlphaEarth 10 m · CBERS-4A ~2 m WRI Brasil · FUNAPE · LAPIG/UFG


62,009
Candidate Points (50k Persistent + 12k Candidates)
701
Field Reference Areas (608 Classified + 93 Unlabeled)
6 Years
Historical Series (2020 – 2025)
7
Pasture Condition Categories
64 Dim
AlphaEarth Latent Embeddings (10 m)
Context and Rationale

Challenges in Mapping Cerrado Pastures & Sample Curation

The high spatial, temporal, and ecological heterogeneity of pastures in the Cerrado biome demands a traceable strategy to expand sample sets from high-quality field references, filter pseudo-labeling errors in complementary representation spaces, and rigorously document uncertainty.

Sample Expansion at Scale
62,009 Candidates across the Cerrado

Structured candidate universe combining 5,453 Sentinel-2 Col. 4 points, 6,556 Landsat validation points, and 50,000 persistent pasture points (2020–2025; ~35% low, ~44% medium, ~21% high vigor). Sensitivity analysis compares the inclusive universe (62,009) against the Collection 11 agreement scenario (56,952 points).

Dual Representation Spaces
64-Dim Embeddings & S2 Time Series

AlphaEarth Foundations annual embeddings (64 latent dimensions at 10 m with L2-renormalization) combined with monthly Sentinel-2 spectral series (NDVI, EVI, NDII, NDWI, PSRI, CAI) extracted via a 60-day moving window using the joint median of valid pixel-date pairs.

Quality Control & Validation
Cosine Similarity, SOM & CBERS-4A

Two-stage filtering via L2 cosine similarity against 608 classified field reference polygons, Top-1/Top-3 md3 concordance, 12×12 SOM clustering with Bayesian smoothing, and stratified ~2 m CBERS-4A visual inspection before Random Forest training.



Methodological Framework & Technical Progress

Methodological Architecture & Quality Assurance Framework

Detailed technical overview of the sampling plan, representation spaces, two-stage filtering, quality control protocols, and machine learning roadmap under the WRI Brasil · FUNAPE · LAPIG/UFG agreement (Report Product 1 — August 17, 2026).

⚠️
Scope & Status Note (Product 1 Progress Report): This documentation details the methodological architecture and progress established as of August 17, 2026. While reference databases, Sentinel-2 multi-index series, and AlphaEarth embeddings are assembled and extracted, all downstream components — including similarity thresholds, SOM neural configurations, visual inspection protocols, final training dataset assembly, and field campaign executions — represent preliminary analyses, work under active development, or planned stages.
📍 Assembled
1. Reference & Candidates
701 Reference Polygons + 62,009 Candidate points
🛰️ Extracted
2. S2 Series & Embeddings
60-day medians & 64-dim L2 vectors (2020–2025)
📐 Preliminary
3. Embedding Similarity
Top-1/Top-3, md3 & Sensitivity Scenarios A–E
🧠 In Development
4. SOM Neural Filtering
12×12 grid (144 neurons) & Bayesian smoothing
🔍 Planning
5. CBERS-4A Inspection
Stratified sample (n ≈ 1,040) on ~2 m imagery
🌲 Planning
6. ML & Field Pilot
Random Forest 10 m & ~100 field localities
🏛️
Institutional Governance · Component 01

1. Governance, Scope & Strategic Goals

The project operates under the technical agreement between WRI Brasil and FUNAPE, with technical execution by LAPIG/UFG (Project GR000153-73077-CARGILL-BRAZIL), under the scientific coordination of Prof. Dr. Laerte G. Ferreira and statistical consulting by Prof. Luis Rodrigo Fernandes Baumann (IME/UFG).

Strategic Objective: Develop the sampling and analytical infrastructure to produce the first version of the new 10-meter resolution pasture map for the Cerrado biome, specifically designed to differentiate among viable pasture condition categories and quantify sampling uncertainty across the biome.
WRI Brasil · FUNAPE LAPIG / UFG 10 m Cerrado Pasture Map
📍
Databases & Independence · Status: Assembled

2. Ground Truth Reference & 62,009 Candidates

Field Reference Base (701 Polygons): 701 in situ polygons (~1 ha). Currently, 608 polygons have consolidated labels across pasture condition categories, while 93 remain unclassified (retained for unsupervised tasks).

Dictionary Reconciliation & Integrity Rule: There is an active inconsistency between an initial 5-class proposal and the 7-class implementation; a single unified dictionary is currently being constructed. An explicit integrity rule prevents presenting category distributions as definitive until this reconciliation is formalized. Polygons include partially intercepted pixels in full (without area-weighting). The 93 unclassified samples may enter unsupervised SOM training, but must not be used to assign neuron labels or vote in Top-k decisions.

Candidate Universe (62,009 Points) & Independence Caveats:

  • 5,453 visually inspected pasture points (MapBiomas Col. 4 S2): Carries methodological and spatial dependence from MapBiomas classification algorithms.
  • 6,556 points (Landsat validation database, ~85k dataset): Critical Caveat: If incorporated into model training, these points lose independence and can no longer serve for validation.
  • 50,000 persistent pasture points (2020–2025): Generated as remote sensing pseudo-labels (stratified proportionally to 2025 vigor: ~35% low, ~44% medium, ~21% high vigor), not in situ ground truth.

Spatial Support & Scenarios: Candidate extraction uses a 3×3 pixel window (900 m² / 0.09 ha), carrying spectral mixing risks at field boundaries. The Agreement Scenario (corroborated by Col. 11) totals 56,952 candidates, while the Inclusive Scenario comprises 62,009 candidates.

701 Reference Polygons 62,009 Candidates 56,952 Agreement 3×3 Window (900 m²) Integrity Rule
🛰️
Dense Remote Sensing · Status: Extracted

3. S2 Time Series & 64-Dim Embeddings

Monthly Sentinel-2 Time Series (2020–2025): Extracted on Google Earth Engine over a centered moving window of ~60 days, excluding observations contaminated by clouds and shadows (cloud/shadow masking audit pending). The monthly value is calculated as the joint median of all valid pixel-date pairs:

Ĩ(i, k, m) = median { I(i, k, p, t) : p ∈ grid(i), t ∈ window(m) }

Covering NDVI, EVI, NDII, NDWI, PSRI, and CAI. A final audit of valid observation counts, edge months, and the resampling of native 20 m bands to 10 m is currently in progress.

AlphaEarth Foundations Latent Embeddings: 64-dimensional annual latent vectors (A00–A63) per 10 m pixel. Because spatial averaging of unit vectors reduces resultant vector norm, an explicit L2-renormalization (||ē|| = 1) will be enforced upon recomputation before calculating cosine similarity. Embeddings are temporally anchored to the field collection year.
Sentinel-2 (10 m) 60-Day Moving Window Joint Median Ĩ(i,k,m) 64-Dim Embeddings L2-Renormalization Pending
🧠
Filtering Architecture · Status: Preliminary & In Dev

4. Cosine Similarity & 12×12 SOM Filtering

Stage 1 (Embedding Similarity — Preliminary): Cosine similarity against 608 classified field references, assigning Top-1 provisional labels, Top-3 nearest neighbors, and concordance index md3 ∈ {1, 2, 3} across five sensitivity scenarios (Scenarios A–E). All thresholds remain exploratory and subject to final calibration.

Stage 2 (Self-Organizing Maps — Under Development): 12×12 Kohonen grid (144 neurons) organizing multivariate time series:

  • Consistency Scope: Because candidate points vastly outnumber reference points, the SOM functions as a spectral-temporal consistency filter, not an independent validation.
  • Dimensionality: Single-index annual vectors are 12 values; with J indices, each annual block is 12×J, expanding to 72×J for the 6-year multi-year series (2020–2025).
  • Bayesian Smoothing: Evaluates neighborhood coherence using Dirichlet-Multinomial Bayesian smoothing (with Normal-Normal conjugate models evaluated as an alternative) to triage candidates (Keep, Inspect, Reject).
  • Quality Controls: Standardizing each variable, missing data handling, monitoring quantization and topographic errors, seed initialization, and strict reference segregation.
Top-1 / Top-3 & md3 Scenarios A–E 12×12 SOM (144 Neurons) 12×J / 72×J Dimensions Dirichlet-Multinomial
🔍
High-Resolution Quality Control · Status: Planning

5. Stratified Inspection on CBERS-4A (~2 m)

High-Resolution Verification: High-resolution visual inspection of pre-filtered candidates and borderline cases using ~2 m nominal resolution pan-sharpened CBERS-4A imagery (WPM sensor).

Scope Limitation: This visual inspection serves as a training-sample quality control filter, does not constitute map validation (which occurs after the cartographic product is completed), and cannot substitute for in situ field observation across all pasture condition categories.

Sample Sizing with Finite Population Correction: Designed using the finite population correction formula based on an illustrative filtered universe of N = 40,000 candidate points:

n₀ = (1.96² · 0.5 · 0.5) / 0.03² ≈ 1,067.11
n = 1,067.11 / (1 + 1,066.11 / 40,000) ≈ 1,039.43 ≈ 1,040 points

Statistical Qualification: The N = 40,000 figure is strictly illustrative, and the 3% margin of error at 95% confidence applies solely to an overall binary proportion (e.g., valid pasture vs. non-pasture / correct vs. incorrect pseudo-label), not to estimating per-category accuracy across the 7 individual condition classes.

Blind Multi-Interpreter Protocol: Candidate pseudo-labels are concealed from interpreters. Inter-interpreter agreement will be quantitatively measured, with systematic adjudication for disagreements.

CBERS-4A (~2 m) Finite Correction (n ≈ 1,040) Binary Proportion Margin (3%) Blind Protocol
🌲
Machine Learning & Field Validation · Status: Planning

6. Random Forest (10 m) & Field Pilot

Machine Learning Pasture Models: Assembly of the final curated training database with documented provenance and confidence flags. Models will employ Random Forest (10 m) partitioned by spatial block cross-validation (spatial k-fold) to prevent spatial autocorrelation leakage.

In Situ Field Campaign & Precision Limits: Planned operational pilot of ~100 field localities across the Cerrado, designed in collaboration with IME/UFG. Logistical priority is allocated to Goiás, with complementary transects across biome gradients.

Statistical Limitation: A sample of n ≈ 100 points yields a margin of error of approximately ±8 percentage points for an expected overall accuracy around 80% (before design effect). Consequently, this field campaign functions as an operational pilot for calibration, not a definitive per-category accuracy assessment for the entire biome.

Independent Validation: Post-classification statistical evaluation combining high-resolution imagery and in situ field observations to estimate overall map accuracy.

Random Forest (10 m) Spatial Block Partitioning ~100 Field Localities Margin ±8 pp (Pilot Scale)
🛡️
Quality Assurance & Risk Controls · Table 5 of Report

7. Methodological Risk Controls & Traceability Framework

To guarantee scientific integrity and reproducibility, the project integrates proactive controls across seven core methodological risk areas (Report Product 1, Table 5):

Risk 01 · Circularity

Preserve candidate provenance; evaluate model scenarios with and without Collection 11 constraints; base final validation on independent high-resolution imagery and field observations.

Risk 02 · Class Imbalance

Track per-category retention rates; calibrate category-specific similarity thresholds; prevent artificial absorption of minority categories.

Risk 03 · Spatial Mixing

Compute homogeneity and edge metrics for the 3×3 candidate windows (900 m²) to detect and flag mixed boundary pixels.

Risk 04 · Temporal Change

Anchor reference labels strictly to their field collection year prior to multi-year propagation across 2020–2025 series.

Risk 05 · Norm Distortion

Explicitly compute and unit-normalize L2 norms on spatial mean embeddings to eliminate non-unit dot product distortion.

Risk 06 · Pseudo-Label Domination

Segregate field reference support from candidate pseudo-label frequencies in the SOM to prevent circular reinforcement loops.

Risk 07 · Regional Inference

Pair the Goiás-centered field pilot with biome-wide stratified satellite inspection (CBERS-4A ~2 m) to ensure balanced statistical inference across all Cerrado gradients.

Table 5 Risk Controls Circularity Prevention Class Imbalance Mitigation Vector Norm Correction Audit-Ready Traceability
📋
Concluding Technical Roadmap

Upcoming Decisions and Technical Deliverables

In alignment with the concluding roadmap of Product 1, the immediate operational phase requires the resolution of eight technical milestones:

1
Class Dictionary Reconciliation

Consolidate a single versioned dictionary for the 7 condition categories, formalizing operational criteria and reconciling the 608 classified field samples.

2
Sentinel-2 Series Audit

Finalize verification of cloud/shadow masks, minimum valid observations, edge month behavior, and resampling of native 20 m bands to 10 m (NDII, PSRI, CAI).

3
Renormalized Cosine Similarity

Complete the recomputation of similarity matrices using verified L2-renormalized embedding means.

4
Threshold & Margin Sensitivity

Finalize empirical thresholds for Top-1, Top-3, md3, and separation margins (s₁ − s₂) across Scenarios A through E.

5
SOM Configuration & Smoothing

Define final SOM grid parameters, evaluate input dimensionality (12×J vs 72×J), and calibrate Dirichlet-Multinomial vs Normal-Normal smoothing.

6
CBERS-4A Sample Draw

Draw the stratified quality control sample (n ≈ 1,040 points) and execute the blind multi-interpreter inspection protocol.

7
Training Dataset Assembly

Construct the approved training dataset and evaluate spectral-temporal separability among condition categories.

8
Model Training & Validation

Train Random Forest 10 m pasture models with spatial block partitioning and execute independent image- and field-based validation.

Roadmap Deliverables Dictionary Harmonization 20 m to 10 m Resampling SOM Calibration CBERS-4A Draw 10 m Model Training
Products & Deliverables

Product 1 Deliverables & Technical Assets

Open data, processing routines, technical reports, and interactive spatial analytics.

🗺️
Candidate Sample Base (62,009 Points)
50k persistent Sentinel-2 samples + 12,009 MapBiomas samples (5,453 Col. 4 S2 + 6,556 Landsat) across 2020–2025
📍
Field Reference Base (701 Polygons)
701 in situ polygons (~1 ha; 608 classified into 7 categories + 93 unlabeled) with extracted embeddings and S2 series
Cosine Similarity & md3 Metrics
64-dimensional L2-renormalized dot product, Top-1/Top-3 matching, and md3 concordance scores in Parquet, CSV, and JSON
💻
Interactive Web Map & Inspector
GPU-accelerated Canvas viewer, multi-filters (50k / 12k, years, CVP vigor, md3 slider), spatial bounding box, and charts
🧠
SOM Topological Clustering
12×12 SOM network organizing multivariate time series with Dirichlet-Multinomial Bayesian smoothing (keep/inspect/reject)
📄
Technical Report (Product 1)
Detailed Sampling Plan and Sampling Improvement Strategies report, sensitivity analyses (Scenarios A–E), and CBERS-4A protocol


Pasture Condition Categories

Pasture Condition Classification Scheme

701 in situ field reference areas (~1 ha polygons) surveyed across the Cerrado biome (608 classified samples and 93 unclassified units retained for unsupervised steps). Dictionary Reconciliation & Integrity Rule: An initial 5-class proposal is being reconciled with the 7 implemented categories to construct a single versioned dictionary; category distributions remain provisional until this formal reconciliation is completed.

Productive Pasture (Pasto Produtivo)
High forage grass coverage, vigorous density, and absence of severe degradation signs.
Pasture with Weeds (Pasto com Ervas)
Pasture with moderate to severe infestation of herbaceous and ruderal invasive plants.
Pasture with Shrubs (Pasto com Lenhosas)
Significant presence of shrub and woody canopy layers across the pasture support area.
Intermediate Pasture (Intermediário)
Intermediate vegetative vigor stage, with mixed coverage of forage and bare soil.
Biological Degradation (Deg. Biológica)
Prominent bare soil, visible erosion signs, and severe loss of carrying capacity.
Natural Regeneration (Reg. Natural)
Areas undergoing natural vegetative recovery and succession of native Cerrado physiognomies.
Miscellaneous (Miscelânea)
Atypical, transitional, or mixed compositions that do not fit into standard categories.

Project Status & Activities

Status of Activities (Product 1 — August 17, 2026)

Tracking of methodological components, current execution status, and next technical deliverables.

Component 01 · Status: Assembled
1. Field Reference Base
701 polygons of approximately 1 ha; 608 classified across 7 pasture condition categories and 93 without a consolidated class.
Component 02 · Status: Extracted
2. Sentinel-2 Time Series
Monthly composites (2020–2025) for references and candidates using ~60-day moving window joint medians; audit of valid observations in progress.
Component 03 · Status: Extracted
3. Annual Latent Embeddings
64-dimensional AlphaEarth Foundations vectors per year (2020–2025); verification and L2 norm correction of mean vectors applied before similarity calculation.
Component 04 · Status: Assembled
4. Candidate Points Universe
62,009 candidates with provenance preserved; sensitivity scenarios structured with and without the 5,057 points that disagree with MapBiomas Collection 11.
Component 05 · Status: Preliminary (Active Milestone)
5. Embedding-Based Filtering
Top-1/Top-3 similarity matches, concordance metric (md3), separation margins, and sensitivity thresholds evaluated across Scenarios A through E.
Component 06 · Status: Under Development
6. SOM-Based Filtering (Self-Organizing Maps)
Initial 12×12 neuron grid (144 units); multivariate dimensionality, normalization, neighborhood, and Dirichlet-Multinomial Bayesian smoothing under evaluation.
Component 07 · Status: Planning
7. CBERS-4A Visual Inspection (~2 m)
Stratified sampling (~1,040 points for 3% margin of error at 95% confidence) after embedding-SOM selection with blind multi-interpreter protocol.
Component 08 · Status: Planning
8. Field Data Collection & Subsequent Validation
Initial campaign of approximately 100 field localities across the Cerrado (logistical priority given to Goiás) for targeted validation of the 10 m map.