Brasa — synthetic wildfire
thermal imagery.
Real labeled wildfire thermal data is scarce — fires are dangerous to instrument, aerial campaigns are expensive, and “ground truth” is usually a threshold drawn on the very pixels a model trains on. Brasa is the synthetic alternative, shipped with the evidence that its physics is real.
Every frame is a fully simulated wildfire standing on real bare earth measured by airborne lidar — 1 m posts, ~10 cm vertical, flown over the Kern Plateau, Sierra Nevada, CA. Rothermel fire spread through a 3-D conifer forest, terrain and crowns casting real shadows, imaged through a physically-modeled thermal camera. Labels are projections of the simulated ground truth: pixel-exact, never thresholded from the image.
full 300-frame bundle (99.4 MB) too — get the bundle →
- release
- 01 · brasa · v1.0
- engine
- a30efa8b · 2026-07-15
- license
- CC BY-NC 4.0
- cite as
- Brasa — ay4la.com/brasa
Run the engine yourself.
The same engine, compiled to WebAssembly and running entirely in your browser — no upload, nothing leaves your machine. Type a seed and it samples a whole wildfire scenario, solves the fire, and images it through the modeled thermal camera in a few seconds.
Then scrub the fire's age, change the hour and wind, toggle the radiometric truth against the deployed-camera view — plateau histogram equalization, the same code, to the line, that our sim-to-real eval runs on real fire — and download the frame in the dataset's exact schema. That AGC is one representative 8-bit-camera pipeline of several, not a specific camera's firmware; a detector trained on one AGC may not transfer to another.
One honest difference: Brasa Live stands on 30 m SRTM terrain, not the 1 m lidar the download stands on — a lidar block is gigabytes and will not cross a browser. Same physics, same sensor, coarser ground.
desktop-first · ~25 MB one-time terrain load
All 12 stages reproduce an independent reference — 9 scored against real data or a reference code, 3 against an analytic law.
Synthetic data is easy to make and hard to trust. Brasa's rule is that no physics stage is believed until it reproduces something we didn't make — a reference radiative-transfer code, a published fire law, a USGS survey, 738 frames of real wildfire thermal imagery. Three are pulled out below, including the one we are worst at:
Real FLAME 3 fires, pooled over all fire frames: 2.71 px. Ours: 2.62 px.
Pooled across different fire sizes, so fill and diameter partly compensate — inside matched diameter bins the agreement is looser. Not fitted; the bin-by-bin breakdown ships in the record.
Max |ΔTb| against libRadtran 2.0.6, the reference radiative-transfer code, at the standard atmospheres the sky model is calibrated from — every zenith angle.
At its reference atmospheres it is the same answer, not approximately right. Blind interpolation to unsampled humidity — leave-one-out across the six-point humidity grid — holds to ≤ 3.7 K.
Below 4 m we are half again as rough as real forest floor — 1.77 K against 1.15 K.
It passes, and it is the soft spot. We say which term we think is missing.
That third number is on the page for the same reason the other two are: it is what the measurement said. It ships in the bundle's certificate too, under a heading called What is known to be wrong. A synthetic dataset that only publishes its best gates is asking to be trusted; one that publishes its worst is telling you where not to.
No black box.
The “measurement” is not a filter over a picture — it's a chain of named, individually-validated effects. Watch one fire frame lose exactly one thing at a time: optical blur (Airy PSF), detector sampling, 40 mK temporal noise, fixed-pattern noise, then the 14-bit ADC. The final panel is bit-identical to the delivered frame.


featureless white blob. that loss is the domain gap, and we model the camera that causes it.
Validated against independent references.
No physics stage is trusted until it reproduces an independent reference. Each figure carries its own method box — reference, metric, threshold — printed by the same code the validation gates run on. In the table below, each gate is tagged by what it is checked against: real data, a reference code, or an analytic law.
- Clear-sky LWIR radiance vs libRadtran RT · reference codemax |dTb| 0.001 K · < 0.5 K · vs libRadtran 2.0.6
- Lidar georeference vs USGS tile corners · real data0.49 m · < 1 m · vs USGS 3DEP tile bboxes
- Fire Tb distribution vs real FLAME · real datawithin 7-11 K · vs FLAME 3
- Fire micro-texture: fill fraction · real data0.035 · 0.03-0.08 · vs 0.048
- Fire micro-texture: within-fire Tb CV · real data0.136 · 0.12-0.18 · vs 0.154
- Fire micro-texture: Tb p10-p90 spread (K) · real data191 · 130-200 · vs 173
- Fire front width, pooled (px) · real data2.62 · vs 2.71
- Shadow thermal lag vs analytic conduction wave · analytic lawreproduces · vs 1-D analytic conduction wave (A0 e^-z/d)
- Rothermel spread vs hand-traced + published grass ROS · analytic lawreproduces · vs Rothermel 1972 / Anderson 1982
- Flaming residence time · analytic lawexact · vs Anderson 1969 (384/sigma)
- Background texture, crown scale (4-20 m) · real data2.28 · <= 2x · vs 2.69
- Background texture, litter scale (< 4 m) · real data1.77 · <= 2x · vs 1.15
The crowns are the stand FIA would describe -- real sizes, and porous.
Every conifer is DBH-sampled from the stand's diameter distribution and sized by the operational FVS Sierra Nevada allometry: crown radius, height and crown base are one consistent tree, all published, none fitted. A DBH-60 cm overstory tree is a ~28 m conifer with a ~4 m-radius crown, not the uniform placeholder earlier versions carried. And each crown is porous -- a medium of needles, not an opaque solid -- so the sun's beam is attenuated through it by Beer-Lambert over the path length L through foliage. Because a crown is convex, that path runs from ~0 at the silhouette to the full depth on the axis: the shadow is dark in the middle and FEATHERS over a few metres, dappling the floor rather than stencilling it. A hard-edged shadow is a step, and a step is broadband -- it sprays crown-scale structure down into the sub-4 m band where our background was too rough.
Nothing here is tuned to the gate. Crown radius, height and crown ratio are the FVS Western Sierra Nevada equations; canopy density is anchored to the reference (Sierra mixed-conifer cover 40-70%; ~52% stand cover here, 91 stems/ha overstory). The leaf area density (1.0 m2/m3, mid published conifer range) was fixed before the gate was re-measured and lands mid-band on two stand-scale cross-checks it was never fitted against: crown LAI 4.0 (published 3-8) and sub-canopy beam transmittance 13.5% (published 5-15%).
Compared at a matched physical scale, against ground that has trees on it.
A texture filter defined in pixels measures a different thing on every camera: 25 px is 3.9 m of ground at FLAME 3's 120 m flight height and 6-29 m at Brasa's 150-900 m, so the same tree-crown shadow can fall on either side of it. This gate's scale split is fixed in metres of ground and converted per frame through that frame's GSD. It is scored in two bands -- litter (< 4 m) and crown (4-20 m) -- because trees add structure at crown scale and not below it. And it is scored against real FORESTED ground: FLAME 3's no-fire frames are treeless marsh, so the reference is taken from its fire frames with the fire masked out, leaving real forest floor under real crowns.
Below 4 metres we are half again as rough as real forest floor.
1.77 K against 1.15 K -- 1.54x. Inside the band, and our weakest gate. The cause is measured, not guessed: turn each effect off in turn and re-render, and the sub-4 m band is made by tree-crown shadows and nothing else.
The search has now cleared every easy answer. Not the painted sub-metre clutter or terrain shadows (0.00 K each). Not a surface-energy-balance term like latent flux: it has no length scale, so it would drag both bands together and break the crown gate to fix the litter one. Not the shadow edge -- making the crowns porous fixed that (1.83x to 1.57x). Not the forest floor -- smoothing the ground micro-relief moves it under 0.05 K. And now not crown GEOMETRY: rebuilt to real FVS allometry (correct size, height and count), the litter band barely moved (1.57x to 1.54x) while the crown band improved to 0.85x of real.
next — That the litter band did not move under correct crowns is the finding. The leading lead -- a lead, not a claim -- is diffuse fill: under a real canopy, skylight and multiple scattering between crowns soften the shadow contrast, so a closed forest floor is more uniformly lit than a model casting discrete, independently-attenuated crown shadows produces. A different physics from anything tried so far, and the next thread.
Does it transfer to real fire?
The test that matters: train a detector on Brasa frames only, then score it on 622 fire + 116 no-fire real FLAME 3 frames the detector never saw. It is run as a matrix — 2 grounds × 3 scene conditions, 6 arms in all — because a single number would hide what the answer depends on. Two sensor domains separate, and they answer differently.
Frame-level fire / no-fire, saturated — and it held there in all 6 arms, so nothing measured below moved it. Shown as a baseline, not a result: absolute temperature makes fire nearly separable, and a plain maximum-temperature threshold scores AUC 1.000 on the same frames. No detector is needed to win this column.
Per-frame auto-gain destroys absolute radiance, so the temperature cue is gone — and the equivalent brightness baseline collapses to AUC 0.500, chance. Whatever score remains is shape and texture, not temperature. This is the honest frontier, and the open research question.
| AGC domain · ground | linework off | as shipped | not drawn |
|---|---|---|---|
| lidar-1m33 scenario windows | 0.804TPR 0.254 | 0.810TPR 0.256 | 0.840TPR 0.482 |
| srtm-30m244 scenario windows | 0.887TPR 0.449 | 0.924TPR 0.600 | 0.942TPR 0.725 |
The published M14 numbers were measured on SRTM ground with linework OFF. Only that arm is a like-for-like comparison; the lead column (linework on) describes a different world and must not be compared to them directly.
AUC 0.774 → 0.887 (+0.113)TPR@≤5% 0.384 → 0.449 (+0.065)
M18 measured linework-on as a regression to 0.203 and its ablation chain attributed the damage to night-background thermal flatness. M20 built that background; measured on the same condition today the arm scores far above the M18 value, so the regression is retired by measurement rather than by assertion.
ablationRoads and waterways earn their keep as world physics and give part of it back as appearance. What they do to the world — break fuel continuity, place ignitions at roadsides, keep trees out of road beds — is worth +0.228 / +0.276 TPR (lidar-1m / srtm-30m). Drawing the ribbons into the frame costs −0.227 / −0.125, so on lidar ground the two very nearly cancel. That residual appearance tax is small, isolated and now measurable — which makes it the next milestone rather than a caveat.
detectionretired, not re-measured — The AP50 0.10-vs-0.28 / ~37% retention figures came from a separate box-level lane fed by a different generator. M7.5 showed that metric is convention-bound — hand-labelling the eval alone made BOTH models score worse — so it is not restated for this build.
distributionthe synthetic tracks the real temperature distribution — background within 7 K, median fire within 11 K, fire↔background contrast within 4.7 K
how to use it — the split above is the guidance. If your camera hands you calibrated radiance, the radiometric column is saturated and Brasa is training data. If it hands you 8-bit auto-gained frames — which is most deployed cameras — the AGC column is where you actually are, and Brasa is pretraining and augmentation data: pretrain on Brasa, then fine-tune on a small real set. Nothing measured here supports replacing real data.
The two grounds are not a controlled comparison: the lidar block yields 33 scenario windows against 244 for SRTM, so terrain fidelity and scene diversity move together here and cannot be separated from these runs. Read the rows as two worlds measured, not as a ranking of grounds. frame-level ROC against FLAME 3's own Fire/No-Fire labels; frame score = max detection confidence. No box-derivation convention anywhere near the headline number (the Track C confound lesson).
provenance — engine 73349c60, one commit past the a30efa8b the gate table above was measured at; that commit adds the ablation switches this matrix needs. Some arms carry a `-dirty` stamp in their raw arm.json: the working tree changed during the run (documentation and analysis tooling), while engine/ did not. Frames are a pure function of (engine code, seed), so those arms are reproducible from this commit; the stamp was over-conservative and its scope has since been corrected. Checked rather than asserted: seeds regenerated from a dirty-stamped arm and byte-diffed against the originals — bit-identical.
the open question — a physically-modeled AGC stage inside the synthesis loop, so a detector trained purely on synthetic frames transfers to real deployed cameras
The sample dataset.
The sample is a deliberate evaluation slice: 48 of the bundle's 300 frames (41 fire / 7 no-fire, 20 night / 28 day, 194 boxes) — sized for inspecting labels, verifying the radiometric format, and smoke-testing a training pipeline before you commit to the full bundle. Every sample file is a byte-identical copy of its bundle counterpart, emitted by the same packager run and byte-compared by a live certificate gate — so the sample you smoke-test and the bundle you train on cannot version-skew.
Each frame is 16-bit radiometric PGM (deci-kelvin — pixel/10 = Tb K, so fire cores are represented, not clipped), with COCO detection labels carrying the physical ground truth: FRP, fire area, plume height, peak Tb, contrast. Frames with no visible fire ship as labeled negatives.








sha256 7dfddf29…d7393142
license: CC BY-NC 4.0 — for research & noncommercial use, with attribution.
commercial use — including deploying models trained on this data — requires a license.
The full bundle — 300 frames.
v1.0: 300 frames (270 train / 30 val) across seeded scenarios — terrain, weather, ignition cause, time of day — with YOLO + COCO labels, the full provenance manifest, and the validation certificate baked in. 99.4 MB zipped (211 MB unpacked), hosted on Hugging Face, same CC BY-NC 4.0 terms as the sample.
sha256 1155b480…2bef656c — full hash on the Hugging Face card.
And v1.0 doesn't have to be the end of it — Brasa is a deterministic generator; every frame regenerates bit-identically from (engine version, seed):
- ›Bigger runs or different scenario mixes — on request, CC BY-NC 4.0
- ›Your sensor: custom LWIR profile (GSD, NETD, ADC) baked into the render
- ›Your scene: wildfire ships today; the radiometry and sensor stack render any LWIR scene — a new scene class is a scoped validation, not a free swap
- ›Commercial use & model deployment — licensed, priced per ask
custom bundles typically ship within days — generation is deterministic and fast.
Three reference sensors.
Every number below is computed from the sensor model — none typed by hand. These three are reference profiles that bracket the model's range — GSD 0.82→13.6 m/px, pixel-limited→diffraction-limited optics: one resolves the fire, one sees it sub-pixel. The v1.0 bundle's own frames are rendered across 150–900 m altitude, not off these exact three. Your sensor spec is a profile away.
| spec | tower-LWIR | airborne-LWIR | longrange-LWIR |
|---|---|---|---|
| role | fixed early-detection tower | airborne nadir mapper | long-range wide-area scanner |
| band | 8–12 µm | 8–12 µm | 8–12 µm |
| range | 1000 m | 2000 m | 20000 m |
| fov | 24° | 35° | 40° |
| detector | 512 px | 1024 px | 1024 px |
| gsd | 0.818 m/px | 1.193 m/px | 13.635 m/px |
| f/# | 1 | 1.2 | 2 |
| regime | pixel-limited | diffraction-limited | diffraction-limited |
| netd | 50 mK | 40 mK | 60 mK |
| adc | 14-bit | 14-bit | 14-bit |
| frame rate | 30 Hz | 60 Hz | 10 Hz |
The honest part.
A number without its caveats is marketing. These ship with every bundle, alongside the validation certificate:
- ›The litter-scale background band runs 1.54x rougher than real forest floor. It passes, and it is where we are weakest. The cause is measured -- crown shadows -- and the next lead is named above.
- ›Two engine stamps on this page, kept apart rather than merged. The gate table was measured at a30efa8b; the sim-to-real transfer matrix and the radiometric distribution gap at 73349c60, the next commit, which adds the ablation switches that matrix needs. The distribution figures changed materially when they were re-measured -- background gap 7.8 -> 6.6 K, fire-to-background contrast 3.5 -> 4.7 K -- and the cause is staleness, not regression: they had not been re-emitted since before M24 replaced a fixed downwelling-longwave constant with a temperature-dependent sky, so the old values described an engine five milestones behind the one they were published beside. The background cooled toward the real reference, which is an improvement. Contrast is fire minus background and the fire did not move at all, so contrast widened by exactly the same 1.2 K -- the metric got worse because the physics got more right.
- ›The background and contrast figures pool day and night frames, and those two populations sit about 31 K apart, so the pooled median is sensitive to how many day pixels survive the 330 K cut that defines background -- a cut that selects on the very quantity it then reports. Measured separately, our night background sits within about 2 K of the real reference and our day background runs about 34 K warmer. Whether that is a daytime defect or a comparison with no counterpart in the reference depends on FLAME 3's acquisition conditions, which we have not confirmed. Reporting this gate split by day and night, and defining background by a non-fire mask instead of a temperature threshold, is open work and is not done.
- ›Labels are fire REGIONS. Residual smouldering hotspots inside a burn scar merge into their scar at a 16 m ground distance rather than each becoming an object; a genuine spot fire, clear of the burn, stays its own box. A positive box marks an active-combustion region -- flaming or still-smouldering -- not a new ignition, so a warm burn scar with live embers is a true positive, not a false one.
- ›FLAME 3 measures one world (forested prescribed understory burns), so open grass and brush fire behaviour cannot be scored against it -- only asserted from Rothermel.
- ›NETD is quoted at 300 K; the 14-bit ADC is linear in radiance over a fixed 260-900 K span (headroom so fire cores are not clipped), so near ambient the quantization step is coarser than the NETD -- the NETD is the detector spec, not the delivered background floor.
- ›The sensor chain is static-platform -- optical PSF, detector sampling, NETD, fixed-pattern noise (PRNU/DSNU/column), then the 14-bit ADC -- with no motion or scan smear modelled, so along-track resolution on the moving airborne and long-range profiles is optics- and detector-limited only.
- ›Background warm clutter is modelled -- solar-loaded roads and waterways, terrain insolation, and material-dependent surface temperature -- but no non-fire source is driven to fire-adjacent apparent temperature, so frame-level fire/no-fire is near-separable by absolute temperature (this is why the radiometric AUC saturates at 1.000, a baseline rather than a result). A hard negative at ember-or-fire brightness -- a sun-baked rock or hot artefact pushed into the fire Tb range -- is a next scene addition, not a solved case.
deterministic provenance: every frame regenerates bit-identically from (engine version, seed). the numbers on this page are imported from the engine's machine-derived results files — never re-typed.