Adaptive Geometry-Aware Fourier Neural Operator — with the complete controlled-evidence stack, extended depth sweep to 16, a second PDE family, a deformation baseline, a direct measurement of geometric forgetting, and a fully programmatic research paper…
Model Card
By Sehaj Randhir Singh, published under cc-by-4.0, revision 05b93b94ab6f.
Adaptive Geometry-Aware Fourier Neural Operator — with the complete controlled-evidence stack, extended depth sweep to 16, a second PDE family, a deformation baseline, a direct measurement of geometric forgetting, and a fully programmatic research paper (paper/agfnopaper.pdf). (mode truncation discards everything above the cut). A zero-gated, SDF-derived multiplicative modulation of the spectral weights restores the truncated band by spectral convolution — and the paper measures the whole story: diagnosis (proposition), fix (mechanism), consequence (probe). = 0.971× FNO's global error — the gain is NOT extra parameters (5.10M vs 4.81M) or channels (identical 3-channel inputs). −56% ring.…
Read Sehaj Randhir Singh's full model card
AGF-NO: Restoring the Truncated Band in Spectral Operators
Adaptive Geometry-Aware Fourier Neural Operator — with the complete
controlled-evidence stack, extended depth sweep to 16, a second PDE family,
a deformation baseline, a direct measurement of geometric forgetting, and a
fully programmatic research paper (paper/agfno_paper.pdf).
Core claim: FNO's global mixing channel is structurally band-limited (mode truncation discards everything above the cut). A zero-gated, SDF-derived multiplicative modulation of the spectral weights restores the truncated band by spectral convolution — and the paper measures the whole story: diagnosis (proposition), fix (mechanism), consequence (probe).
1. Attribution result (obstacle Darcy, 3 seeds, identical budgets)
| Variant | rel-L² (global) | rel-L² (ring) | Wall fidelity |
|---|---|---|---|
| FNO baseline | 0.1232 ± 0.0012 | 0.2497 ± 0.0007 | 0.0665 |
| AGF-NO gates-frozen (capacity control) | 0.1197 | 0.2484 | 0.0709 |
| AGF-NO spec-only | 0.1003 | 0.1361 | 0.0069 |
| AGF-NO (full mechanism) | 0.0805 ± 0.0019 | 0.1106 ± 0.0046 | 0.0051 |
| FNO, no penalty | 0.1250 | 0.2577 | 0.0821 |
| AGF-NO, no penalty | 0.0818 | 0.1147 | 0.0091 |
- Capacity control: frozen network (identical graph, gates hard-zeroed) = 0.971× FNO's global error — the gain is NOT extra parameters (5.10M vs 4.81M) or channels (identical 3-channel inputs).
- Mechanism: full AGF-NO beats its own frozen self by −33% global / −56% ring. ~33σ seed separation vs FNO.
- Loss: removing the boundary penalty moves ring error by −3.7%; the advantage is architectural.
2. Depth to 16: collapse vs plateau (the paper's headline)
| depth 4 | depth 6 | depth 8 | depth 16 | |
|---|---|---|---|---|
| FNO rel-L² (global) | 0.1283 | 0.1392 | 0.1444 | 1.0053 (collapse) |
| AGF-NO rel-L² (global) | 0.0813 | 0.1050 | 0.1190 | 0.0846 (depth-stable) |
| FNO probe R² (final block) | 0.567 | 0.475 | 0.502 | 0.0003 (total) |
| AGF-NO probe R² (final block) | 0.626 | 0.607 | 0.652 | 0.633 (plateau) |
At depth 16 the vanilla FNO is worse than predicting the mean, and its latent carries zero linearly decodable geometry — behavioural and representational collapse coincide, as the truncation-band-limitation proposition requires. AGF-NO's probe plateaus within 0.045 R² across depths 4–16 and its ring error improves with depth (0.1163 → 0.1057): depth becomes a resource instead of a liability.
Three-seed replication (all cells re-run, mean ± std):
| depth 4 | depth 8 | depth 16 | |
|---|---|---|---|
| FNO rel-L² | 0.1250 ± 0.0030 | 0.1449 ± 0.0040 | 1.0054 ± 0.0016 |
| AGF-NO rel-L² | 0.0801 ± 0.0010 | 0.1109 ± 0.0058 | 0.0841 ± 0.0012 |
The collapse is essentially deterministic: 11.96× error gap at depth 16 with zero overlap across seeds; FNO's worst-case seed is 1.0036 (still collapsed), AGF-NO's seed spread is 1.4% of its mean. Forgetting at depth is not a tail event — it is what the architecture does. AGF-NO's best cell is depth 4; the honest claim is depth-stable vs depth-collapsing.
3-D replication (the phenomenon is dimension-independent)
The full stack lifts to 3-D obstacle Darcy — spheres and solid tori (multiply-connected, so deformation-based escapes stay structurally unavailable), analytic SDFs, float64 PCG ground truth (every stored field satisfies the discrete PDE to <1e-6 relative residual), zero-gate capacity control unchanged, res 32 → zero-shot 48.
Measured (3 seeds × {fno, agfno} × depth {4, 16}, experiments4/):
| rel-L² (3-D) | depth 4 | depth 16 |
|---|---|---|
| FNO global | 0.144 ± 0.020 | 1.008 ± 0.000 — collapsed, zero seed overlap |
| AGF-NO global | 0.077 ± 0.002 | 0.113 ± 0.003 — depth-stable |
| FNO near-wall | 0.229 ± 0.007 | 0.907 |
| AGF-NO near-wall | 0.085 ± 0.002 | 0.124 (1.45× its depth-4 value) |
The 2-D collapse reproduces unchanged in 3-D: FNO's probe-decodable SDF
information falls 0.494 → 0.020 (R²) from depth 4 to 16 — geometry is
essentially erased from its features — while AGF-NO holds 0.494 at depth 16
(24× more decodable geometry, gate G3). Near-wall error ratio at depth 16:
7.3× (gate G1; FNO's wall violation also explodes 0.034 → 0.771). The
only honest caveat, carried into the paper: AGF-NO's near-wall error does
grow 1.45× from depth 4 to 16 (gate G2 measures stability, not constancy)
— the plateau is not perfectly flat in 3-D, but 0.124 vs FNO's 0.907 is not
a close call. All pre-registered gates pass (ALL_GATES_PASS: true).
3. The forgetting measurement (the novel artifact)
A closed-form linear probe asks, per block: how much SDF information is still linearly decodable from the latent? Shallow sweep (depths 1–4): FNO falls 0.896 → 0.472 monotonically; AGF-NO falls 0.934 → 0.591. The probe ships with unit-tested null (noise → R² ≈ 0) and sanity (linear embeddings → R² > 0.99) controls.
4. Frequency-resolved analysis (P1/P2)
- SDF spectrum: 99.70% of energy below the architectural cut, 0.013% above — the direct tail is tiny; restoration is dominated by multiplicative mixing (measured >20× high-band regeneration; unit-tested, constant modulation = 0).
- P2 confirmed: FNO's high-band ring error 0.0897 → AGF-NO 0.0513 (−43%), high-over-low ratio halved (0.174 → 0.077). The reduction is exactly in the truncated band near walls.
- P1 consistent: wall perturbation influence more wall-confined in FNO (ring/far 1.66) than AGF-NO (1.32).
5. External validity
- Canonical piecewise-constant Darcy (no penalty, no ring loss): FNO 0.1940 / frozen 0.2030 / AGF-NO 0.1450 — ordering replicates with confounds removed. Published references quoted for scale only (FNO 0.0082 @85², Geo-FNO 0.0068; not comparable).
- Deform-FNO baseline: global 0.1058 (beats FNO) but ring 0.2385 (≈ FNO's blindness) — deformation fixes the simply-connected part, cannot touch multiply-connected walls, exactly as predicted.
- Second PDE family (advection–diffusion past fixed-temperature obstacles, zero-shot rollout to 2T): honest near-null — capacity control passes (frozen 1.025× FNO) but mechanism gain is only −2% T / −3.5% ring / −2.5% at 2T. Interpretation in the paper: the mechanism is a targeted fix for wall-anchored difficulty (Darcy's solution is singular at walls; advected thermal layers are smeared downstream). Reported, not hidden.
- A-priori diagnostic (Δρ): predicts where the mechanism helps. Δρ = truncated-band ring-energy fraction of target minus input, on the model's own mode cut. PDE1: +0.0244 (the solve must create high-band wall content) → mechanism gain 0.445. PDE2: −0.0236 (high-band steps are given in the input; diffusion smooths) → gain 0.965 (null). Equal magnitude, opposite sign, matching the observed gains. The naive statistic (target-only ρ) inverts the ranking — documented as a negative result. A new benchmark: compute Δρ before training; it costs two FFTs.
6. Headline run (2,000 samples, 300 epochs, T4)
| Metric | FNO | AGF-NO | Gain |
|---|---|---|---|
| rel-L² (global) | 0.1004 | 0.0541 | −46% |
| rel-L² (ring) | 0.2334 | 0.0627 | −73% |
| Wall fidelity | 0.0769 | 0.0129 | 5.9× |
| Zero-shot 2× super-res | 0.4666 | 0.2295 | −51% |
Repository contents
paper/agfno_paper.pdf— the full 10-page research paper (every number programmatically generated from the JSONs; macros pipeline included).fno_*.pt,agfno_*.pt— headline checkpoints.experiments/— controlled suite: per-variant/per-seed results, ablations, depth sweep, probe (probe/forgetting_probe.json).experiments2/— extended suite: PDE2 matrix + rollout, depth 4/6/8/16, Geo-FNO baseline, extended probe, frequency analysis.experiments3/— 3-seed depth sweep (error bars, replication table).diagnostic/— the a-priori Δρ statistic on both PDE families.experiments4/— 3-D replication: sphere+torus Darcy, depth 4/16, 3 seeds, probe, zero-shot SR to res 48.benchmark/— canonical piecewise-constant Darcy results.results.json,reeval_results.json— headline metrics + local cross-validation (matches remote to 4 decimals).
Reproducibility
One-click public reruns (self-contained kernels, no internet needed): full run · controlled suite · benchmark · extended suite · 3-seed depth sweep · Δρ diagnostic · 3-D replication. Data are byte-deterministic (seed-pinned); 41 unit tests cover the architecture (zero-gate ≡ FNO equivalence, both dimensions), the 2-D and 3-D solvers, SDFs, losses, probe controls, and the band-restoration property.
Honest limitations
Two PDE families, two dimensions (2-D and 3-D synthetic); no external reimplementation at matched settings; probe measures linear decodability (lower bound); PDE2 result shows the mechanism's scope is wall-anchored difficulty; obstacles are static (3-D tori restore multiply-connectedness, but not deforming boundaries over time).
Identity and Version
- Repository
- Sejibeji/agfno-darcy
- Publisher
- Sehaj Randhir Singh
- Task
- Not stated by the source
- Modality
- Other
- Library
- pytorch
- Parameters
- Not stated by the source
- Languages
- pde
- Revision
- 05b93b94ab6f924997f34b7cde1a069640d5b350
- First published
- 2026-09-09
- Last updated
- 2026-09-18
Files and Weights
71 files, 158.6 MB in total. The weights are 4 files totalling 156.7 MB in pt.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| agfno_best.pt | Weights | 40.2 MB | 1871bb0056de |
| agfno_final.pt | Weights | 40.2 MB | 48ae9b04aa09 |
| fno_best.pt | Weights | 38.1 MB | eb80488e2bf7 |
| fno_final.pt | Weights | 38.1 MB | 9ad0104118ce |
| benchmark/summary.json | Configuration | 2.1 KB | — |
| diagnostic/diagnostic.json | Configuration | 1.5 KB | — |
| diagnostic/diagnostic_rho_kernel.json | Configuration | 851 B | — |
| experiments/summary.json | Configuration | 4.7 KB | — |
| experiments2/analysis/frequency_analysis.json | Configuration | 3.9 KB | — |
| experiments2/probe_ext/forgetting_probe.json | Configuration | 12.3 KB | — |
| experiments2/runs/experiments2/pde2_agfno_frozen_s0/run_results.json | Configuration | 267 B | — |
| experiments2/runs/experiments2/pde2_agfno_s0/run_results.json | Configuration | 263 B | — |
| experiments2/runs/experiments2/pde2_agfno_s1/run_results.json | Configuration | 262 B | — |
| experiments2/runs/experiments2/pde2_agfno_s2/run_results.json | Configuration | 264 B | — |
| experiments2/runs/experiments2/pde2_fno_s0/run_results.json | Configuration | 260 B | — |
| experiments2/runs/experiments2/pde2_fno_s1/run_results.json | Configuration | 259 B | — |
| experiments2/runs/experiments2/pde2_fno_s2/run_results.json | Configuration | 260 B | — |
| experiments2/runs/experiments2/summary.json | Configuration | 4.1 KB | — |
| experiments3/runs/experiments3/d16_agfno_s0/run_results.json | Configuration | 156 B | — |
| experiments3/runs/experiments3/d16_agfno_s1/run_results.json | Configuration | 159 B | — |
| experiments3/runs/experiments3/d16_agfno_s2/run_results.json | Configuration | 158 B | — |
| experiments3/runs/experiments3/d16_fno_s0/run_results.json | Configuration | 151 B | — |
| experiments3/runs/experiments3/d16_fno_s1/run_results.json | Configuration | 152 B | — |
| experiments3/runs/experiments3/d16_fno_s2/run_results.json | Configuration | 152 B | — |
| experiments3/runs/experiments3/d4_agfno_s0/run_results.json | Configuration | 157 B | — |
| experiments3/runs/experiments3/d4_agfno_s1/run_results.json | Configuration | 156 B | — |
| experiments3/runs/experiments3/d4_agfno_s2/run_results.json | Configuration | 158 B | — |
| experiments3/runs/experiments3/d4_fno_s0/run_results.json | Configuration | 151 B | — |
| experiments3/runs/experiments3/d4_fno_s1/run_results.json | Configuration | 153 B | — |
| experiments3/runs/experiments3/d4_fno_s2/run_results.json | Configuration | 154 B | — |
| experiments3/runs/experiments3/d8_agfno_s0/run_results.json | Configuration | 155 B | — |
| experiments3/runs/experiments3/d8_agfno_s1/run_results.json | Configuration | 157 B | — |
| experiments3/runs/experiments3/d8_agfno_s2/run_results.json | Configuration | 156 B | — |
| experiments3/runs/experiments3/d8_fno_s0/run_results.json | Configuration | 153 B | — |
| experiments3/runs/experiments3/d8_fno_s1/run_results.json | Configuration | 154 B | — |
| experiments3/runs/experiments3/d8_fno_s2/run_results.json | Configuration | 154 B | — |
| experiments3/runs/experiments3/summary.json | Configuration | 5.0 KB | — |
| experiments4/runs/experiments4/d16_agfno_s0/run_results.json | Configuration | 351 B | — |
| experiments4/runs/experiments4/d16_agfno_s1/run_results.json | Configuration | 350 B | — |
| experiments4/runs/experiments4/d16_agfno_s2/run_results.json | Configuration | 352 B | — |
| experiments4/runs/experiments4/d16_fno_s0/run_results.json | Configuration | 344 B | — |
| experiments4/runs/experiments4/d16_fno_s1/run_results.json | Configuration | 348 B | — |
| experiments4/runs/experiments4/d16_fno_s2/run_results.json | Configuration | 344 B | — |
| experiments4/runs/experiments4/d4_agfno_s0/run_results.json | Configuration | 349 B | — |
| experiments4/runs/experiments4/d4_agfno_s1/run_results.json | Configuration | 351 B | — |
| experiments4/runs/experiments4/d4_agfno_s2/run_results.json | Configuration | 351 B | — |
| experiments4/runs/experiments4/d4_fno_s0/run_results.json | Configuration | 345 B | — |
| experiments4/runs/experiments4/d4_fno_s1/run_results.json | Configuration | 349 B | — |
| experiments4/runs/experiments4/d4_fno_s2/run_results.json | Configuration | 342 B | — |
| experiments4/runs/experiments4/summary.json | Configuration | 5.7 KB | — |
| probe/forgetting_probe.json | Configuration | 4.3 KB | — |
| reeval_results.json | Configuration | 761 B | — |
| results.json | Configuration | 14.6 KB | — |
| train_log.json | Configuration | 13.9 KB | — |
| README.md | Documentation | 9.9 KB | — |
| benchmark/benchmark_comparison.png | Other | 88.0 KB | — |
| comparison.png | Other | 128.3 KB | 503334366c59 |
| experiments/ablation_matrix.png | Other | 67.4 KB | — |
| experiments/forgetting_curve.png | Other | 88.4 KB | — |
| experiments/penalty_ablation.png | Other | 55.7 KB | — |
| experiments2/forgetting_curve_ext.png | Other | 80.7 KB | — |
| experiments2/pde2_bars.png | Other | 38.9 KB | — |
| experiments2/probe_ext/forgetting_probe.png | Other | 104.3 KB | 813de3b26e0a |
| experiments3/forgetting_curve_seeds.png | Other | 75.2 KB | — |
| experiments4/paper/agfno_paper.pdf | Other | 410.1 KB | 711bec880147 |
| experiments4/runs/experiments4/forgetting3d.png | Other | 70.5 KB | — |
| paper/agfno_paper.pdf | Other | 422.9 KB | 12d2fc9ea06c |
| probe/forgetting_probe.png | Other | 94.9 KB | — |
| super_resolution.png | Other | 29.9 KB | — |
| training_curves.png | Other | 82.5 KB | — |
| .gitattributes | Repository | 1.8 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- Open weights, no gate
- Download size
- 156.7 MB
Released by Sehaj Randhir Singh through its official repository on Hugging Face. Read the license.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 156.7 MB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About agfno-darcy
Can I use agfno-darcy commercially?
Yes. agfno-darcy is released under Creative Commons Attribution 4.0. CC BY 4.0 permits sharing and adapting the work, including commercially, provided the creator is credited and changes are indicated.