diff --git a/CLAUDE.md b/CLAUDE.md index 7a0ebc6..d9d583e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -493,8 +493,9 @@ should be *all three ported, two more added*; and `Renderer::Impl`'s destructor The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC collision operator**, written from scratch in Rust (not a port): flow past bodies in -a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall -models, and GIF output locked to *physical* flow time. `Cargo.toml` declares +a channel, two interchangeable backends, a chain of nested AMR levels (`--levels`; +the old single `--refine`/`--patch` is its one-link case), sub-grid wall models with +an optional Spalding wall function, and GIF output locked to *physical* flow time. `Cargo.toml` declares **Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the method authors' papers in `docs/origins/`; formula references in the code follow the 2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models @@ -512,14 +513,16 @@ defaults** — most defaults are the outcome of a documented measurement, not a cd docs/theory/2d_solver cargo build --release # with the GPU backend (default feature `gpu`) cargo build --release --no-default-features # CPU only, no wgpu -cargo test --release # 29 fast tests -cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s) +cargo test --release # 36 fast tests +cargo test --release -- --include-ignored # + 5 long CPU runs (~35 s) cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic ``` -33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the -4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two -diagnostics, not four benchmarks. +41 `#[test]` functions exist (23 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`, 3 in +`main.rs`); 5 are `#[ignore]`d — four in `cpu.rs` (two reference benchmarks, two +diagnostics) and `levels_match_single_patch_on_steady_cylinder` in `main.rs`. The +`main.rs` tests build a real `Spec` through `Cli::try_parse_from` + `build_spec` — the +only tests that drive a whole `cpu::Sim`. **Never build `--no-default-features` last.** Both builds write the same `target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and @@ -551,7 +554,7 @@ tests. | `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. | | `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. | | `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. | -| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. | +| `src/main.rs` | CLI, problem assembly (incl. `parse_levels`), step loop, live output, report, CSV. Owns the `Spec` / `PatchSpec` / `LevelGeo` / `StepRec` / `FieldKind` contract shared by both backends. | | `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. | ## Invariants (read before editing) @@ -598,6 +601,33 @@ tests. layout has none, which is the whole reason it is kept. Both give identical numbers; the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to anything other than `0`. +- **Levels are a chain, stepped recursively.** `Spec.patches: Vec` (each + box in its *parent's* coordinates) replaced `refine`/`patch`; `Spec::levels()` derives + every level's size, scale, origin and τ (`τ_k = r_k(τ_{k−1} − ½) + ½`). Both + backends step level k, then r_k substeps of k + 1 (frame interpolated in time), + then restrict; force is taken on the finest level at every one of its substeps, + walls/inlet/outlet/sponges exist on L0 only. `--size`, `--nx`, `--ny`, `--steps` + and Re are all in **L0** cells/steps, so L0 is the level closest to τ → ½. The + single-patch path is bitwise identical to the pre-1.4.0 code on both backends + (verified on five setups) — keep it that way when touching the recursion. +- **GPU force slots scale with substeps.** `results` holds `HIST·SLOT` (SLOT = 28: + 12 stats + 12 per-body force + ρ at the probe + 3 spare) plus `substeps·MAXB·4` + force work slots. Before 1.4.0 there were slots for exactly two substeps, so with + `--refine ≥ 3` force of later substeps fell off the buffer and GPU Cd came out + r/2 times too low — any GPU result with `--refine 3` from ≤ 1.3.0 is wrong. +- **β is per node**, not per column (`math::beta_field`), so `--sponge-side` can add + lateral sponges; with `--sponge-side 0` the values equal the old column profile. +- **Wall function (`--wall-function spalding`) uses the same (u_node, τ_w) for every + scheme.** Geometry (`cpu::WfNode`: normal from the SDF gradient, y_w, a sample + point one cell further along the normal, its bilinear stencil) is built once and + uploaded. `grad`/`hrr` take u_node into the target velocity and u_τ²/ν into the + pressure tensor; `bouzidi` gets a **stress-matched** wall slip + `u_s = u_node − u_τ²·y_w/ν`, identically zero in the viscous sublayer. Two earlier + slip formulas were wrong and are documented in the README (velocity-matched slip + made the wall nearly free-slip, Cd 0.016; resolved node velocity broke the no-slip + limit by 2 %). The moment-wall and wall-function pipelines (group 3: wall nodes, + `GWf`, slips) are only built when needed and require 11 storage bindings per stage + (27 with split populations). Every channel run reports y⁺ (`yplus_*`, `u_tau_mean`). - **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but per-body force is bucketed — bodies beyond the fourth have their forces merged (`min(body, MAXB-1)` in three places), and the CLI warns when that happens. @@ -612,12 +642,20 @@ tests. ## Validation campaign (`bench/`) -115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90 -machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs -of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and -`summary.json` into its own folder under `out/` (untracked — **those are results, -never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no -animation at all. +205 runs in eleven groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9, +J 65, K 25), ≈101 machine-hours — ≈98 GPU-hours over 198 runs plus ≈3 CPU-hours over +the 7 CPU runs of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, +`series.csv` and `summary.json` into its own folder under `out/` (untracked — **those +are results, never clean them**). Only **59 of the 205** pass `--gif`; J and K never do. + +**J** compares wall schemes with the wall function (Bouzidi without it vs Bouzidi, +Grad and HRR with Spalding) on cylinders and NACA 0012 over Re and resolution; **K** +compares outer-domain layouts built with `--levels` (buffer sizes, ×4/×8 jump vs ×2 +steps, fine-wake length, lateral sponges) against a uniform fine reference. Both +measure against a reference *inside* the group; `bench/compare.py` prints the tables. +J's Re/D were chosen against the **measured stability limit τ − ½ ≈ 6·10⁻⁴** (long +runs: Re 10⁴ dies at D 16 and 32, survives at D 64; Re 5000 at D 32 survives). The +older README claim "Re ≈ 5·10⁴ at D = 16" came from short runs and is wrong. ```sh cd bench @@ -637,7 +675,7 @@ Note the campaign has barely started: `out/` currently holds four finished runs ## Deployment -Published image **`notbigghost/kbc2d:1.3.0`** (also `latest`); the server needs only +Published image **`notbigghost/kbc2d:1.4.0`** (also `latest`); the server needs only `docker-compose.server.yml` (the only compose file with no `build:` section), not the sources. `ENTRYPOINT` is `bench/entrypoint.sh` and `CMD` defaults to `--dry-run`: without `KBC2D_SHARD` it just forwards to `run_campaign.sh`, so a stray `docker run` @@ -645,7 +683,7 @@ prints an estimate instead of starting a 100-hour job. **Five shards** (since 1.3.0): every run in `scenarios.json` carries `"shard": 1..5`, assigned by LPT in `gen_scenarios.py` (`--shards N`) on a dzn-shaped time model -(`DZN_MLUPS`), ≈24 h each. `docker-compose.shards.yml` + `shard.bat N` target someone +(`DZN_MLUPS`), ≈27 h each. `docker-compose.shards.yml` + `shard.bat N` target someone else's Windows/WSL2 PC: with `KBC2D_SHARD` set, the entrypoint picks the GPU path (`/dev/dxg` → dzn, else NVIDIA toolkit), requires a non-CPU adapter in `vulkaninfo`, runs `parity.py` once (marker `out/parity_shardN.ok`), then @@ -658,7 +696,7 @@ the card first (`--profile check run --rm vulkan`) — a missing GPU is better discovered in a minute than in an hour. The `check` profile also carries `preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file). -Two version caveats: the `1.3.0` tag lives only in the compose files and prose — +Two version caveats: the `1.4.0` tag lives only in the compose files and prose — `Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults `ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose file sets `platform:` and the Dockerfile sets no `--platform`.