CLAUDE.md: kbc2d 1.4.0 — цепочка уровней (--levels), пристеночная функция Spalding, исправление силы на GPU при --refine>=3, группы J и K, измеренный предел устойчивости

This commit is contained in:
2026-10-02 12:25:51 +03:00
parent 3882af3886
commit 5448e4334e
+55 -17
View File
@@ -493,8 +493,9 @@ should be *all three ported, two more added*; and `Renderer::Impl`'s destructor
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
collision operator**, written from scratch in Rust (not a port): flow past bodies in collision operator**, written from scratch in Rust (not a port): flow past bodies in
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall a channel, two interchangeable backends, a chain of nested AMR levels (`--levels`;
models, and GIF output locked to *physical* flow time. `Cargo.toml` declares the old single `--refine`/`--patch` is its one-link case), sub-grid wall models with
an optional Spalding wall function, and GIF output locked to *physical* flow time. `Cargo.toml` declares
**Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the **Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the
method authors' papers in `docs/origins/`; formula references in the code follow the method authors' papers in `docs/origins/`; formula references in the code follow the
2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models 2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models
@@ -512,14 +513,16 @@ defaults** — most defaults are the outcome of a documented measurement, not a
cd docs/theory/2d_solver cd docs/theory/2d_solver
cargo build --release # with the GPU backend (default feature `gpu`) cargo build --release # with the GPU backend (default feature `gpu`)
cargo build --release --no-default-features # CPU only, no wgpu cargo build --release --no-default-features # CPU only, no wgpu
cargo test --release # 29 fast tests cargo test --release # 36 fast tests
cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s) cargo test --release -- --include-ignored # + 5 long CPU runs (~35 s)
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
``` ```
33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the 41 `#[test]` functions exist (23 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`, 3 in
4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two `main.rs`); 5 are `#[ignore]`d — four in `cpu.rs` (two reference benchmarks, two
diagnostics, not four benchmarks. diagnostics) and `levels_match_single_patch_on_steady_cylinder` in `main.rs`. The
`main.rs` tests build a real `Spec` through `Cli::try_parse_from` + `build_spec` — the
only tests that drive a whole `cpu::Sim`.
**Never build `--no-default-features` last.** Both builds write the same **Never build `--no-default-features` last.** Both builds write the same
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and `target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
@@ -551,7 +554,7 @@ tests.
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. | | `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. |
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. | | `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. |
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. | | `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. | | `src/main.rs` | CLI, problem assembly (incl. `parse_levels`), step loop, live output, report, CSV. Owns the `Spec` / `PatchSpec` / `LevelGeo` / `StepRec` / `FieldKind` contract shared by both backends. |
| `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. | | `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. |
## Invariants (read before editing) ## Invariants (read before editing)
@@ -598,6 +601,33 @@ tests.
layout has none, which is the whole reason it is kept. Both give identical numbers; layout has none, which is the whole reason it is kept. Both give identical numbers;
the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to
anything other than `0`. anything other than `0`.
- **Levels are a chain, stepped recursively.** `Spec.patches: Vec<PatchSpec>` (each
box in its *parent's* coordinates) replaced `refine`/`patch`; `Spec::levels()` derives
every level's size, scale, origin and τ (`τ_k = r_k(τ_{k−1} − ½) + ½`). Both
backends step level k, then r_k substeps of k + 1 (frame interpolated in time),
then restrict; force is taken on the finest level at every one of its substeps,
walls/inlet/outlet/sponges exist on L0 only. `--size`, `--nx`, `--ny`, `--steps`
and Re are all in **L0** cells/steps, so L0 is the level closest to τ → ½. The
single-patch path is bitwise identical to the pre-1.4.0 code on both backends
(verified on five setups) — keep it that way when touching the recursion.
- **GPU force slots scale with substeps.** `results` holds `HIST·SLOT` (SLOT = 28:
12 stats + 12 per-body force + ρ at the probe + 3 spare) plus `substeps·MAXB·4`
force work slots. Before 1.4.0 there were slots for exactly two substeps, so with
`--refine ≥ 3` force of later substeps fell off the buffer and GPU Cd came out
r/2 times too low — any GPU result with `--refine 3` from ≤ 1.3.0 is wrong.
- **β is per node**, not per column (`math::beta_field`), so `--sponge-side` can add
lateral sponges; with `--sponge-side 0` the values equal the old column profile.
- **Wall function (`--wall-function spalding`) uses the same (u_node, τ_w) for every
scheme.** Geometry (`cpu::WfNode`: normal from the SDF gradient, y_w, a sample
point one cell further along the normal, its bilinear stencil) is built once and
uploaded. `grad`/`hrr` take u_node into the target velocity and u_τ²/ν into the
pressure tensor; `bouzidi` gets a **stress-matched** wall slip
`u_s = u_node − u_τ²·y_w/ν`, identically zero in the viscous sublayer. Two earlier
slip formulas were wrong and are documented in the README (velocity-matched slip
made the wall nearly free-slip, Cd 0.016; resolved node velocity broke the no-slip
limit by 2 %). The moment-wall and wall-function pipelines (group 3: wall nodes,
`GWf`, slips) are only built when needed and require 11 storage bindings per stage
(27 with split populations). Every channel run reports y⁺ (`yplus_*`, `u_tau_mean`).
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but - **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
per-body force is bucketed — bodies beyond the fourth have their forces merged per-body force is bucketed — bodies beyond the fourth have their forces merged
(`min(body, MAXB-1)` in three places), and the CLI warns when that happens. (`min(body, MAXB-1)` in three places), and the CLI warns when that happens.
@@ -612,12 +642,20 @@ tests.
## Validation campaign (`bench/`) ## Validation campaign (`bench/`)
115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90 205 runs in eleven groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9,
machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs J 65, K 25), ≈101 machine-hours — ≈98 GPU-hours over 198 runs plus ≈3 CPU-hours over
of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and the 7 CPU runs of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`,
`summary.json` into its own folder under `out/` (untracked — **those are results, `series.csv` and `summary.json` into its own folder under `out/` (untracked — **those
never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no are results, never clean them**). Only **59 of the 205** pass `--gif`; J and K never do.
animation at all.
**J** compares wall schemes with the wall function (Bouzidi without it vs Bouzidi,
Grad and HRR with Spalding) on cylinders and NACA 0012 over Re and resolution; **K**
compares outer-domain layouts built with `--levels` (buffer sizes, ×4/×8 jump vs ×2
steps, fine-wake length, lateral sponges) against a uniform fine reference. Both
measure against a reference *inside* the group; `bench/compare.py` prints the tables.
J's Re/D were chosen against the **measured stability limit τ − ½ ≈ 6·10⁻⁴** (long
runs: Re 10⁴ dies at D 16 and 32, survives at D 64; Re 5000 at D 32 survives). The
older README claim "Re ≈ 5·10⁴ at D = 16" came from short runs and is wrong.
```sh ```sh
cd bench cd bench
@@ -637,7 +675,7 @@ Note the campaign has barely started: `out/` currently holds four finished runs
## Deployment ## Deployment
Published image **`notbigghost/kbc2d:1.3.0`** (also `latest`); the server needs only Published image **`notbigghost/kbc2d:1.4.0`** (also `latest`); the server needs only
`docker-compose.server.yml` (the only compose file with no `build:` section), not `docker-compose.server.yml` (the only compose file with no `build:` section), not
the sources. `ENTRYPOINT` is `bench/entrypoint.sh` and `CMD` defaults to `--dry-run`: the sources. `ENTRYPOINT` is `bench/entrypoint.sh` and `CMD` defaults to `--dry-run`:
without `KBC2D_SHARD` it just forwards to `run_campaign.sh`, so a stray `docker run` without `KBC2D_SHARD` it just forwards to `run_campaign.sh`, so a stray `docker run`
@@ -645,7 +683,7 @@ prints an estimate instead of starting a 100-hour job.
**Five shards** (since 1.3.0): every run in `scenarios.json` carries `"shard": 1..5`, **Five shards** (since 1.3.0): every run in `scenarios.json` carries `"shard": 1..5`,
assigned by LPT in `gen_scenarios.py` (`--shards N`) on a dzn-shaped time model assigned by LPT in `gen_scenarios.py` (`--shards N`) on a dzn-shaped time model
(`DZN_MLUPS`), ≈24 h each. `docker-compose.shards.yml` + `shard.bat N` target someone (`DZN_MLUPS`), ≈27 h each. `docker-compose.shards.yml` + `shard.bat N` target someone
else's Windows/WSL2 PC: with `KBC2D_SHARD` set, the entrypoint picks the GPU path else's Windows/WSL2 PC: with `KBC2D_SHARD` set, the entrypoint picks the GPU path
(`/dev/dxg` → dzn, else NVIDIA toolkit), requires a non-CPU adapter in `vulkaninfo`, (`/dev/dxg` → dzn, else NVIDIA toolkit), requires a non-CPU adapter in `vulkaninfo`,
runs `parity.py` once (marker `out/parity_shardN.ok`), then runs `parity.py` once (marker `out/parity_shardN.ok`), then
@@ -658,7 +696,7 @@ the card first (`--profile check run --rm vulkan`) — a missing GPU is better
discovered in a minute than in an hour. The `check` profile also carries discovered in a minute than in an hour. The `check` profile also carries
`preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file). `preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file).
Two version caveats: the `1.3.0` tag lives only in the compose files and prose — Two version caveats: the `1.4.0` tag lives only in the compose files and prose —
`Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults `Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults
`ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose `ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose
file sets `platform:` and the Dockerfile sets no `--platform`. file sets `platform:` and the Dockerfile sets no `--platform`.