CLAUDE.md: kbc2d 1.4.0 — цепочка уровней (--levels), пристеночная функция Spalding, исправление силы на GPU при --refine>=3, группы J и K, измеренный предел устойчивости
This commit is contained in:
@@ -493,8 +493,9 @@ should be *all three ported, two more added*; and `Renderer::Impl`'s destructor
|
||||
|
||||
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
|
||||
collision operator**, written from scratch in Rust (not a port): flow past bodies in
|
||||
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall
|
||||
models, and GIF output locked to *physical* flow time. `Cargo.toml` declares
|
||||
a channel, two interchangeable backends, a chain of nested AMR levels (`--levels`;
|
||||
the old single `--refine`/`--patch` is its one-link case), sub-grid wall models with
|
||||
an optional Spalding wall function, and GIF output locked to *physical* flow time. `Cargo.toml` declares
|
||||
**Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the
|
||||
method authors' papers in `docs/origins/`; formula references in the code follow the
|
||||
2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models
|
||||
@@ -512,14 +513,16 @@ defaults** — most defaults are the outcome of a documented measurement, not a
|
||||
cd docs/theory/2d_solver
|
||||
cargo build --release # with the GPU backend (default feature `gpu`)
|
||||
cargo build --release --no-default-features # CPU only, no wgpu
|
||||
cargo test --release # 29 fast tests
|
||||
cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s)
|
||||
cargo test --release # 36 fast tests
|
||||
cargo test --release -- --include-ignored # + 5 long CPU runs (~35 s)
|
||||
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
|
||||
```
|
||||
|
||||
33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the
|
||||
4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two
|
||||
diagnostics, not four benchmarks.
|
||||
41 `#[test]` functions exist (23 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`, 3 in
|
||||
`main.rs`); 5 are `#[ignore]`d — four in `cpu.rs` (two reference benchmarks, two
|
||||
diagnostics) and `levels_match_single_patch_on_steady_cylinder` in `main.rs`. The
|
||||
`main.rs` tests build a real `Spec` through `Cli::try_parse_from` + `build_spec` — the
|
||||
only tests that drive a whole `cpu::Sim`.
|
||||
|
||||
**Never build `--no-default-features` last.** Both builds write the same
|
||||
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
|
||||
@@ -551,7 +554,7 @@ tests.
|
||||
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. |
|
||||
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. |
|
||||
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
|
||||
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. |
|
||||
| `src/main.rs` | CLI, problem assembly (incl. `parse_levels`), step loop, live output, report, CSV. Owns the `Spec` / `PatchSpec` / `LevelGeo` / `StepRec` / `FieldKind` contract shared by both backends. |
|
||||
| `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. |
|
||||
|
||||
## Invariants (read before editing)
|
||||
@@ -598,6 +601,33 @@ tests.
|
||||
layout has none, which is the whole reason it is kept. Both give identical numbers;
|
||||
the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to
|
||||
anything other than `0`.
|
||||
- **Levels are a chain, stepped recursively.** `Spec.patches: Vec<PatchSpec>` (each
|
||||
box in its *parent's* coordinates) replaced `refine`/`patch`; `Spec::levels()` derives
|
||||
every level's size, scale, origin and τ (`τ_k = r_k(τ_{k−1} − ½) + ½`). Both
|
||||
backends step level k, then r_k substeps of k + 1 (frame interpolated in time),
|
||||
then restrict; force is taken on the finest level at every one of its substeps,
|
||||
walls/inlet/outlet/sponges exist on L0 only. `--size`, `--nx`, `--ny`, `--steps`
|
||||
and Re are all in **L0** cells/steps, so L0 is the level closest to τ → ½. The
|
||||
single-patch path is bitwise identical to the pre-1.4.0 code on both backends
|
||||
(verified on five setups) — keep it that way when touching the recursion.
|
||||
- **GPU force slots scale with substeps.** `results` holds `HIST·SLOT` (SLOT = 28:
|
||||
12 stats + 12 per-body force + ρ at the probe + 3 spare) plus `substeps·MAXB·4`
|
||||
force work slots. Before 1.4.0 there were slots for exactly two substeps, so with
|
||||
`--refine ≥ 3` force of later substeps fell off the buffer and GPU Cd came out
|
||||
r/2 times too low — any GPU result with `--refine 3` from ≤ 1.3.0 is wrong.
|
||||
- **β is per node**, not per column (`math::beta_field`), so `--sponge-side` can add
|
||||
lateral sponges; with `--sponge-side 0` the values equal the old column profile.
|
||||
- **Wall function (`--wall-function spalding`) uses the same (u_node, τ_w) for every
|
||||
scheme.** Geometry (`cpu::WfNode`: normal from the SDF gradient, y_w, a sample
|
||||
point one cell further along the normal, its bilinear stencil) is built once and
|
||||
uploaded. `grad`/`hrr` take u_node into the target velocity and u_τ²/ν into the
|
||||
pressure tensor; `bouzidi` gets a **stress-matched** wall slip
|
||||
`u_s = u_node − u_τ²·y_w/ν`, identically zero in the viscous sublayer. Two earlier
|
||||
slip formulas were wrong and are documented in the README (velocity-matched slip
|
||||
made the wall nearly free-slip, Cd 0.016; resolved node velocity broke the no-slip
|
||||
limit by 2 %). The moment-wall and wall-function pipelines (group 3: wall nodes,
|
||||
`GWf`, slips) are only built when needed and require 11 storage bindings per stage
|
||||
(27 with split populations). Every channel run reports y⁺ (`yplus_*`, `u_tau_mean`).
|
||||
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
|
||||
per-body force is bucketed — bodies beyond the fourth have their forces merged
|
||||
(`min(body, MAXB-1)` in three places), and the CLI warns when that happens.
|
||||
@@ -612,12 +642,20 @@ tests.
|
||||
|
||||
## Validation campaign (`bench/`)
|
||||
|
||||
115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90
|
||||
machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs
|
||||
of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and
|
||||
`summary.json` into its own folder under `out/` (untracked — **those are results,
|
||||
never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no
|
||||
animation at all.
|
||||
205 runs in eleven groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9,
|
||||
J 65, K 25), ≈101 machine-hours — ≈98 GPU-hours over 198 runs plus ≈3 CPU-hours over
|
||||
the 7 CPU runs of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`,
|
||||
`series.csv` and `summary.json` into its own folder under `out/` (untracked — **those
|
||||
are results, never clean them**). Only **59 of the 205** pass `--gif`; J and K never do.
|
||||
|
||||
**J** compares wall schemes with the wall function (Bouzidi without it vs Bouzidi,
|
||||
Grad and HRR with Spalding) on cylinders and NACA 0012 over Re and resolution; **K**
|
||||
compares outer-domain layouts built with `--levels` (buffer sizes, ×4/×8 jump vs ×2
|
||||
steps, fine-wake length, lateral sponges) against a uniform fine reference. Both
|
||||
measure against a reference *inside* the group; `bench/compare.py` prints the tables.
|
||||
J's Re/D were chosen against the **measured stability limit τ − ½ ≈ 6·10⁻⁴** (long
|
||||
runs: Re 10⁴ dies at D 16 and 32, survives at D 64; Re 5000 at D 32 survives). The
|
||||
older README claim "Re ≈ 5·10⁴ at D = 16" came from short runs and is wrong.
|
||||
|
||||
```sh
|
||||
cd bench
|
||||
@@ -637,7 +675,7 @@ Note the campaign has barely started: `out/` currently holds four finished runs
|
||||
|
||||
## Deployment
|
||||
|
||||
Published image **`notbigghost/kbc2d:1.3.0`** (also `latest`); the server needs only
|
||||
Published image **`notbigghost/kbc2d:1.4.0`** (also `latest`); the server needs only
|
||||
`docker-compose.server.yml` (the only compose file with no `build:` section), not
|
||||
the sources. `ENTRYPOINT` is `bench/entrypoint.sh` and `CMD` defaults to `--dry-run`:
|
||||
without `KBC2D_SHARD` it just forwards to `run_campaign.sh`, so a stray `docker run`
|
||||
@@ -645,7 +683,7 @@ prints an estimate instead of starting a 100-hour job.
|
||||
|
||||
**Five shards** (since 1.3.0): every run in `scenarios.json` carries `"shard": 1..5`,
|
||||
assigned by LPT in `gen_scenarios.py` (`--shards N`) on a dzn-shaped time model
|
||||
(`DZN_MLUPS`), ≈24 h each. `docker-compose.shards.yml` + `shard.bat N` target someone
|
||||
(`DZN_MLUPS`), ≈27 h each. `docker-compose.shards.yml` + `shard.bat N` target someone
|
||||
else's Windows/WSL2 PC: with `KBC2D_SHARD` set, the entrypoint picks the GPU path
|
||||
(`/dev/dxg` → dzn, else NVIDIA toolkit), requires a non-CPU adapter in `vulkaninfo`,
|
||||
runs `parity.py` once (marker `out/parity_shardN.ok`), then
|
||||
@@ -658,7 +696,7 @@ the card first (`--profile check run --rm vulkan`) — a missing GPU is better
|
||||
discovered in a minute than in an hour. The `check` profile also carries
|
||||
`preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file).
|
||||
|
||||
Two version caveats: the `1.3.0` tag lives only in the compose files and prose —
|
||||
Two version caveats: the `1.4.0` tag lives only in the compose files and prose —
|
||||
`Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults
|
||||
`ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose
|
||||
file sets `platform:` and the Dockerfile sets no `--platform`.
|
||||
|
||||
Reference in New Issue
Block a user