CLAUDE.md: заметки Claude Code по всему проекту
Ветка-сирота: только этот файл, по коммиту на редакцию. Общего предка с ветками кода нет, поэтому она никуда не сливается и в диффах не появляется. Рабочие ветки держат файл в .gitignore, так что одна копия переживает переключения; здесь эталон, резервная копия и история правок. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B9Gcr11JJJyf8NnWsXjDzZ
This commit is contained in:
@@ -0,0 +1,494 @@
|
|||||||
|
# CLAUDE.md
|
||||||
|
|
||||||
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||||
|
|
||||||
|
## Read this first: the file is shared across branches
|
||||||
|
|
||||||
|
`CLAUDE.md` is **deliberately untracked on every working branch** — `/CLAUDE.md`
|
||||||
|
sits in `.gitignore` on `main`, `dev`, `research` and `rust`. One physical file
|
||||||
|
therefore survives every `git checkout`, so these notes stay put while the tree
|
||||||
|
around them changes, and editing them never shows up in `git status`.
|
||||||
|
|
||||||
|
The master copy lives on the **orphan branch `notes`**, which holds this one file
|
||||||
|
and nothing else, one commit per revision — that branch is both the backup and the
|
||||||
|
change history. It shares no ancestor with any code branch, so it never merges into
|
||||||
|
anything and never appears in a diff.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git show notes:CLAUDE.md > CLAUDE.md # restore it in a fresh clone
|
||||||
|
git log --oneline notes # history of these notes
|
||||||
|
git diff notes~1 notes # what changed in the last revision
|
||||||
|
```
|
||||||
|
|
||||||
|
After editing this file, publish the new revision — this touches neither the
|
||||||
|
working tree nor the checked-out branch:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
blob=$(git hash-object -w CLAUDE.md)
|
||||||
|
tree=$(printf '100644 blob %s\tCLAUDE.md\n' "$blob" | git mktree)
|
||||||
|
git branch -f notes "$(git commit-tree "$tree" -p notes -m 'CLAUDE.md: what changed')"
|
||||||
|
git push CFDManager notes
|
||||||
|
```
|
||||||
|
|
||||||
|
`git clean -xfd` deletes the working copy, since it is untracked — restore it with
|
||||||
|
the first command above.
|
||||||
|
|
||||||
|
The consequence to keep in mind: **this file describes the whole project, but no
|
||||||
|
single branch contains all of it.** Before assuming a directory exists, check
|
||||||
|
`git branch --show-current` against the map below — `rust/` and
|
||||||
|
`docs/theory/2d_solver/` each live on one branch only.
|
||||||
|
|
||||||
|
## Branch map
|
||||||
|
|
||||||
|
| branch | what it adds | trees present |
|
||||||
|
|---|---|---|
|
||||||
|
| `main`, `dev` | baseline, nothing beyond the initial import (both at the same commit) | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) |
|
||||||
|
| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images. Has unpushed commits. | `+ docs/theory/2d_solver/` |
|
||||||
|
| `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` |
|
||||||
|
|
||||||
|
Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD
|
||||||
|
solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are
|
||||||
|
unrelated projects that merely both happen to be Rust. Neither is part of the CMake
|
||||||
|
build.
|
||||||
|
|
||||||
|
Untracked leftovers you will see in the working tree regardless of branch:
|
||||||
|
`build/`, `rust/target/`, `docs/theory/2d_solver/out/` (campaign results —
|
||||||
|
**the user's data, do not clean**), `pipeline_cache.bin`, `imgui.ini`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# SimVulcan — C++/Vulkan editor (all branches)
|
||||||
|
|
||||||
|
A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY,
|
||||||
|
XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj`
|
||||||
|
loading with selectable display modes (solid, wireframe, solid + wireframe). It
|
||||||
|
runs **no simulation** — it is a rendering skeleton with an ImGui interface (a
|
||||||
|
Viewport control panel and a Mesh load panel). Unchanged since the initial commit
|
||||||
|
on every branch.
|
||||||
|
|
||||||
|
## Build
|
||||||
|
|
||||||
|
Prerequisites: **Vulkan SDK 1.3.290+** (provides `glslangValidator` for offline
|
||||||
|
shader compilation), **CMake 3.26+**, **Ninja**, a C++20 compiler (MSVC 19.36+,
|
||||||
|
gcc 11+, clang 14+). On macOS, Vulkan is via MoltenVK (Apple Silicon only).
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cmake --preset windows-msvc-release
|
||||||
|
cmake --build --preset windows-msvc-release
|
||||||
|
```
|
||||||
|
|
||||||
|
Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`,
|
||||||
|
`linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under
|
||||||
|
`build/<presetName>/`). The first configure fetches dependencies via FetchContent
|
||||||
|
(GLFW, GLM, volk, vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog,
|
||||||
|
tinyobjloader, Catch2) and needs network access — it clones nine repositories with
|
||||||
|
full history (**775 MB**, ~15 min on a slow link) and **has no shared cache**, so
|
||||||
|
every new build directory downloads all of it again. `VK_NO_PROTOTYPES` is set
|
||||||
|
project-wide; Vulkan entry points load through **volk**.
|
||||||
|
|
||||||
|
Warnings come from `simv_set_warnings` (`/W4 /permissive-`, or
|
||||||
|
`-Wall -Wextra -Wpedantic -Wshadow -Wold-style-cast …`); they are **not** errors.
|
||||||
|
`cmake/Sanitizers.cmake` defines `simv_enable_sanitizers` (Debug-only ASan/UBSan)
|
||||||
|
but no target currently calls it — wire it in manually when chasing memory bugs.
|
||||||
|
|
||||||
|
## Running
|
||||||
|
|
||||||
|
The executable resolves SPIR-V relative to the working directory (`FindSpvPath`
|
||||||
|
probes `spirv/<rel>` and `current_path()/spirv/<rel>`). The `SimVulcan` POST_BUILD
|
||||||
|
step copies the compiled `spirv/` tree and `assets/` next to the executable, so
|
||||||
|
**run from the executable's own directory** (`build/<preset>/src/app/`). `main.cpp`
|
||||||
|
also probes several `../` ancestors for `assets/meshes`. Meshes load at runtime
|
||||||
|
through the ImGui Mesh panel.
|
||||||
|
|
||||||
|
Running writes two files into the CWD: `pipeline_cache.bin` (serialised
|
||||||
|
`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini`. Both are
|
||||||
|
disposable — delete them if pipeline creation or the panel layout misbehaves.
|
||||||
|
|
||||||
|
`ContextOptions::enableValidation` / `enableDebugUtils` default to **true** in
|
||||||
|
every build config, so the Khronos validation layer is requested even in Release;
|
||||||
|
messages (error + warning severity) go through spdlog.
|
||||||
|
|
||||||
|
`run.bat` at the repo root is **stale**: it launches
|
||||||
|
`build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never
|
||||||
|
produce. Do not point users at it without fixing the path first.
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
|
||||||
|
Catch2 unit tests, pure CPU/math (mesh bounds + welding). No GPU required.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cmake --build --preset windows-msvc-debug --target simv_tests
|
||||||
|
ctest --preset windows-msvc-debug
|
||||||
|
```
|
||||||
|
|
||||||
|
`windows-msvc-debug` is the only preset with a `testPreset` — for the other
|
||||||
|
configs, invoke `ctest` in `build/<preset>/` directly.
|
||||||
|
|
||||||
|
Single test / subset — either through CTest (each `TEST_CASE` is registered
|
||||||
|
individually by `catch_discover_tests`):
|
||||||
|
|
||||||
|
```sh
|
||||||
|
ctest --preset windows-msvc-debug -R "WeldVertices" -V
|
||||||
|
```
|
||||||
|
|
||||||
|
or by running the binary with a Catch2 name or tag filter:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]"
|
||||||
|
./build/windows-msvc-debug/tests/simv_tests.exe --list-tests
|
||||||
|
```
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
### Library / target map
|
||||||
|
|
||||||
|
```
|
||||||
|
SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor
|
||||||
|
simv_core → simv_vk (App owns Window + vk::Renderer; no Vulkan calls)
|
||||||
|
simv_editor → simv_core, simv_mesh (Camera, input, ImGui panels; no Vulkan)
|
||||||
|
simv_mesh → simv_core (CPU mesh only; no Vulkan)
|
||||||
|
simv_vk → third-party (volk, vk-bootstrap, VMA, GLFW, GLM, spdlog, imgui)
|
||||||
|
simv_shaders → glslangValidator (GLSL → SPIR-V)
|
||||||
|
```
|
||||||
|
|
||||||
|
Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`,
|
||||||
|
`simv::editor`. Each library exports `src/` as its include root, so includes are
|
||||||
|
written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
|
||||||
|
|
||||||
|
### Frame loop
|
||||||
|
|
||||||
|
`main.cpp` creates `core::App`, which owns the `core::Window` and a
|
||||||
|
`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` calls the UI
|
||||||
|
callback (between ImGui NewFrame/Render), then records the scene: grid + mesh into
|
||||||
|
one dynamic-rendering pass with a depth attachment, followed by ImGui, then
|
||||||
|
presents (2 frames in flight, sync2 submits).
|
||||||
|
|
||||||
|
`main.cpp` wires the editor via the UI callback: it draws the panels, applies
|
||||||
|
mouse input to the `editor::Camera`, and pushes the resulting state into the
|
||||||
|
renderer (`SetViewProj`, `SetRenderMode`, `SetGridVisible`). Mesh loads go through
|
||||||
|
`MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`.
|
||||||
|
|
||||||
|
### Mesh load path
|
||||||
|
|
||||||
|
`MeshLoadPanel` lists `*.obj` in the mesh directory and kicks off
|
||||||
|
`mesh::LoadObjAsync` (worker thread, `std::future`). The future is **drained on
|
||||||
|
the main thread** at the top of `MeshLoadPanel::Draw`, so the `OnLoaded` callback
|
||||||
|
— and therefore the GPU upload — always runs on the render thread. Loader
|
||||||
|
exceptions surface as the panel's status string.
|
||||||
|
|
||||||
|
`main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs
|
||||||
|
the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius.
|
||||||
|
Despite the file name, `mesh/MeshDecimator.h` implements **only** spatial-hash
|
||||||
|
vertex welding (collapse near-duplicates, drop degenerate triangles, recompute
|
||||||
|
bounds) — there is no LOD/decimation.
|
||||||
|
|
||||||
|
### Non-obvious invariants (read before editing)
|
||||||
|
|
||||||
|
- **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no
|
||||||
|
Vulkan API calls. `vk::Renderer`'s public header is deliberately Vulkan-free
|
||||||
|
(pImpl + glm/`RenderMode` only) so `App` can own it without pulling in volk. Keep
|
||||||
|
it that way — do not leak `Vk*` types into the public interfaces of those modules.
|
||||||
|
Note this is **convention only**: every header is visible, so nothing stops a
|
||||||
|
`#include <volk.h>` in `mesh/`; the compiler will not catch it.
|
||||||
|
- **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one
|
||||||
|
`vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth
|
||||||
|
attachment. The ImGui backend is initialised with `depthAttachmentFormat` set so
|
||||||
|
its pipeline matches the pass; the depth buffer is recreated with the swapchain.
|
||||||
|
- **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and
|
||||||
|
a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small depth bias so
|
||||||
|
the overlay sits on top of the fill. The device feature is requested in `Context`.
|
||||||
|
Culling is off (`VK_CULL_MODE_NONE`) — loaded models may have mixed winding.
|
||||||
|
- **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance)
|
||||||
|
the world origin; loaded models are recentred there via a translate-only model
|
||||||
|
matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip,
|
||||||
|
isolated to `Camera.cpp`).
|
||||||
|
- **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on
|
||||||
|
the transfer queue). `GridRenderer` builds a static host-visible line buffer once.
|
||||||
|
Both scene renderers use only a push constant — no descriptor sets.
|
||||||
|
- **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous
|
||||||
|
`MeshPC { mat4 mvp; vec4 color; }` must stay byte-identical to the
|
||||||
|
`push_constant` block in `mesh.vert`/`mesh.frag`; `color.a` is a *flag*, not
|
||||||
|
alpha (1 = flat-shaded, 0 = constant colour for wireframe). `GridRenderer` pushes
|
||||||
|
a bare `mat4` (vertex stage only). Change either side and you must change both.
|
||||||
|
- **Swapchain recreation rebuilds sync objects.** `RecreateSwapDependent` recreates
|
||||||
|
the per-frame `imageAvailable` semaphores (a failed acquire can leave one
|
||||||
|
signalled) *and* the per-swapchain-image `renderFinished` semaphores (the image
|
||||||
|
count may change), then re-ensures both pipelines. Keep that ordering if you
|
||||||
|
touch resize handling.
|
||||||
|
- **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle`
|
||||||
|
before touching `GpuMesh` — acceptable because loads are rare; do not copy that
|
||||||
|
pattern into per-frame paths.
|
||||||
|
|
||||||
|
### Known defects (found by the Rust port, all still present here)
|
||||||
|
|
||||||
|
Porting the editor surfaced four real bugs that nobody was looking for. They are
|
||||||
|
unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md`
|
||||||
|
(branch `rust`).
|
||||||
|
|
||||||
|
1. **The CMake target graph has a cycle.** `core::App` owns `vk::Renderer`, and
|
||||||
|
`vk::Renderer` takes a `core::Window&`. It links only because `simv_vk` does
|
||||||
|
**not** link `simv_core` — it sees the headers through
|
||||||
|
`target_include_directories(simv_vk PUBLIC ..)` and everything resolves at
|
||||||
|
executable link time. CMake never complains.
|
||||||
|
2. **Mesh upload runs mid-frame.** `MeshLoadPanel::Draw` calls `OnLoaded` from
|
||||||
|
inside `Renderer::DrawFrame`, i.e. *after* `vkAcquireNextImageKHR`; the handler
|
||||||
|
calls `SetMeshCpu` → `vkDeviceWaitIdle`. Legal, but it waits for device idle in
|
||||||
|
the middle of command recording.
|
||||||
|
3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and
|
||||||
|
`Renderer`'s separate `hasMesh` flag must be kept consistent by hand.
|
||||||
|
4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category
|
||||||
|
`MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category
|
||||||
|
immediately after; the category throttles at 0.5 s and both are `Info`, so the
|
||||||
|
second message is always dropped. Post-weld counts have therefore never appeared
|
||||||
|
in any log.
|
||||||
|
|
||||||
|
Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** (~200 lines).
|
||||||
|
`GridRenderer`, `GpuMesh` and the depth attachment all call VMA directly.
|
||||||
|
|
||||||
|
### Shaders
|
||||||
|
|
||||||
|
GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders`
|
||||||
|
compiles each to `build/<preset>/spirv/editor/<name>.spv` targeting `vulkan1.3`,
|
||||||
|
with `shaders/` as the `-I` root (so `#include "common/foo.glsl"` would resolve).
|
||||||
|
Adding a shader under that globbed dir is picked up automatically
|
||||||
|
(`CONFIGURE_DEPENDS`). `mesh.frag` reconstructs a flat normal from screen-space
|
||||||
|
derivatives, so the vertex stream carries only positions (tight `float3`, one
|
||||||
|
binding, one attribute). **The Rust port on branch `rust` compiles this same
|
||||||
|
directory** — a shader change affects both versions.
|
||||||
|
|
||||||
|
### Logging
|
||||||
|
|
||||||
|
`simv::core::Logger` (spdlog-backed, singleton) with category-based throttling.
|
||||||
|
Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`,
|
||||||
|
`Vulkan`, `MeshIO`, `UI`, `Test`. Per-category level, throttle interval and
|
||||||
|
on/off are settable at runtime (`SetMinLevel` / `SetThrottle` / `SetEnabled`).
|
||||||
|
Some low-level Vulkan code still calls `spdlog::` directly — which is the only
|
||||||
|
reason the GPU name survives startup (it bypasses category throttling; see
|
||||||
|
defect 4 above).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# `rust/` — Rust port of the editor (branch `rust` only)
|
||||||
|
|
||||||
|
A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions
|
||||||
|
sit side by side and build independently. Same Vulkan 1.3, dynamic rendering,
|
||||||
|
synchronization2, and the same `shaders/` and `assets/` directories. Needs **Rust
|
||||||
|
1.82+** (tested on 1.97.1) and the Vulkan SDK for `glslangValidator` only.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cd rust
|
||||||
|
cargo build --release
|
||||||
|
./target/release/SimVulcan # runnable from any directory
|
||||||
|
cargo test # mesh bounds, welding, Cube.obj, camera clip space
|
||||||
|
```
|
||||||
|
|
||||||
|
**Work on the code with plain `cargo build`, not `--release`.** The release profile
|
||||||
|
sets `lto = "thin"`, which makes a one-line edit re-optimise and relink the whole
|
||||||
|
binary: 52 s versus ~5 s in debug.
|
||||||
|
|
||||||
|
Five crates mirror the CMake target map, and the "all Vulkan in one place"
|
||||||
|
invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have
|
||||||
|
neither `ash` nor `simv-vk` among their dependencies:
|
||||||
|
|
||||||
|
```
|
||||||
|
crates/simv-core/ logger, window → log, chrono, winit
|
||||||
|
crates/simv-mesh/ .obj reading, welding, bounds → glam, tobj
|
||||||
|
crates/simv-vk/ all Vulkan + build.rs (GLSL→SPIR-V) → ash, gpu-allocator,
|
||||||
|
egui-ash-renderer
|
||||||
|
crates/simv-editor/ camera, input, panels → egui
|
||||||
|
crates/simv-app/ App and entry point → all of the above
|
||||||
|
```
|
||||||
|
|
||||||
|
Structural differences from C++, each with a reason (full list in
|
||||||
|
`docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the
|
||||||
|
target-graph cycle); the UI closure **returns** a `FrameState` instead of calling
|
||||||
|
setters (it is invoked from a renderer method, so it cannot borrow the renderer);
|
||||||
|
the loaded mesh is uploaded *after* the frame, not from mid-recording; `Buffer` and
|
||||||
|
`Image` are actually used; no pImpl (private module fields give the same
|
||||||
|
isolation); winit owns the event loop, so `ShouldClose`/`PollEvents` are gone.
|
||||||
|
SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime
|
||||||
|
shader lookup and no requirement to run from a particular directory.
|
||||||
|
|
||||||
|
The UI toolkit differs **by design**: egui instead of Dear ImGui, because the
|
||||||
|
`imgui` crate is bindings and would compile ~40k lines of C++, defeating the point
|
||||||
|
of comparing ecosystems.
|
||||||
|
|
||||||
|
## Comparison result (`docs/rust_vs_cpp.md`)
|
||||||
|
|
||||||
|
The document deliberately picks no winner; **no measurement differs by an order of
|
||||||
|
magnitude.** Machine: i5-1135G7 / Iris Xe / Windows 11, both builds release.
|
||||||
|
|
||||||
|
| | C++ | Rust |
|
||||||
|
|---|---|---|
|
||||||
|
| lines of code | 2434 | 2962 (+22 %) |
|
||||||
|
| release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) |
|
||||||
|
| release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) |
|
||||||
|
| debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** |
|
||||||
|
| release exe | **1.18 MB** | 5.80 MB |
|
||||||
|
| dependency sources | 775 MB **per build dir** | **80 MB** shared registry |
|
||||||
|
| startup to renderer ready | 1068 ms | **754 ms** |
|
||||||
|
| CPU at 60 Hz | **11.2 %** | 12.6 % |
|
||||||
|
|
||||||
|
Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap
|
||||||
|
is egui rebuilding its layout every frame, not the language. The +22 % line count
|
||||||
|
is almost entirely the **missing vk-bootstrap replacement**: `Context` + `Swapchain`
|
||||||
|
go from 301 to 606 lines, while the rest of the Vulkan layer matches nearly line
|
||||||
|
for line. The one sharp build cell (52 s) is the price of thin LTO, not of Rust —
|
||||||
|
with LTO off it is 5.5 s against 4.4 s.
|
||||||
|
|
||||||
|
Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj`
|
||||||
|
differs by 0.13 %, entirely due to the reading libraries (tobj drops vertices no
|
||||||
|
face references; fan triangulation of n-gons versus tinyobjloader's earcut).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# `docs/theory/2d_solver/` — kbc2d, Rust LBM solver (branch `research` only)
|
||||||
|
|
||||||
|
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
|
||||||
|
collision operator**, written from scratch in Rust (not a port): flow past bodies in
|
||||||
|
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall
|
||||||
|
models, and GIF output locked to *physical* flow time. Needs **Rust 1.75+**.
|
||||||
|
Physics is anchored to the method authors' papers in `docs/origins/`; formula
|
||||||
|
references in the code follow the 2D paper (arXiv:1507.02509).
|
||||||
|
|
||||||
|
`README.md` there is the authoritative status/validation log (~600 lines: what is
|
||||||
|
verified, against which equation, with what measured numbers) and `bench/README.md`
|
||||||
|
covers the validation campaign. **Read them before changing physics or defaults** —
|
||||||
|
most defaults are the outcome of a documented measurement, not a guess.
|
||||||
|
|
||||||
|
## Build / test / run
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cd docs/theory/2d_solver
|
||||||
|
cargo build --release # with the GPU backend (default feature `gpu`)
|
||||||
|
cargo build --release --no-default-features # CPU only, no wgpu
|
||||||
|
cargo test --release # 29 fast tests
|
||||||
|
cargo test --release -- --include-ignored # + 4 long paper benchmarks (~18 s)
|
||||||
|
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
|
||||||
|
```
|
||||||
|
|
||||||
|
**Never build `--no-default-features` last.** Both builds write the same
|
||||||
|
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
|
||||||
|
`--backend gpu` then refuses to run. Order: no-default-features first, normal
|
||||||
|
build second.
|
||||||
|
|
||||||
|
```sh
|
||||||
|
./target/release/kbc2d --shape cylinder --size 24 --re 150 \
|
||||||
|
--nx 480 --ny 240 --steps 40000 --sponge-len 32 \
|
||||||
|
--gif wake.gif --gif-field vorticity --verbose full
|
||||||
|
```
|
||||||
|
|
||||||
|
`--help` groups every key by role (Physics, Grid, Body, Time, Scheme, Animation,
|
||||||
|
Output). Note `--time <seconds>` as an alternative to `--steps`: a step is not a
|
||||||
|
fixed slice of time (δt = u_lat·δx/u_phys), so refining the cell silently shortens
|
||||||
|
a fixed step count.
|
||||||
|
|
||||||
|
## File roles — one concern each
|
||||||
|
|
||||||
|
| file | owns |
|
||||||
|
|---|---|
|
||||||
|
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, unit conversion. Tests against the papers' formulas live here. `pub type R = f64`. |
|
||||||
|
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. |
|
||||||
|
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
|
||||||
|
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. |
|
||||||
|
| `src/gif.rs` | encoding, palettes, normalisation, and the physical-time frame timing. |
|
||||||
|
|
||||||
|
## Invariants (read before editing)
|
||||||
|
|
||||||
|
- **Physics is written twice, topology once.** `gpu.rs` mirrors `math.rs` by hand
|
||||||
|
in WGSL; masks, Bouzidi links and the patch frame come from `cpu::Geom` /
|
||||||
|
`cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the
|
||||||
|
GPU silently running Bouzidi for `--wall staircase`, caught only because two
|
||||||
|
models produced bit-identical output where they had to differ. Backend parity is
|
||||||
|
a hard requirement (CPU f64 vs GPU f32 agree to 4–5 significant digits; all four
|
||||||
|
wall models agree to 0.007 % on Cd) and `bench/parity.py` checks it.
|
||||||
|
- **The backends have different step contracts.** `Backend` in `main.rs` is an
|
||||||
|
enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` /
|
||||||
|
`flush(out)` push a *batch* — results accumulate in a 128-slot ring and sync once
|
||||||
|
per batch. Adding a per-step GPU readback outside `StepRec` reintroduces a
|
||||||
|
`map_async` + `poll(Wait)` per step, which cost ~8× throughput before batching.
|
||||||
|
- **`GREL = 1e-8` is a relative threshold** — a fraction of ⟨Δh|Δh⟩, not an
|
||||||
|
absolute one. ⟨Δh|Δh⟩ is quadratic in non-equilibrium and physically tiny
|
||||||
|
(~1e-7…1e-9), so an absolute threshold fires almost everywhere and silently
|
||||||
|
substitutes γ = 2, i.e. plain LBGK instead of KBC. The report prints the degenerate-
|
||||||
|
node fraction; on a healthy threshold it must be ~0 (except the very first step).
|
||||||
|
- **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no
|
||||||
|
hardware helps. Convergence studies therefore run on CPU, everything else on GPU;
|
||||||
|
below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude.
|
||||||
|
Reductions and force sums use compensated (Kahan–Neumaier) summation — that is
|
||||||
|
where f32 was losing most of its precision.
|
||||||
|
- **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` /
|
||||||
|
`wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The
|
||||||
|
workgroup → node mapping stays exactly linear, which is what lets the reductions
|
||||||
|
keep working; verified to 4096×4096.
|
||||||
|
- **Population storage has two binding layouts.** One combined binding normally;
|
||||||
|
**nine bindings, one per direction**, when the adapter's max binding size is small
|
||||||
|
(dzn/WSL2 caps a binding at 128 MiB while allowing a 2047 MiB buffer). Chosen from
|
||||||
|
adapter limits, no `switch` in the accessors; both give identical numbers, split
|
||||||
|
costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS=1` to compare on one
|
||||||
|
card.
|
||||||
|
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
|
||||||
|
per-body force is bucketed — bodies beyond the fourth have their forces merged.
|
||||||
|
- **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps
|
||||||
|
a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge
|
||||||
|
cannot remove it, since a standing mode has its pressure node exactly where the
|
||||||
|
sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but higher bulk viscosity,
|
||||||
|
which damps longitudinal acoustics); `--wall hrr` (**provisional** — chosen on
|
||||||
|
scheme structure, pending campaign group C; `bouzidi` resolves geometry 4× better).
|
||||||
|
|
||||||
|
## Validation campaign (`bench/`)
|
||||||
|
|
||||||
|
115 runs, ≈90 GPU-hours in nine groups; each run writes logs, series, a
|
||||||
|
machine-readable summary and a GIF into its own folder under `out/` (untracked —
|
||||||
|
**those are results, never clean them**).
|
||||||
|
|
||||||
|
```sh
|
||||||
|
cd bench
|
||||||
|
python preflight.py # start every scenario for two steps — catches typos
|
||||||
|
./run_campaign.sh --calibrate # measure this machine's MLUPS (hour estimates need it)
|
||||||
|
./run_campaign.sh --dry-run # cost estimate
|
||||||
|
./run_campaign.sh --resume # run, skipping what is already done
|
||||||
|
```
|
||||||
|
|
||||||
|
Run `preflight.py` for real — it caught an entire group failing on a GPU limit and
|
||||||
|
nine runs passing `--body-x` twice, both of which would otherwise have surfaced
|
||||||
|
twenty hours into a server campaign.
|
||||||
|
|
||||||
|
## Deployment
|
||||||
|
|
||||||
|
Published image **`notbigghost/kbc2d:1.2.0`** (`linux/amd64`); the server needs
|
||||||
|
only `docker-compose.server.yml`, not the sources. `ENTRYPOINT` is the campaign
|
||||||
|
driver and `CMD` defaults to `--dry-run`, so a stray `docker run` prints an
|
||||||
|
estimate instead of starting a 100-hour job. Check the card first
|
||||||
|
(`--profile check run --rm vulkan`) — a missing GPU is better discovered in a
|
||||||
|
minute than in an hour.
|
||||||
|
|
||||||
|
Two environment traps, both already handled in the image but overridable from
|
||||||
|
outside: NVIDIA Container Toolkit only injects the Vulkan ICD when
|
||||||
|
`NVIDIA_DRIVER_CAPABILITIES` contains `graphics` (with `compute` alone wgpu sees no
|
||||||
|
adapter), and **WSL2 has no NVIDIA Vulkan driver for Linux at all** — the card
|
||||||
|
arrives over `/dev/dxg`. That path instead uses **dzn** (Mesa's Vulkan→D3D12
|
||||||
|
translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build
|
||||||
|
dzn), needs no NVIDIA runtime, and requires
|
||||||
|
`WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports
|
||||||
|
`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default. Use
|
||||||
|
`docker-compose.wsl.yml` there.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# `docs/theory/` — Python prototype (all branches)
|
||||||
|
|
||||||
|
The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow
|
||||||
|
with a ×2 nested AMR patch and SDF+Bouzidi boundaries, **GPU/CuPy only, no CPU
|
||||||
|
fallback**. `kbc2d` is a deliberate rewrite of this, not a port, and the two differ
|
||||||
|
in two documented places (no collision inside the body; restriction skips fine
|
||||||
|
source nodes inside the body). Still live: it owns the notebook's figures.
|
||||||
|
|
||||||
|
- `solver_2x_sdf/README.md` — maps every file to the physics it owns; entry points
|
||||||
|
`run.py`, `run_blockage.py`, `run_factors.py`.
|
||||||
|
- `demos_gpu/` — demo runs producing the notebook's figures/GIFs; they import
|
||||||
|
physics from `solver_2x_sdf` and never duplicate it.
|
||||||
|
- `kbc_lbm.ipynb` — the write-up; embeds pre-rendered artefacts, executes nothing.
|
||||||
|
- `docs/origins/` — the source PDFs behind both implementations.
|
||||||
|
|
||||||
|
Do not fold any of this into the CMake build, and do not treat it as dead code.
|
||||||
Reference in New Issue
Block a user