CLAUDE.md: сверка всех разделов с кодом, исправлены расхождения

Сплошная проверка файла против дерева проекта: каждое фактическое
утверждение сверено с исходниками, CMake, манифестами Cargo и git.

Ветки и заголовок:
- research не имеет неотправленных коммитов (все пять веток совпадают
  с CFDManager), main/dev стоят на коммит дальше начального импорта;
- research добавляет ещё и docs/origins/Grad's_aproximation.pdf;
- pipeline_cache.bin и imgui.ini лежат не в корне, а рядом с
  исполняемым файлом; добавлены прочие незаметные остатки;
- сказано, что вне research в docs/theory/2d_solver/ остаются только
  артефакты, и как читать решатель через git без переключения ветки;
- build/ в рабочем дереве собран из чужого каталога и содержит 262 МБ
  nlohmann_json, который проект не объявляет.

C++-редактор:
- панель называется Mesh, а не Mesh load;
- версия Vulkan SDK и версии компиляторов нигде не проверяются;
- 775 МБ зависимостей — измерение по чужому build/, на деле ~400 МБ;
- VK_NO_PROTOTYPES задан на двух целях, а не на проекте целиком;
- ASan без UBSan на MSVC; уточнён список предупреждений;
- карта целей дополнена опущенными библиотеками и видимостью связей;
- RecreateSwapDependent пересоздаёт ещё и буфер глубины, первым;
- дефект 1: цикл существует на уровне исходников, граф CMake ацикличен —
  именно поэтому CMake и молчит;
- дефект 2: ожидание простоя происходит до vkBeginCommandBuffer,
  а не посреди записи команд;
- мёртвый код Buffer/Image — 362 строки, а не ~200;
- glob шейдеров ловит только *.vert и *.frag, суффикс стадии остаётся;
- из категорий журнала используется лишь MeshIO, рантайм-настройки
  не вызываются ниоткуда, весь vk/ пишет через сырой spdlog.

Rust-порт:
- объявленный минимум 1.82 недостижим: залоченный egui 0.36.1 требует
  1.95, на 1.82 обрывается резолвинг;
- перечни зависимостей крейтов приведены полностью;
- строк кода 2976, а не 2962 (устарела строка mesh на 17 строк);
- утверждение «ни одно измерение не отличается на порядок» неверно:
  52.3/4.4 = 11.9x;
- рост +22 % на 56 % объясняется заменой vk-bootstrap, а не «почти весь»;
- assets/ не копируется, откат на текущий каталог сохранён;
- перечислены ошибки самого docs/rust_vs_cpp.md.

kbc2d:
- GREL — доля от полной неравновесной нормы, den > GREL*nrm;
- 0.007 % по Cd — это сходимость бэкендов для каждой модели стенки,
  а не согласие четырёх моделей между собой;
- компенсированное суммирование только в статистике: ядро сил считает
  обычным f32;
- раскладок привязок две и восемнадцать, switch — признак расщеплённой;
- заголовки --help русские;
- гифку пишут 59 прогонов из 115; ≈87 GPU-часов плюс ≈3 CPU-часа;
- тег образа 1.2.0 не отражён ни в Cargo.toml, ни в Dockerfile.

Python-прототип:
- в каталоге четыре поколения решателя, а не одно; добавлена таблица,
  какое живое, какое опорное, какое мёртвое;
- «только CuPy без отката» верно не для всего дерева;
- ноутбук состоит из 20 markdown-ячеек без единой ячейки кода, ссылки
  относительные, четыре из них битые;
- отличий от kbc2d три, а не два, и второе сформулировано наоборот;
- README solver_2x_sdf описывает несуществующий переключатель
  collision и сам объявляет свои числа устаревшими.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-16 14:01:46 +03:00
co-authored by Claude Opus 5
parent 6902b887d7
commit 137876193d
+427 -199
View File
@@ -49,18 +49,40 @@ single branch contains all of it.** Before assuming a directory exists, check
| branch | what it adds | trees present | | branch | what it adds | trees present |
|---|---|---| |---|---|---|
| `main`, `dev` | baseline, nothing beyond the initial import (both at the same commit) | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) | | `main`, `dev` | baseline; both at the same commit, one past the initial import — and that commit only untracks `CLAUDE.md` | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) |
| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images. Has unpushed commits. | `+ docs/theory/2d_solver/` | | `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images | `+ docs/theory/2d_solver/`, `+ docs/origins/Grad's_aproximation.pdf` |
| `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` | | `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` |
All five branches (including `notes`) are fully pushed to the `CFDManager` remote;
verify with `git ls-remote CFDManager` rather than trusting a stale tracking ref.
Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD
solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are
unrelated projects that merely both happen to be Rust. Neither is part of the CMake unrelated projects that merely both happen to be Rust. Neither is part of the CMake
build. build — `CMakeLists.txt` never mentions either tree.
Untracked leftovers you will see in the working tree regardless of branch: **A checkout away from `research` does not remove `docs/theory/2d_solver/`** — the
`build/`, `rust/target/`, `docs/theory/2d_solver/out/` (campaign results — tracked sources vanish but the untracked artefacts stay, so on `rust` that directory
**the user's data, do not clean**), `pipeline_cache.bin`, `imgui.ini`. contains only `out/`, `target/`, `wake.gif` and `demo_wake.gif` and looks like a
mutilated copy of the solver. To read the solver from another branch, go through
git rather than the filesystem:
```sh
git show research:docs/theory/2d_solver/src/math.rs
git archive research docs/theory/2d_solver | tar -x -C /tmp/kbc2d # to grep it
```
Untracked leftovers you will see regardless of branch: `build/`, `rust/target/`,
`docs/theory/2d_solver/target/`, the `__pycache__` dirs under `docs/theory/`, and
`docs/theory/2d_solver/out/` plus the two ~9 MB GIFs beside it (campaign results —
**the user's data, do not clean**). `pipeline_cache.bin` and `imgui.ini` are written
next to whichever executable ran, i.e. in `build/<preset>/src/app/` and
`rust/pipeline_cache.bin` — not at the repo root.
**`build/` in this working tree is foreign and stale.** Its
`CMakeCache.txt` records `CMAKE_HOME_DIRECTORY=C:/Users/ivan/UAS/SimV4(Not_KBC)/SimVulcan`,
and `_deps/` carries a 262 MB `nlohmann_json` clone that this project never declares.
Do not measure anything from it without re-configuring first.
--- ---
@@ -69,15 +91,25 @@ Untracked leftovers you will see in the working tree regardless of branch:
A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY, A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY,
XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj` XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj`
loading with selectable display modes (solid, wireframe, solid + wireframe). It loading with selectable display modes (solid, wireframe, solid + wireframe). It
runs **no simulation** — it is a rendering skeleton with an ImGui interface (a runs **no simulation** — no compute pipeline or `.comp` shader exists anywhere, and
Viewport control panel and a Mesh load panel). Unchanged since the initial commit `Renderer.cpp` disables async compute explicitly. It is a rendering skeleton with
on every branch. two ImGui windows, titled **`Viewport`** and **`Mesh`** (the class behind the second
is `MeshLoadPanel`, which is why the docs keep calling it the "mesh load panel").
Unchanged since the initial commit on every branch — `git log --all -- src` returns
exactly one commit.
## Build ## Build
Prerequisites: **Vulkan SDK 1.3.290+** (provides `glslangValidator` for offline Prerequisites: **Vulkan SDK** (for `glslangValidator`), **CMake 3.26+**, **Ninja**,
shader compilation), **CMake 3.26+**, **Ninja**, a C++20 compiler (MSVC 19.36+, a C++20 compiler. Only CMake 3.26 and C++20 are actually enforced
gcc 11+, clang 14+). On macOS, Vulkan is via MoltenVK (Apple Silicon only). (`CMakeLists.txt:1`, `:9-11`). `cmake/VulkanSetup.cmake:3` is
`find_package(Vulkan REQUIRED COMPONENTS glslangValidator)` with **no version
argument**, so the "SDK 1.3.290+" in `README.md` is a comment, not a check; the
de-facto floor comes from the pinned volk `1.3.295` and vk-bootstrap `v1.3.295`.
The compiler minimums (MSVC 19.36+, gcc 11+, clang 14+) are likewise README prose,
enforced nowhere. On macOS the only hard constraint is
`CMAKE_OSX_ARCHITECTURES: arm64` in the preset; MoltenVK is mentioned in a
`message(STATUS)`.
```sh ```sh
cmake --preset windows-msvc-release cmake --preset windows-msvc-release
@@ -86,42 +118,67 @@ cmake --build --preset windows-msvc-release
Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`, Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`,
`linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under `linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under
`build/<presetName>/`). The first configure fetches dependencies via FetchContent `build/<presetName>/`). None of them pins a compiler — they inherit whatever the
(GLFW, GLM, volk, vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog, environment provides, so `README.md`'s "MSVC / clang-cl" is aspirational.
tinyobjloader, Catch2) and needs network access — it clones nine repositories with
full history (**775 MB**, ~15 min on a slow link) and **has no shared cache**, so
every new build directory downloads all of it again. `VK_NO_PROTOTYPES` is set
project-wide; Vulkan entry points load through **volk**.
Warnings come from `simv_set_warnings` (`/W4 /permissive-`, or The first configure fetches nine repositories via FetchContent (GLFW, GLM, volk,
`-Wall -Wextra -Wpedantic -Wshadow -Wold-style-cast …`); they are **not** errors. vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog, tinyobjloader, Catch2) and
`cmake/Sanitizers.cmake` defines `simv_enable_sanitizers` (Debug-only ASan/UBSan) needs network access. There is no `GIT_SHALLOW`, so each is a clone with full
but no target currently calls it — wire it in manually when chasing memory bugs. history, and **no shared cache** is configured (the only `FETCHCONTENT_*` variable
set is `FETCHCONTENT_QUIET OFF`) — every new build directory downloads all of it
again. Budget **~400 MB of sources / ~510 MB of `_deps` once built**, not the
775 MB quoted in `docs/rust_vs_cpp.md`: that figure was measured on the foreign
`build/` tree described above and includes 262 MB of `nlohmann_json` that nothing
here declares.
`VK_NO_PROTOTYPES` is **not** project-wide — there is no `add_compile_definitions`
at project scope. It is set `PUBLIC` on exactly two targets, `imgui`
(`CMakeLists.txt:114`) and `simv_vk` (`src/vk/CMakeLists.txt:29`). Because
`simv_core` links `simv_vk` **PRIVATE**, the definition never reaches `simv_mesh`,
`simv_editor` or `simv_tests`. Vulkan entry points load through **volk**.
Warnings come from `simv_set_warnings` (`cmake/CompilerWarnings.cmake`), called by
all six targets: `/W4 /permissive- /Zc:preprocessor /Zc:__cplusplus /wd4100` plus
the defines `_CRT_SECURE_NO_WARNINGS`, `NOMINMAX`, `WIN32_LEAN_AND_MEAN` on MSVC;
`-Wall -Wextra -Wpedantic -Wno-unused-parameter -Wshadow -Wnon-virtual-dtor
-Wold-style-cast -Wcast-align -Wunused -Woverloaded-virtual` elsewhere. They are
**not** errors — no `/WX` or `-Werror` anywhere. `cmake/Sanitizers.cmake` defines
`simv_enable_sanitizers` (Debug-only; ASan **only** on MSVC, ASan+UBSan elsewhere)
but no target calls it — wire it in manually when chasing memory bugs.
## Running ## Running
The executable resolves SPIR-V relative to the working directory (`FindSpvPath` The executable resolves SPIR-V relative to the working directory (`FindSpvPath`,
probes `spirv/<rel>` and `current_path()/spirv/<rel>`). The `SimVulcan` POST_BUILD `src/vk/Shader.cpp:24-33`, probes `spirv/<rel>` then `current_path()/spirv/<rel>` —
step copies the compiled `spirv/` tree and `assets/` next to the executable, so the second probe is a no-op, since the first is already CWD-relative). Two
**run from the executable's own directory** (`build/<preset>/src/app/`). `main.cpp` `SimVulcan` POST_BUILD steps copy `assets/` and the compiled `spirv/` tree next to
also probes several `../` ancestors for `assets/meshes`. Meshes load at runtime the executable, so **run from the executable's own directory**
through the ImGui Mesh panel. (`build/<preset>/src/app/`). `main.cpp:40-50` additionally probes four `../`
ancestors for `assets/meshes`. Meshes load at runtime through the `Mesh` panel.
Running writes two files into the CWD: `pipeline_cache.bin` (serialised Running writes two files into the CWD: `pipeline_cache.bin` (serialised
`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini`. Both are `VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini` (no
disposable — delete them if pipeline creation or the panel layout misbehaves. `IniFilename` is ever set, so ImGui's default applies). Both are disposable —
delete them if pipeline creation or the panel layout misbehaves.
`ContextOptions::enableValidation` / `enableDebugUtils` default to **true** in `ContextOptions::enableValidation` / `enableDebugUtils` default to **true**
every build config, so the Khronos validation layer is requested even in Release; (`src/vk/Context.h:15-16`) and nothing overrides them, so the Khronos validation
messages (error + warning severity) go through spdlog. layer is requested even in Release. The severity mask is error + warning only
(`Context.cpp:56-58`), which makes the `INFO` branch of the callback dead code.
Those messages go to **raw `spdlog::`**, i.e. spdlog's default logger — not the
`"simv"` logger that `core::Logger` creates.
`run.bat` at the repo root is **stale**: it launches `run.bat` at the repo root is **stale**: it launches
`build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never `build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never
produce. Do not point users at it without fixing the path first. produce — while its own error message tells the user to run those presets. Do not
point users at it without fixing the path first.
## Tests ## Tests
Catch2 unit tests, pure CPU/math (mesh bounds + welding). No GPU required. Catch2 unit tests, mesh bounds + welding. No GPU is touched at runtime, but the
binary is **not** free of Vulkan: static-lib propagation through
`simv_mesh → simv_core → (PRIVATE) simv_vk` makes `simv_tests.exe` link
`simv_vk`, `vk-bootstrap`, `imgui`, `volk` and `glfw3`.
```sh ```sh
cmake --build --preset windows-msvc-debug --target simv_tests cmake --build --preset windows-msvc-debug --target simv_tests
@@ -131,16 +188,14 @@ ctest --preset windows-msvc-debug
`windows-msvc-debug` is the only preset with a `testPreset` — for the other `windows-msvc-debug` is the only preset with a `testPreset` — for the other
configs, invoke `ctest` in `build/<preset>/` directly. configs, invoke `ctest` in `build/<preset>/` directly.
Single test / subset — either through CTest (each `TEST_CASE` is registered There are exactly three `TEST_CASE`s, all in `tests/unit/test_mesh.cpp`:
individually by `catch_discover_tests`): `Mesh::RecalculateBounds finds AABB` `[mesh]`, `WeldVertices collapses
near-duplicate vertices` and `WeldVertices drops degenerate triangles`
```sh (both `[mesh][decimator]`). Each is registered individually by
ctest --preset windows-msvc-debug -R "WeldVertices" -V `catch_discover_tests`, so either filter works:
```
or by running the binary with a Catch2 name or tag filter:
```sh ```sh
ctest --preset windows-msvc-debug -R "WeldVertices" -V # matches the last two
./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]" ./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]"
./build/windows-msvc-debug/tests/simv_tests.exe --list-tests ./build/windows-msvc-debug/tests/simv_tests.exe --list-tests
``` ```
@@ -149,15 +204,23 @@ or by running the binary with a Catch2 name or tag filter:
### Library / target map ### Library / target map
Arrows below show the **project** edges; each target also links third-party
libraries, and the visibility keywords matter more than they look.
``` ```
SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor (all PRIVATE)
simv_core → simv_vk (App owns Window + vk::Renderer; no Vulkan calls) + glm, spdlog; add_dependencies(… simv_shaders)
simv_editor → simv_core, simv_mesh (Camera, input, ImGui panels; no Vulkan) simv_core → PUBLIC spdlog · PRIVATE simv_vk, glfw (App owns Window + Renderer)
simv_mesh → simv_core (CPU mesh only; no Vulkan) simv_editor → PUBLIC simv_core, simv_mesh, imgui, glm (no Vulkan)
simv_vk → third-party (volk, vk-bootstrap, VMA, GLFW, GLM, spdlog, imgui) simv_mesh → PUBLIC simv_core, glm, tinyobjloader (no Vulkan)
simv_shaders → glslangValidator (GLSL → SPIR-V) simv_vk → PUBLIC Vulkan::Headers, volk, vk-bootstrap, VMA, glfw, glm, spdlog
PRIVATE imgui
simv_shaders → glslangValidator (GLSL → SPIR-V), an ALL custom target
``` ```
The `simv_core → simv_vk` edge being **PRIVATE** is what keeps `VK_NO_PROTOTYPES`
(and volk) from reaching the rest of the tree — see the Build section.
Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`, Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`,
`simv::editor`. Each library exports `src/` as its include root, so includes are `simv::editor`. Each library exports `src/` as its include root, so includes are
written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`). written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
@@ -165,66 +228,77 @@ written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
### Frame loop ### Frame loop
`main.cpp` creates `core::App`, which owns the `core::Window` and a `main.cpp` creates `core::App`, which owns the `core::Window` and a
`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` calls the UI `vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` acquires a
callback (between ImGui NewFrame/Render), then records the scene: grid + mesh into swapchain image, calls the UI callback (between ImGui NewFrame/Render), *then*
one dynamic-rendering pass with a depth attachment, followed by ImGui, then begins the command buffer and records the scene: grid + mesh into one
presents (2 frames in flight, sync2 submits). dynamic-rendering pass with a depth attachment, followed by ImGui, then presents
(2 frames in flight, `vkQueueSubmit2` / `VkImageMemoryBarrier2`).
The acquire-before-callback ordering is the root of defect 2 below — keep it in
mind before moving work into the callback.
`main.cpp` wires the editor via the UI callback: it draws the panels, applies `main.cpp` wires the editor via the UI callback: it draws the panels, applies
mouse input to the `editor::Camera`, and pushes the resulting state into the mouse input to the `editor::Camera`, and pushes the resulting state into the
renderer (`SetViewProj`, `SetRenderMode`, `SetGridVisible`). Mesh loads go through renderer (`SetRenderMode`, `SetGridVisible`, `SetViewProj`). Mesh loads go through
`MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`. `MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`.
### Mesh load path ### Mesh load path
`MeshLoadPanel` lists `*.obj` in the mesh directory and kicks off `MeshLoadPanel` lists `*.obj` in the mesh directory (extension match is
`mesh::LoadObjAsync` (worker thread, `std::future`). The future is **drained on **case-insensitive**, so `.OBJ` shows up too) and kicks off `mesh::LoadObjAsync`
the main thread** at the top of `MeshLoadPanel::Draw`, so the `OnLoaded` callback (`std::async(std::launch::async, …)`, `std::future`). The future is **drained on
— and therefore the GPU upload — always runs on the render thread. Loader the main thread** at the top of `MeshLoadPanel::Draw`, before the first
exceptions surface as the panel's status string. `ImGui::Begin`, so the `OnLoaded` callback — and therefore the GPU upload — always
runs on the render thread. Loader exceptions surface as the panel's status string.
`main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs `main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs
the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius. the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius.
Despite the file name, `mesh/MeshDecimator.h` implements **only** spatial-hash Despite the file name, `mesh/MeshDecimator.h` declares **one** function and
vertex welding (collapse near-duplicates, drop degenerate triangles, recompute implements **only** spatial-hash vertex welding (round-to-nearest cell hash,
bounds) — there is no LOD/decimation. collapse near-duplicates, drop degenerate triangles, recompute bounds) — there is
no LOD/decimation.
### Non-obvious invariants (read before editing) ### Non-obvious invariants (read before editing)
- **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no - **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no
Vulkan API calls. `vk::Renderer`'s public header is deliberately Vulkan-free Vulkan API calls — a grep for `volk|vulkan|Vk[A-Z]` across them returns zero hits.
(pImpl + glm/`RenderMode` only) so `App` can own it without pulling in volk. Keep `vk::Renderer`'s public header is deliberately Vulkan-free (pImpl + glm/`RenderMode`
it that way — do not leak `Vk*` types into the public interfaces of those modules. only) so `App` can own it without pulling in volk. Keep it that way — do not leak
Note this is **convention only**: every header is visible, so nothing stops a `Vk*` types into the public interfaces of those modules. Note this is **convention
only**: every library exports `src/` publicly, so nothing stops a
`#include <volk.h>` in `mesh/`; the compiler will not catch it. `#include <volk.h>` in `mesh/`; the compiler will not catch it.
- **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one - **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one
`vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth `vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth
attachment. The ImGui backend is initialised with `depthAttachmentFormat` set so attachment (depth `storeOp` is `DONT_CARE`). The ImGui backend is initialised with
its pipeline matches the pass; the depth buffer is recreated with the swapchain. `depthAttachmentFormat` set so its pipeline matches the pass.
- **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and - **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and
a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small depth bias so a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small *static* depth
the overlay sits on top of the fill. The device feature is requested in `Context`. bias (`constantFactor` and `slopeFactor` both `-1.0`) so the overlay sits on top
Culling is off (`VK_CULL_MODE_NONE`) — loaded models may have mixed winding. of the fill. The device feature is requested in `Context`. Culling is off
(`VK_CULL_MODE_NONE`) — loaded models may have mixed winding.
- **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance) - **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance)
the world origin; loaded models are recentred there via a translate-only model the world origin; loaded models are recentred there via a translate-only model
matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip, matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip,
isolated to `Camera.cpp`). isolated to `Camera.cpp` — the macro appears nowhere else in the tree).
- **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on - **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on
the transfer queue). `GridRenderer` builds a static host-visible line buffer once. the dedicated transfer queue). `GridRenderer` builds a static host-visible mapped
Both scene renderers use only a push constant — no descriptor sets. line buffer once in its constructor. Both scene renderers create pipeline layouts
with one push-constant range and **no** `pSetLayouts` — no descriptor sets at all.
- **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous - **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous
`MeshPC { mat4 mvp; vec4 color; }` must stay byte-identical to the `MeshPC { mat4 mvp; vec4 color; }` (range `VERTEX|FRAGMENT`) must stay
`push_constant` block in `mesh.vert`/`mesh.frag`; `color.a` is a *flag*, not byte-identical to the `push_constant` block in `mesh.vert`/`mesh.frag`; `color.a`
alpha (1 = flat-shaded, 0 = constant colour for wireframe). `GridRenderer` pushes is a *flag*, not alpha — `mesh.frag` does `mix(base, base*shade, pc.color.a)`, so
a bare `mat4` (vertex stage only). Change either side and you must change both. 1 = flat-shaded and 0 = constant colour for wireframe. `GridRenderer` pushes a
- **Swapchain recreation rebuilds sync objects.** `RecreateSwapDependent` recreates bare `mat4` (vertex stage only). Change either side and you must change both.
the per-frame `imageAvailable` semaphores (a failed acquire can leave one - **Swapchain recreation rebuilds depth *and* sync objects.** `RecreateSwapDependent`
signalled) *and* the per-swapchain-image `renderFinished` semaphores (the image runs `WaitIdle` → `swap->Recreate()` → **`CreateDepth`** → per-frame
count may change), then re-ensures both pipelines. Keep that ordering if you `imageAvailable` semaphores (a failed acquire can leave one signalled) → per-
touch resize handling. swapchain-image `renderFinished` semaphores (the image count may change) → re-ensure
both pipelines. Keep that ordering if you touch resize handling.
- **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle` - **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle`
before touching `GpuMesh` — acceptable because loads are rare; do not copy that before touching `GpuMesh` — acceptable because loads are rare; do not copy that
pattern into per-frame paths. pattern into per-frame paths. `SetMeshCpu` early-returns on an empty mesh *before*
the wait.
### Known defects (found by the Rust port, all still present here) ### Known defects (found by the Rust port, all still present here)
@@ -232,46 +306,71 @@ Porting the editor surfaced four real bugs that nobody was looking for. They are
unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md` unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md`
(branch `rust`). (branch `rust`).
1. **The CMake target graph has a cycle.** `core::App` owns `vk::Renderer`, and 1. **A source-level dependency cycle that the CMake graph hides.** `core::App` owns
`vk::Renderer` takes a `core::Window&`. It links only because `simv_vk` does `vk::Renderer`, and `vk::Renderer` takes a `core::Window&` — so the *sources*
**not** link `simv_core` — it sees the headers through are mutually dependent. The **CMake graph is acyclic**, which is exactly why
`target_include_directories(simv_vk PUBLIC ..)` and everything resolves at nothing complains: `simv_vk` links no project target at all, seeing
executable link time. CMake never complains. `core/Window.h` through `target_include_directories(simv_vk PUBLIC ..)` and
2. **Mesh upload runs mid-frame.** `MeshLoadPanel::Draw` calls `OnLoaded` from getting `Window`'s symbols at executable link time. Cargo rejects the same shape
inside `Renderer::DrawFrame`, i.e. *after* `vkAcquireNextImageKHR`; the handler outright, which is what surfaced it.
calls `SetMeshCpu` → `vkDeviceWaitIdle`. Legal, but it waits for device idle in 2. **Mesh upload runs inside the frame.** `MeshLoadPanel::Draw` calls `OnLoaded`
the middle of command recording. from the UI callback, which `Renderer::DrawFrame` invokes *after*
`vkAcquireNextImageKHR`; the handler calls `SetMeshCpu` → `vkDeviceWaitIdle`.
Legal, and command recording has not started yet (`vkBeginCommandBuffer` comes
later) — but it waits for device idle while holding an acquired swapchain image.
3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and 3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and
`Renderer`'s separate `hasMesh` flag must be kept consistent by hand. `Renderer`'s separate `hasMesh` flag must be kept consistent by hand;
`MeshRenderer::Draw` then re-checks `IsValid()` a third time.
4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category 4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category
`MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category `MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category
immediately after; the category throttles at 0.5 s and both are `Info`, so the immediately after; the category throttles at 0.5 s and both are `Info`, so the
second message is always dropped. Post-weld counts have therefore never appeared second message is always dropped (only Warn/Error/Critical bypass the throttle).
in any log. Post-weld counts have therefore never appeared in any log.
Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** (~200 lines). Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** — 362 lines
`GridRenderer`, `GpuMesh` and the depth attachment all call VMA directly. across four files, still compiled into `simv_vk`. `GridRenderer`, `GpuMesh` and the
depth attachment all call VMA directly. `Image::Desc` even defaults to
`VK_IMAGE_TYPE_3D` and an `R32G32B32A32_SFLOAT` storage image, a leftover from some
compute/CFD design. Note `README.md:53-54` still lists both as part of the working
Vulkan stack — that line is wrong.
### Shaders ### Shaders
GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders` GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders`
compiles each to `build/<preset>/spirv/editor/<name>.spv` targeting `vulkan1.3`, compiles each to `build/<preset>/spirv/editor/<name>.<stage>.spv` — the stage suffix
with `shaders/` as the `-I` root (so `#include "common/foo.glsl"` would resolve). is **kept**, so the real artefacts are `mesh.vert.spv`, `mesh.frag.spv` and so on —
Adding a shader under that globbed dir is picked up automatically targeting `vulkan1.3`, with `shaders/` as the `-I` root (so `#include
(`CONFIGURE_DEPENDS`). `mesh.frag` reconstructs a flat normal from screen-space "common/foo.glsl"` would resolve). The glob is `CONFIGURE_DEPENDS`, so a new shader
derivatives, so the vertex stream carries only positions (tight `float3`, one is picked up automatically, but it matches **only `editor/*.vert` and
binding, one attribute). **The Rust port on branch `rust` compiles this same `editor/*.frag`**; a `.comp` or `.geom` dropped there is silently ignored.
directory** — a shader change affects both versions.
`mesh.frag` reconstructs a flat normal from screen-space derivatives, so the mesh
vertex stream carries only positions (tight `float3`, one binding, one attribute).
The *grid* pipeline is different — two attributes, position + colour.
**The Rust port on branch `rust` compiles this same directory** (`build.rs` reaches
`../../../shaders`) with the same `glslangValidator` flags — a shader change affects
both versions.
### Logging ### Logging
`simv::core::Logger` (spdlog-backed, singleton) with category-based throttling. `simv::core::Logger` (spdlog-backed, singleton) with category-based throttling.
Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`, Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`,
`Vulkan`, `MeshIO`, `UI`, `Test`. Per-category level, throttle interval and `Vulkan`, `MeshIO`, `UI`, `Test` (printed as `core`, `vk`, `mesh`, `ui`, `test`).
on/off are settable at runtime (`SetMinLevel` / `SetThrottle` / `SetEnabled`).
Some low-level Vulkan code still calls `spdlog::` directly — which is the only Two things the API surface does not tell you:
reason the GPU name survives startup (it bypasses category throttling; see
defect 4 above). - **Only `MeshIO` is ever used.** The entire tree contains three `LogFmt` call
sites — two in `ObjLoader.cpp`, one in `main.cpp`. `Core`, `Vulkan`, `UI` and
`Test` have zero.
- **The runtime knobs are never turned.** `SetMinLevel` / `SetThrottle` /
`SetEnabled` exist but have no callers, and no flag, env var or UI control is
wired to them. They are settable programmatically, not configurable.
And the whole of `vk/` logs through **raw `spdlog::`**, never through `LogFmt` —
as do `core/App.cpp` and `main.cpp`'s fatal handler. That bypasses not just the
category throttle but the `"simv"` logger entirely (different pattern, unaffected
by `Logger::Shutdown()`), which is the only reason the GPU name survives startup;
see defect 4 above.
--- ---
@@ -279,42 +378,67 @@ defect 4 above).
A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions
sit side by side and build independently. Same Vulkan 1.3, dynamic rendering, sit side by side and build independently. Same Vulkan 1.3, dynamic rendering,
synchronization2, and the same `shaders/` and `assets/` directories. Needs **Rust synchronization2, and the same `shaders/` directory.
1.82+** (tested on 1.97.1) and the Vulkan SDK for `glslangValidator` only.
**The declared MSRV is wrong.** `rust/Cargo.toml:13` says `rust-version = "1.82"`
(and `rust/README.md` repeats it), but the locked `egui`/`egui-winit`/`epaint`
0.36.1 each declare `rust-version = "1.95"`, so Cargo refuses to resolve on 1.82.
**The real minimum is 1.95**; the tree is tested on 1.97.1. The Vulkan SDK is needed
for `glslangValidator` only.
```sh ```sh
cd rust cd rust
cargo build --release cargo build --release
./target/release/SimVulcan # runnable from any directory ./target/release/SimVulcan # SimVulcan.exe on Windows; package is simv-app
cargo test # mesh bounds, welding, Cube.obj, camera clip space cargo test # 13 tests + 1 #[ignore]d diagnostic
``` ```
**Work on the code with plain `cargo build`, not `--release`.** The release profile **Work on the code with plain `cargo build`, not `--release`.** The release profile
sets `lto = "thin"`, which makes a one-line edit re-optimise and relink the whole sets `lto = "thin"` with `codegen-units = 1`, which makes a one-line edit
binary: 52 s versus ~5 s in debug. re-optimise and relink the whole binary: 52 s versus ~5 s in debug.
The tests cover mesh bounds, welding (three cases), `Cube.obj` loading and its error
path, camera clip space and zoom/pitch clamping, and logger category mapping. They
need no GPU but are **not** pure math: the loader tests read real files from
`assets/meshes` via `$CARGO_MANIFEST_DIR`, so they require the repo checkout. The
ignored `weld_counts_for_bundled_assets` is the diagnostic that produced the loader
parity table below.
Five crates mirror the CMake target map, and the "all Vulkan in one place" Five crates mirror the CMake target map, and the "all Vulkan in one place"
invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have
neither `ash` nor `simv-vk` among their dependencies: neither `ash` nor `simv-vk` among their dependencies. Third-party deps in full
(each crate also depends on the workspace crates shown):
``` ```
crates/simv-core/ logger, window → log, chrono, winit crates/simv-core/ logger, window → log, chrono, winit
crates/simv-mesh/ .obj reading, welding, bounds → glam, tobj crates/simv-mesh/ .obj, welding, bounds → glam, tobj, log, thiserror
crates/simv-vk/ all Vulkan + build.rs (GLSL→SPIR-V) → ash, gpu-allocator, (+ simv-core)
egui-ash-renderer crates/simv-vk/ all Vulkan + build.rs → ash, ash-window, raw-window-handle,
crates/simv-editor/ camera, input, panels → egui (GLSL→SPIR-V) gpu-allocator, egui, egui-winit,
crates/simv-app/ App and entry point → all of the above egui-ash-renderer, winit, glam,
bytemuck, log, thiserror
(+ simv-core, simv-mesh)
crates/simv-editor/ camera, input, panels → egui, glam, log
(+ simv-core, simv-mesh)
crates/simv-app/ App and entry point → the four crates above
+ egui, winit, glam, log, anyhow
``` ```
Structural differences from C++, each with a reason (full list in Structural differences from C++, each with a reason (full list in
`docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the `docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the
target-graph cycle); the UI closure **returns** a `FrameState` instead of calling target-graph cycle); the UI closure **returns** a `FrameState` instead of calling
setters (it is invoked from a renderer method, so it cannot borrow the renderer); setters (it is invoked from a renderer method, so it cannot borrow the renderer) —
the loaded mesh is uploaded *after* the frame, not from mid-recording; `Buffer` and `Renderer`'s public API has no `Set*` methods at all; the loaded mesh is uploaded
`Image` are actually used; no pImpl (private module fields give the same *after* `draw_frame` returns, not from inside it; `Buffer` and `Image` are actually
isolation); winit owns the event loop, so `ShouldClose`/`PollEvents` are gone. used; no pImpl (private module fields give the same isolation); winit owns the
event loop, so `ShouldClose`/`PollEvents` are gone.
SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime
shader lookup and no requirement to run from a particular directory. shader lookup. Two runtime path dependencies remain, though, so "runs from any
directory" is not quite true: `assets/meshes` is searched from the executable's
ancestors **and then the CWD's**, and `pipeline_cache.bin` is written next to the
exe. Unlike the C++ build, nothing copies `assets/` — the port finds the repo-root
copy by walking up from `rust/target/<profile>/`.
The UI toolkit differs **by design**: egui instead of Dear ImGui, because the The UI toolkit differs **by design**: egui instead of Dear ImGui, because the
`imgui` crate is bindings and would compile ~40k lines of C++, defeating the point `imgui` crate is bindings and would compile ~40k lines of C++, defeating the point
@@ -322,30 +446,46 @@ of comparing ecosystems.
## Comparison result (`docs/rust_vs_cpp.md`) ## Comparison result (`docs/rust_vs_cpp.md`)
The document deliberately picks no winner; **no measurement differs by an order of The document deliberately picks no winner. Machine: i5-1135G7 / Iris Xe /
magnitude.** Machine: i5-1135G7 / Iris Xe / Windows 11, both builds release. Windows 11, both builds release.
| | C++ | Rust | | | C++ | Rust |
|---|---|---| |---|---|---|
| lines of code | 2434 | 2962 (+22 %) | | lines of code | 2434 | 2976 (+22 %) |
| release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) | | release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) |
| release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) | | release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) |
| debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** | | debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** |
| release exe | **1.18 MB** | 5.80 MB | | release exe | **1.18 MB** | 5.80 MB |
| dependency sources | 775 MB **per build dir** | **80 MB** shared registry | | dependency sources | ~400 MB **per build dir** | **80 MB** shared registry |
| startup to renderer ready | 1068 ms | **754 ms** | | startup to renderer ready (best of 3) | 1068 ms | **754 ms** |
| CPU at 60 Hz | **11.2 %** | 12.6 % | | CPU at 60 Hz | **11.2 %** | 12.6 % |
Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap
is egui rebuilding its layout every frame, not the language. The +22 % line count is egui rebuilding its layout every frame, not the language. The one sharp build
is almost entirely the **missing vk-bootstrap replacement**: `Context` + `Swapchain` cell (52 s) is the price of thin LTO, not of Rust — with LTO off it is 5.5 s
go from 301 to 606 lines, while the rest of the Vulkan layer matches nearly line against 4.4 s. Note the document's summary claim that *no* measurement differs by
for line. The one sharp build cell (52 s) is the price of thin LTO, not of Rust — an order of magnitude is false on its own numbers: 52.3 / 4.4 = **11.9×**. Every
with LTO off it is 5.5 s against 4.4 s. other cell stays under 10×.
The +22 % line count is **mostly, not almost entirely**, the missing vk-bootstrap
replacement: `Context` + `Swapchain` go from 301 to 606 lines, which is 305 of the
542-line overhang (56 %); the whole Vulkan layer accounts for 346 (64 %). The rest
is real — mesh +117, editor +104, app +90, core −64, and 176 in-module test lines
against C++'s 51.
Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj` Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj`
differs by 0.13 %, entirely due to the reading libraries (tobj drops vertices no differs by 0.13 %, entirely due to the reading libraries (tobj drops the 35
face references; fan triangulation of n-gons versus tinyobjloader's earcut). vertices no face references; fan triangulation of n-gons versus tinyobjloader's
earcut).
**Known errors in `docs/rust_vs_cpp.md`** (numbers below are the measured truth):
line-count total is 2976 / 494 comments, not 2962 / 492 — the `mesh` row is stale by
the 17 lines of the ignored diagnostic test; "775 MB per build dir" is the foreign
`build/` tree (see the Build section) and so is the matching "10 packages in the
C++ graph" — there are nine; the Rust package graph is 82, not 76; `plane.obj` has
119 faces with five or more vertices, not 121; "three Catch2 tests ported verbatim"
should be *all three ported, two more added*; and `Renderer::Impl`'s destructor is
22 lines, not thirty.
--- ---
@@ -354,14 +494,17 @@ face references; fan triangulation of n-gons versus tinyobjloader's earcut).
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
collision operator**, written from scratch in Rust (not a port): flow past bodies in collision operator**, written from scratch in Rust (not a port): flow past bodies in
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall
models, and GIF output locked to *physical* flow time. Needs **Rust 1.75+**. models, and GIF output locked to *physical* flow time. `Cargo.toml` declares
Physics is anchored to the method authors' papers in `docs/origins/`; formula **Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the
references in the code follow the 2D paper (arXiv:1507.02509). method authors' papers in `docs/origins/`; formula references in the code follow the
2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models
citing Dorschner et al. *JFM* 801 (2016) for `grad` and Malaspinas 2015 /
Coreixas et al. *PRE* 96 for `hrr`.
`README.md` there is the authoritative status/validation log (~600 lines: what is `README.md` there is the authoritative status/validation log (609 lines: what is
verified, against which equation, with what measured numbers) and `bench/README.md` verified, against which equation, with what measured numbers) and `bench/README.md`
covers the validation campaign. **Read them before changing physics or defaults** — (248 lines) covers the validation campaign. **Read them before changing physics or
most defaults are the outcome of a documented measurement, not a guess. defaults** — most defaults are the outcome of a documented measurement, not a guess.
## Build / test / run ## Build / test / run
@@ -370,10 +513,14 @@ cd docs/theory/2d_solver
cargo build --release # with the GPU backend (default feature `gpu`) cargo build --release # with the GPU backend (default feature `gpu`)
cargo build --release --no-default-features # CPU only, no wgpu cargo build --release --no-default-features # CPU only, no wgpu
cargo test --release # 29 fast tests cargo test --release # 29 fast tests
cargo test --release -- --include-ignored # + 4 long paper benchmarks (~18 s) cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s)
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
``` ```
33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the
4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two
diagnostics, not four benchmarks.
**Never build `--no-default-features` last.** Both builds write the same **Never build `--no-default-features` last.** Both builds write the same
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and `target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
`--backend gpu` then refuses to run. Order: no-default-features first, normal `--backend gpu` then refuses to run. Order: no-default-features first, normal
@@ -385,20 +532,27 @@ build second.
--gif wake.gif --gif-field vorticity --verbose full --gif wake.gif --gif-field vorticity --verbose full
``` ```
`--help` groups every key by role (Physics, Grid, Body, Time, Scheme, Animation, `--help` groups every key by role, under **Russian** headings —
Output). Note `--time <seconds>` as an alternative to `--steps`: a step is not a `Физика, Сетка, Тело, Время, Схема, Анимация, Вывод`. Note `--time <seconds>` as an
fixed slice of time (δt = u_lat·δx/u_phys), so refining the cell silently shortens alternative to `--steps` (the two conflict): a step is not a fixed slice of time
a fixed step count. (`dt = u_lat * dx / u_phys`, literally `math.rs`'s `Units::new`), so refining the
cell silently shortens a fixed step count.
Defaults worth knowing beyond the three discussed below: `--backend` is **`cpu`**,
`--collision kbc`, `--outlet extrapolate`, `--case channel`, `--refine 2`.
## File roles — one concern each ## File roles — one concern each
`src/` holds exactly these five files — no `lib.rs`, no `benches/`, no integration
tests.
| file | owns | | file | owns |
|---|---| |---|---|
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, unit conversion. Tests against the papers' formulas live here. `pub type R = f64`. | | `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. |
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. | | `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. |
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. | | `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. | | `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. |
| `src/gif.rs` | encoding, palettes, normalisation, and the physical-time frame timing. | | `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. |
## Invariants (read before editing) ## Invariants (read before editing)
@@ -407,47 +561,63 @@ a fixed step count.
`cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the `cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the
GPU silently running Bouzidi for `--wall staircase`, caught only because two GPU silently running Bouzidi for `--wall staircase`, caught only because two
models produced bit-identical output where they had to differ. Backend parity is models produced bit-identical output where they had to differ. Backend parity is
a hard requirement (CPU f64 vs GPU f32 agree to 4–5 significant digits; all four a hard requirement and `bench/parity.py` checks it: CPU f64 vs GPU f32 agree to
wall models agree to 0.007 % on Cd) and `bench/parity.py` checks it. 4–5 significant digits, and for **each** wall model the two backends agree to
0.007 % on Cd. (That number is *backend* parity, not agreement between wall
models — those differ from one another by ~1 %.)
- **The backends have different step contracts.** `Backend` in `main.rs` is an - **The backends have different step contracts.** `Backend` in `main.rs` is an
enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` / enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` /
`flush(out)` push a *batch* — results accumulate in a 128-slot ring and sync once `flush(out)` push a *batch* — results accumulate in a 128-slot ring (`HIST`,
per batch. Adding a per-step GPU readback outside `StepRec` reintroduces a declared identically in Rust and WGSL) and sync once per batch. Adding a per-step
`map_async` + `poll(Wait)` per step, which cost ~8× throughput before batching. GPU readback outside `StepRec` reintroduces a `map_async` + `poll(Wait)` per step,
- **`GREL = 1e-8` is a relative threshold** — a fraction of ⟨Δh|Δh⟩, not an which cost ~8× throughput before batching (measured 863 → 6715 steps/s).
absolute one. ⟨Δh|Δh⟩ is quadratic in non-equilibrium and physically tiny - **`GREL = 1e-8` is a relative threshold.** The test is `den > GREL * nrm`, where
(~1e-7…1e-9), so an absolute threshold fires almost everywhere and silently `den = ⟨Δh|Δh⟩` and `nrm = ⟨Δ|Δ⟩` — so GREL is a fraction of the **full**
substitutes γ = 2, i.e. plain LBGK instead of KBC. The report prints the degenerate- non-equilibrium norm, not of ⟨Δh|Δh⟩ itself. ⟨Δh|Δh⟩ is quadratic in
node fraction; on a healthy threshold it must be ~0 (except the very first step). non-equilibrium and physically tiny (~1e-7…1e-9), so an absolute threshold fires
almost everywhere and silently substitutes γ = 2, i.e. plain LBGK instead of KBC.
The report prints the degenerate-node fraction; on a healthy threshold it must be
~0 (except the very first step).
- **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no - **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no
hardware helps. Convergence studies therefore run on CPU, everything else on GPU; hardware helps. Convergence studies therefore run on CPU, everything else on GPU;
below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude. below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude.
Reductions and force sums use compensated (Kahan–Neumaier) summation — that is The **statistics reductions** (ρ, ⟨γ⟩, node and degenerate counts) use compensated
where f32 was losing most of its precision. Kahan–Neumaier summation via `kadd`. The **force/torque kernel does not** — it
accumulates in plain f32 and reduces with a naive tree. Both `README.md:366` and
earlier revisions of this file claim otherwise; the code is the authority.
- **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` / - **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` /
`wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The `wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The
workgroup → node mapping stays exactly linear, which is what lets the reductions workgroup → node mapping stays exactly linear, which is what lets the reductions
keep working; verified to 4096×4096. keep working; verified to 4096×4096.
- **Population storage has two binding layouts.** One combined binding normally; - **Population storage has two binding layouts.** Normally two bindings (`f` and
**nine bindings, one per direction**, when the adapter's max binding size is small `post`). When the adapter's max binding size is small (dzn/WSL2 caps a binding at
(dzn/WSL2 caps a binding at 128 MiB while allowing a 2047 MiB buffer). Chosen from 128 MiB while allowing a 2047 MiB buffer), each array is split across nine
adapter limits, no `switch` in the accessors; both give identical numbers, split bindings — one per direction, **18 in total** — raising the ceiling from 3.73 M to
costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS=1` to compare on one 33.5 M nodes. The choice comes from adapter limits. The split layout is the one
card. built out of `switch` statements in `fget`/`fset`/`pget`/`pset`; the combined
layout has none, which is the whole reason it is kept. Both give identical numbers;
the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to
anything other than `0`.
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but - **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
per-body force is bucketed — bodies beyond the fourth have their forces merged. per-body force is bucketed — bodies beyond the fourth have their forces merged
(`min(body, MAXB-1)` in three places), and the CLI warns when that happens.
- **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps - **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps
a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge
cannot remove it, since a standing mode has its pressure node exactly where the cannot remove it, since a standing mode has its pressure node exactly where the
sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but higher bulk viscosity, sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but ~4× higher bulk
which damps longitudinal acoustics); `--wall hrr` (**provisional** — chosen on viscosity, which damps longitudinal acoustics); `--wall hrr` (**provisional** —
scheme structure, pending campaign group C; `bouzidi` resolves geometry 4× better). chosen on scheme structure, pending campaign group C; `bouzidi` resolves geometry
~4× better). Accepted values: `--init uniform|rest`, `--kbc-model n1|n2`,
`--wall hrr|bouzidi|grad|staircase`.
## Validation campaign (`bench/`) ## Validation campaign (`bench/`)
115 runs, ≈90 GPU-hours in nine groups; each run writes logs, series, a 115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90
machine-readable summary and a GIF into its own folder under `out/` (untracked — machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs
**those are results, never clean them**). of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and
`summary.json` into its own folder under `out/` (untracked — **those are results,
never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no
animation at all.
```sh ```sh
cd bench cd bench
@@ -457,18 +627,28 @@ python preflight.py # start every scenario for two steps — catches
./run_campaign.sh --resume # run, skipping what is already done ./run_campaign.sh --resume # run, skipping what is already done
``` ```
Run `preflight.py` for real — it caught an entire group failing on a GPU limit and `run_campaign.sh` also takes `--smoke`, `--group`, `--only` and `--budget-hours`;
nine runs passing `--body-x` twice, both of which would otherwise have surfaced `run_campaign.ps1` is the Windows twin. Run `preflight.py` for real — it caught
twenty hours into a server campaign. group I failing on a GPU limit and group F's nine runs passing `--body-x` twice,
both of which would otherwise have surfaced twenty hours into a server campaign.
Note the campaign has barely started: `out/` currently holds four finished runs
(`A01`, `A02`, `A05`, `C09`) plus `campaign.log` and `summary.csv`.
## Deployment ## Deployment
Published image **`notbigghost/kbc2d:1.2.0`** (`linux/amd64`); the server needs Published image **`notbigghost/kbc2d:1.2.0`**; the server needs only
only `docker-compose.server.yml`, not the sources. `ENTRYPOINT` is the campaign `docker-compose.server.yml` (the only compose file with no `build:` section), not
driver and `CMD` defaults to `--dry-run`, so a stray `docker run` prints an the sources. `ENTRYPOINT` is the campaign driver and `CMD` defaults to `--dry-run`,
estimate instead of starting a 100-hour job. Check the card first so a stray `docker run` prints an estimate instead of starting a 100-hour job. Check
(`--profile check run --rm vulkan`) — a missing GPU is better discovered in a the card first (`--profile check run --rm vulkan`) — a missing GPU is better
minute than in an hour. discovered in a minute than in an hour. The `check` profile also carries
`preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file).
Two version caveats: the `1.2.0` tag lives only in the compose files and prose —
`Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults
`ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose
file sets `platform:` and the Dockerfile sets no `--platform`.
Two environment traps, both already handled in the image but overridable from Two environment traps, both already handled in the image but overridable from
outside: NVIDIA Container Toolkit only injects the Vulkan ICD when outside: NVIDIA Container Toolkit only injects the Vulkan ICD when
@@ -478,7 +658,8 @@ arrives over `/dev/dxg`. That path instead uses **dzn** (Mesa's Vulkan→D3D12
translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build
dzn), needs no NVIDIA runtime, and requires dzn), needs no NVIDIA runtime, and requires
`WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports `WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports
`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default. Use `conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default (the variable
is read only because `gpu.rs` builds the instance with `.with_env()`). Use
`docker-compose.wsl.yml` there. `docker-compose.wsl.yml` there.
--- ---
@@ -486,16 +667,63 @@ dzn), needs no NVIDIA runtime, and requires
# `docs/theory/` — Python prototype (all branches) # `docs/theory/` — Python prototype (all branches)
The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow
with a ×2 nested AMR patch and SDF+Bouzidi boundaries, **GPU/CuPy only, no CPU with a nested AMR patch and SDF+Bouzidi boundaries. `kbc2d` is a deliberate rewrite
fallback**. `kbc2d` is a deliberate rewrite of this, not a port, and the two differ of this, not a port. Do not fold any of it into the CMake build, and do not treat it
in two documented places (no collision inside the body; restriction skips fine as dead code — but do not treat it as one program either.
source nodes inside the body). Still live: it owns the notebook's figures.
- `solver_2x_sdf/README.md` — maps every file to the physics it owns; entry points **This directory holds four generations of the same solver.** CLAUDE.md used to name
`run.py`, `run_blockage.py`, `run_factors.py`. only the newest, which made the rest look like clutter. Oldest to newest:
- `demos_gpu/` — demo runs producing the notebook's figures/GIFs; they import
physics from `solver_2x_sdf` and never duplicate it.
- `kbc_lbm.ipynb` — the write-up; embeds pre-rendered artefacts, executes nothing.
- `docs/origins/` — the source PDFs behind both implementations.
Do not fold any of this into the CMake build, and do not treat it as dead code. | generation | files | status |
|---|---|---|
| the original write-up | `_legacy/kbc_lbm.ipynb` (64 cells, executed, outputs stored) + `_legacy/kbc_lbm.py` | superseded |
| CPU reference prototypes | `_amr_anims.py`, `_amr_anims_sdf.py` (pure NumPy) | superseded, but these are the "verified CPU prototypes" the GPU core is checked against |
| GPU core + runners | `amr_gpu_core.py` (**has a NumPy CPU fallback**) driven by `amr_cylinder_gpu.py` (no SDF) and `amr_cylinder_sdf_gpu.py` (SDF+Bouzidi); both self-titled "ОСНОВНОЙ", three modes: none / 2× / nested 2×+4× | superseded as a model, **still the producer of committed figures 11e–11g and four `*_cyl*.gif` animations** — deleting it orphans notebook images |
| optimised rewrite | `amr_opt/` (`*_opt.py`: no NumPy fallback, `cp.fuse`, precomputed link masks, CUDA Graphs) | dead intermediate — produced nothing that is committed |
| **the current model** | `solver_2x_sdf/` (17 modules + README), demos in `demos_gpu/` | live |
So "**GPU/CuPy only, no CPU fallback**" is true of `solver_2x_sdf/` (`backend.py`
raises outright), `demos_gpu/` and `amr_opt/` — but **not** of `amr_gpu_core.py`,
which silently falls back to NumPy, nor of `_amr_anims*.py` and `_legacy/`, which are
CPU-only.
Also committed and easy to mistake for junk: `figures/` (14 PNG) and `anim/`
(10 GIF, 55 MB) are tracked **deliberately** so the notebook renders on a fresh
clone; `_amr_probe_data*.npz` are probe scalars one runner writes and the other
reads; `factors.csv` is a hand-preserved snapshot of the Experiment-№1 registry
(nothing writes that path — `run_factors.py` writes `solver_2x_sdf/out/factors.csv`,
which is gitignored).
- `solver_2x_sdf/README.md` — maps every file to the physics it owns (17 rows, 17
modules, exact); entry points `run.py`, `run_blockage.py`, `run_factors.py`.
**Two defects in it:** it suggests switching `cfg.collision="bgk"` or registering
a TRT operator in a `get` registry — neither the config field nor the registry
exists, and the same README says two paragraphs later that the operator is fixed;
and it declares **all of its own quantitative results stale** pending a re-run
after the `GREL` fix, which also invalidates notebook §22–24.
- `demos_gpu/` — 8 scripts + README producing 17 of the notebook's artefacts. They
import the model's operators from `solver_2x_sdf` and add only what the model
deliberately lacks: `_common.py` reimplements a polynomial equilibrium, BGK, the
H-function and a γ-field for visualisation. `figures_static.py` is pure matplotlib
and imports no physics at all.
- `kbc_lbm.ipynb` — the write-up. It does not merely "execute nothing": it has **20
cells, all markdown, zero code cells**. Artefacts are relative-path links, not
embedded data, so it is only readable in place — and **4 of its 26 image links are
dead**, all pointing into the gitignored `solver_2x_sdf/out/`.
- `docs/origins/` — the six source PDFs behind both implementations (seven on
`research`, which adds Grad's approximation).
## How kbc2d differs from this prototype
Documented in one paragraph at `docs/theory/2d_solver/README.md:390` — **on branch
`research` only**; nothing under `docs/theory/` mentions the divergence. There are
**three** differences, not two:
1. kbc2d does not collide inside the body (those populations are fictitious; γ
statistics are then gathered strictly over fluid). The Python solver *does*
collide there — `kbc_collide` takes no fluid mask.
2. kbc2d's restriction additionally skips coarse nodes whose *fine source* node lies
inside the body. The Python restriction already gates on the **coarse
destination** mask.
3. kbc2d's CPU backend is f64 and its GPU backend f32; the Python side is fp32 by
default, f64 only via `AMR_FP64`.