diff --git a/CLAUDE.md b/CLAUDE.md index 81e1f27..184e293 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -49,18 +49,40 @@ single branch contains all of it.** Before assuming a directory exists, check | branch | what it adds | trees present | |---|---|---| -| `main`, `dev` | baseline, nothing beyond the initial import (both at the same commit) | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) | -| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images. Has unpushed commits. | `+ docs/theory/2d_solver/` | +| `main`, `dev` | baseline; both at the same commit, one past the initial import — and that commit only untracks `CLAUDE.md` | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) | +| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images | `+ docs/theory/2d_solver/`, `+ docs/origins/Grad's_aproximation.pdf` | | `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` | +All five branches (including `notes`) are fully pushed to the `CFDManager` remote; +verify with `git ls-remote CFDManager` rather than trusting a stale tracking ref. + Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are unrelated projects that merely both happen to be Rust. Neither is part of the CMake -build. +build — `CMakeLists.txt` never mentions either tree. -Untracked leftovers you will see in the working tree regardless of branch: -`build/`, `rust/target/`, `docs/theory/2d_solver/out/` (campaign results — -**the user's data, do not clean**), `pipeline_cache.bin`, `imgui.ini`. +**A checkout away from `research` does not remove `docs/theory/2d_solver/`** — the +tracked sources vanish but the untracked artefacts stay, so on `rust` that directory +contains only `out/`, `target/`, `wake.gif` and `demo_wake.gif` and looks like a +mutilated copy of the solver. To read the solver from another branch, go through +git rather than the filesystem: + +```sh +git show research:docs/theory/2d_solver/src/math.rs +git archive research docs/theory/2d_solver | tar -x -C /tmp/kbc2d # to grep it +``` + +Untracked leftovers you will see regardless of branch: `build/`, `rust/target/`, +`docs/theory/2d_solver/target/`, the `__pycache__` dirs under `docs/theory/`, and +`docs/theory/2d_solver/out/` plus the two ~9 MB GIFs beside it (campaign results — +**the user's data, do not clean**). `pipeline_cache.bin` and `imgui.ini` are written +next to whichever executable ran, i.e. in `build//src/app/` and +`rust/pipeline_cache.bin` — not at the repo root. + +**`build/` in this working tree is foreign and stale.** Its +`CMakeCache.txt` records `CMAKE_HOME_DIRECTORY=C:/Users/ivan/UAS/SimV4(Not_KBC)/SimVulcan`, +and `_deps/` carries a 262 MB `nlohmann_json` clone that this project never declares. +Do not measure anything from it without re-configuring first. --- @@ -69,15 +91,25 @@ Untracked leftovers you will see in the working tree regardless of branch: A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY, XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj` loading with selectable display modes (solid, wireframe, solid + wireframe). It -runs **no simulation** — it is a rendering skeleton with an ImGui interface (a -Viewport control panel and a Mesh load panel). Unchanged since the initial commit -on every branch. +runs **no simulation** — no compute pipeline or `.comp` shader exists anywhere, and +`Renderer.cpp` disables async compute explicitly. It is a rendering skeleton with +two ImGui windows, titled **`Viewport`** and **`Mesh`** (the class behind the second +is `MeshLoadPanel`, which is why the docs keep calling it the "mesh load panel"). +Unchanged since the initial commit on every branch — `git log --all -- src` returns +exactly one commit. ## Build -Prerequisites: **Vulkan SDK 1.3.290+** (provides `glslangValidator` for offline -shader compilation), **CMake 3.26+**, **Ninja**, a C++20 compiler (MSVC 19.36+, -gcc 11+, clang 14+). On macOS, Vulkan is via MoltenVK (Apple Silicon only). +Prerequisites: **Vulkan SDK** (for `glslangValidator`), **CMake 3.26+**, **Ninja**, +a C++20 compiler. Only CMake 3.26 and C++20 are actually enforced +(`CMakeLists.txt:1`, `:9-11`). `cmake/VulkanSetup.cmake:3` is +`find_package(Vulkan REQUIRED COMPONENTS glslangValidator)` with **no version +argument**, so the "SDK 1.3.290+" in `README.md` is a comment, not a check; the +de-facto floor comes from the pinned volk `1.3.295` and vk-bootstrap `v1.3.295`. +The compiler minimums (MSVC 19.36+, gcc 11+, clang 14+) are likewise README prose, +enforced nowhere. On macOS the only hard constraint is +`CMAKE_OSX_ARCHITECTURES: arm64` in the preset; MoltenVK is mentioned in a +`message(STATUS)`. ```sh cmake --preset windows-msvc-release @@ -86,42 +118,67 @@ cmake --build --preset windows-msvc-release Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`, `linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under -`build//`). The first configure fetches dependencies via FetchContent -(GLFW, GLM, volk, vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog, -tinyobjloader, Catch2) and needs network access — it clones nine repositories with -full history (**775 MB**, ~15 min on a slow link) and **has no shared cache**, so -every new build directory downloads all of it again. `VK_NO_PROTOTYPES` is set -project-wide; Vulkan entry points load through **volk**. +`build//`). None of them pins a compiler — they inherit whatever the +environment provides, so `README.md`'s "MSVC / clang-cl" is aspirational. -Warnings come from `simv_set_warnings` (`/W4 /permissive-`, or -`-Wall -Wextra -Wpedantic -Wshadow -Wold-style-cast …`); they are **not** errors. -`cmake/Sanitizers.cmake` defines `simv_enable_sanitizers` (Debug-only ASan/UBSan) -but no target currently calls it — wire it in manually when chasing memory bugs. +The first configure fetches nine repositories via FetchContent (GLFW, GLM, volk, +vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog, tinyobjloader, Catch2) and +needs network access. There is no `GIT_SHALLOW`, so each is a clone with full +history, and **no shared cache** is configured (the only `FETCHCONTENT_*` variable +set is `FETCHCONTENT_QUIET OFF`) — every new build directory downloads all of it +again. Budget **~400 MB of sources / ~510 MB of `_deps` once built**, not the +775 MB quoted in `docs/rust_vs_cpp.md`: that figure was measured on the foreign +`build/` tree described above and includes 262 MB of `nlohmann_json` that nothing +here declares. + +`VK_NO_PROTOTYPES` is **not** project-wide — there is no `add_compile_definitions` +at project scope. It is set `PUBLIC` on exactly two targets, `imgui` +(`CMakeLists.txt:114`) and `simv_vk` (`src/vk/CMakeLists.txt:29`). Because +`simv_core` links `simv_vk` **PRIVATE**, the definition never reaches `simv_mesh`, +`simv_editor` or `simv_tests`. Vulkan entry points load through **volk**. + +Warnings come from `simv_set_warnings` (`cmake/CompilerWarnings.cmake`), called by +all six targets: `/W4 /permissive- /Zc:preprocessor /Zc:__cplusplus /wd4100` plus +the defines `_CRT_SECURE_NO_WARNINGS`, `NOMINMAX`, `WIN32_LEAN_AND_MEAN` on MSVC; +`-Wall -Wextra -Wpedantic -Wno-unused-parameter -Wshadow -Wnon-virtual-dtor +-Wold-style-cast -Wcast-align -Wunused -Woverloaded-virtual` elsewhere. They are +**not** errors — no `/WX` or `-Werror` anywhere. `cmake/Sanitizers.cmake` defines +`simv_enable_sanitizers` (Debug-only; ASan **only** on MSVC, ASan+UBSan elsewhere) +but no target calls it — wire it in manually when chasing memory bugs. ## Running -The executable resolves SPIR-V relative to the working directory (`FindSpvPath` -probes `spirv/` and `current_path()/spirv/`). The `SimVulcan` POST_BUILD -step copies the compiled `spirv/` tree and `assets/` next to the executable, so -**run from the executable's own directory** (`build//src/app/`). `main.cpp` -also probes several `../` ancestors for `assets/meshes`. Meshes load at runtime -through the ImGui Mesh panel. +The executable resolves SPIR-V relative to the working directory (`FindSpvPath`, +`src/vk/Shader.cpp:24-33`, probes `spirv/` then `current_path()/spirv/` — +the second probe is a no-op, since the first is already CWD-relative). Two +`SimVulcan` POST_BUILD steps copy `assets/` and the compiled `spirv/` tree next to +the executable, so **run from the executable's own directory** +(`build//src/app/`). `main.cpp:40-50` additionally probes four `../` +ancestors for `assets/meshes`. Meshes load at runtime through the `Mesh` panel. Running writes two files into the CWD: `pipeline_cache.bin` (serialised -`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini`. Both are -disposable — delete them if pipeline creation or the panel layout misbehaves. +`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini` (no +`IniFilename` is ever set, so ImGui's default applies). Both are disposable — +delete them if pipeline creation or the panel layout misbehaves. -`ContextOptions::enableValidation` / `enableDebugUtils` default to **true** in -every build config, so the Khronos validation layer is requested even in Release; -messages (error + warning severity) go through spdlog. +`ContextOptions::enableValidation` / `enableDebugUtils` default to **true** +(`src/vk/Context.h:15-16`) and nothing overrides them, so the Khronos validation +layer is requested even in Release. The severity mask is error + warning only +(`Context.cpp:56-58`), which makes the `INFO` branch of the callback dead code. +Those messages go to **raw `spdlog::`**, i.e. spdlog's default logger — not the +`"simv"` logger that `core::Logger` creates. `run.bat` at the repo root is **stale**: it launches `build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never -produce. Do not point users at it without fixing the path first. +produce — while its own error message tells the user to run those presets. Do not +point users at it without fixing the path first. ## Tests -Catch2 unit tests, pure CPU/math (mesh bounds + welding). No GPU required. +Catch2 unit tests, mesh bounds + welding. No GPU is touched at runtime, but the +binary is **not** free of Vulkan: static-lib propagation through +`simv_mesh → simv_core → (PRIVATE) simv_vk` makes `simv_tests.exe` link +`simv_vk`, `vk-bootstrap`, `imgui`, `volk` and `glfw3`. ```sh cmake --build --preset windows-msvc-debug --target simv_tests @@ -131,16 +188,14 @@ ctest --preset windows-msvc-debug `windows-msvc-debug` is the only preset with a `testPreset` — for the other configs, invoke `ctest` in `build//` directly. -Single test / subset — either through CTest (each `TEST_CASE` is registered -individually by `catch_discover_tests`): - -```sh -ctest --preset windows-msvc-debug -R "WeldVertices" -V -``` - -or by running the binary with a Catch2 name or tag filter: +There are exactly three `TEST_CASE`s, all in `tests/unit/test_mesh.cpp`: +`Mesh::RecalculateBounds finds AABB` `[mesh]`, `WeldVertices collapses +near-duplicate vertices` and `WeldVertices drops degenerate triangles` +(both `[mesh][decimator]`). Each is registered individually by +`catch_discover_tests`, so either filter works: ```sh +ctest --preset windows-msvc-debug -R "WeldVertices" -V # matches the last two ./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]" ./build/windows-msvc-debug/tests/simv_tests.exe --list-tests ``` @@ -149,15 +204,23 @@ or by running the binary with a Catch2 name or tag filter: ### Library / target map +Arrows below show the **project** edges; each target also links third-party +libraries, and the visibility keywords matter more than they look. + ``` -SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor -simv_core → simv_vk (App owns Window + vk::Renderer; no Vulkan calls) -simv_editor → simv_core, simv_mesh (Camera, input, ImGui panels; no Vulkan) -simv_mesh → simv_core (CPU mesh only; no Vulkan) -simv_vk → third-party (volk, vk-bootstrap, VMA, GLFW, GLM, spdlog, imgui) -simv_shaders → glslangValidator (GLSL → SPIR-V) +SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor (all PRIVATE) + + glm, spdlog; add_dependencies(… simv_shaders) +simv_core → PUBLIC spdlog · PRIVATE simv_vk, glfw (App owns Window + Renderer) +simv_editor → PUBLIC simv_core, simv_mesh, imgui, glm (no Vulkan) +simv_mesh → PUBLIC simv_core, glm, tinyobjloader (no Vulkan) +simv_vk → PUBLIC Vulkan::Headers, volk, vk-bootstrap, VMA, glfw, glm, spdlog + PRIVATE imgui +simv_shaders → glslangValidator (GLSL → SPIR-V), an ALL custom target ``` +The `simv_core → simv_vk` edge being **PRIVATE** is what keeps `VK_NO_PROTOTYPES` +(and volk) from reaching the rest of the tree — see the Build section. + Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`, `simv::editor`. Each library exports `src/` as its include root, so includes are written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`). @@ -165,66 +228,77 @@ written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`). ### Frame loop `main.cpp` creates `core::App`, which owns the `core::Window` and a -`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` calls the UI -callback (between ImGui NewFrame/Render), then records the scene: grid + mesh into -one dynamic-rendering pass with a depth attachment, followed by ImGui, then -presents (2 frames in flight, sync2 submits). +`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` acquires a +swapchain image, calls the UI callback (between ImGui NewFrame/Render), *then* +begins the command buffer and records the scene: grid + mesh into one +dynamic-rendering pass with a depth attachment, followed by ImGui, then presents +(2 frames in flight, `vkQueueSubmit2` / `VkImageMemoryBarrier2`). + +The acquire-before-callback ordering is the root of defect 2 below — keep it in +mind before moving work into the callback. `main.cpp` wires the editor via the UI callback: it draws the panels, applies mouse input to the `editor::Camera`, and pushes the resulting state into the -renderer (`SetViewProj`, `SetRenderMode`, `SetGridVisible`). Mesh loads go through +renderer (`SetRenderMode`, `SetGridVisible`, `SetViewProj`). Mesh loads go through `MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`. ### Mesh load path -`MeshLoadPanel` lists `*.obj` in the mesh directory and kicks off -`mesh::LoadObjAsync` (worker thread, `std::future`). The future is **drained on -the main thread** at the top of `MeshLoadPanel::Draw`, so the `OnLoaded` callback -— and therefore the GPU upload — always runs on the render thread. Loader -exceptions surface as the panel's status string. +`MeshLoadPanel` lists `*.obj` in the mesh directory (extension match is +**case-insensitive**, so `.OBJ` shows up too) and kicks off `mesh::LoadObjAsync` +(`std::async(std::launch::async, …)`, `std::future`). The future is **drained on +the main thread** at the top of `MeshLoadPanel::Draw`, before the first +`ImGui::Begin`, so the `OnLoaded` callback — and therefore the GPU upload — always +runs on the render thread. Loader exceptions surface as the panel's status string. `main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius. -Despite the file name, `mesh/MeshDecimator.h` implements **only** spatial-hash -vertex welding (collapse near-duplicates, drop degenerate triangles, recompute -bounds) — there is no LOD/decimation. +Despite the file name, `mesh/MeshDecimator.h` declares **one** function and +implements **only** spatial-hash vertex welding (round-to-nearest cell hash, +collapse near-duplicates, drop degenerate triangles, recompute bounds) — there is +no LOD/decimation. ### Non-obvious invariants (read before editing) - **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no - Vulkan API calls. `vk::Renderer`'s public header is deliberately Vulkan-free - (pImpl + glm/`RenderMode` only) so `App` can own it without pulling in volk. Keep - it that way — do not leak `Vk*` types into the public interfaces of those modules. - Note this is **convention only**: every header is visible, so nothing stops a + Vulkan API calls — a grep for `volk|vulkan|Vk[A-Z]` across them returns zero hits. + `vk::Renderer`'s public header is deliberately Vulkan-free (pImpl + glm/`RenderMode` + only) so `App` can own it without pulling in volk. Keep it that way — do not leak + `Vk*` types into the public interfaces of those modules. Note this is **convention + only**: every library exports `src/` publicly, so nothing stops a `#include ` in `mesh/`; the compiler will not catch it. - **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one `vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth - attachment. The ImGui backend is initialised with `depthAttachmentFormat` set so - its pipeline matches the pass; the depth buffer is recreated with the swapchain. + attachment (depth `storeOp` is `DONT_CARE`). The ImGui backend is initialised with + `depthAttachmentFormat` set so its pipeline matches the pass. - **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and - a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small depth bias so - the overlay sits on top of the fill. The device feature is requested in `Context`. - Culling is off (`VK_CULL_MODE_NONE`) — loaded models may have mixed winding. + a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small *static* depth + bias (`constantFactor` and `slopeFactor` both `-1.0`) so the overlay sits on top + of the fill. The device feature is requested in `Context`. Culling is off + (`VK_CULL_MODE_NONE`) — loaded models may have mixed winding. - **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance) the world origin; loaded models are recentred there via a translate-only model matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip, - isolated to `Camera.cpp`). + isolated to `Camera.cpp` — the macro appears nowhere else in the tree). - **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on - the transfer queue). `GridRenderer` builds a static host-visible line buffer once. - Both scene renderers use only a push constant — no descriptor sets. + the dedicated transfer queue). `GridRenderer` builds a static host-visible mapped + line buffer once in its constructor. Both scene renderers create pipeline layouts + with one push-constant range and **no** `pSetLayouts` — no descriptor sets at all. - **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous - `MeshPC { mat4 mvp; vec4 color; }` must stay byte-identical to the - `push_constant` block in `mesh.vert`/`mesh.frag`; `color.a` is a *flag*, not - alpha (1 = flat-shaded, 0 = constant colour for wireframe). `GridRenderer` pushes - a bare `mat4` (vertex stage only). Change either side and you must change both. -- **Swapchain recreation rebuilds sync objects.** `RecreateSwapDependent` recreates - the per-frame `imageAvailable` semaphores (a failed acquire can leave one - signalled) *and* the per-swapchain-image `renderFinished` semaphores (the image - count may change), then re-ensures both pipelines. Keep that ordering if you - touch resize handling. + `MeshPC { mat4 mvp; vec4 color; }` (range `VERTEX|FRAGMENT`) must stay + byte-identical to the `push_constant` block in `mesh.vert`/`mesh.frag`; `color.a` + is a *flag*, not alpha — `mesh.frag` does `mix(base, base*shade, pc.color.a)`, so + 1 = flat-shaded and 0 = constant colour for wireframe. `GridRenderer` pushes a + bare `mat4` (vertex stage only). Change either side and you must change both. +- **Swapchain recreation rebuilds depth *and* sync objects.** `RecreateSwapDependent` + runs `WaitIdle` → `swap->Recreate()` → **`CreateDepth`** → per-frame + `imageAvailable` semaphores (a failed acquire can leave one signalled) → per- + swapchain-image `renderFinished` semaphores (the image count may change) → re-ensure + both pipelines. Keep that ordering if you touch resize handling. - **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle` before touching `GpuMesh` — acceptable because loads are rare; do not copy that - pattern into per-frame paths. + pattern into per-frame paths. `SetMeshCpu` early-returns on an empty mesh *before* + the wait. ### Known defects (found by the Rust port, all still present here) @@ -232,46 +306,71 @@ Porting the editor surfaced four real bugs that nobody was looking for. They are unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md` (branch `rust`). -1. **The CMake target graph has a cycle.** `core::App` owns `vk::Renderer`, and - `vk::Renderer` takes a `core::Window&`. It links only because `simv_vk` does - **not** link `simv_core` — it sees the headers through - `target_include_directories(simv_vk PUBLIC ..)` and everything resolves at - executable link time. CMake never complains. -2. **Mesh upload runs mid-frame.** `MeshLoadPanel::Draw` calls `OnLoaded` from - inside `Renderer::DrawFrame`, i.e. *after* `vkAcquireNextImageKHR`; the handler - calls `SetMeshCpu` → `vkDeviceWaitIdle`. Legal, but it waits for device idle in - the middle of command recording. +1. **A source-level dependency cycle that the CMake graph hides.** `core::App` owns + `vk::Renderer`, and `vk::Renderer` takes a `core::Window&` — so the *sources* + are mutually dependent. The **CMake graph is acyclic**, which is exactly why + nothing complains: `simv_vk` links no project target at all, seeing + `core/Window.h` through `target_include_directories(simv_vk PUBLIC ..)` and + getting `Window`'s symbols at executable link time. Cargo rejects the same shape + outright, which is what surfaced it. +2. **Mesh upload runs inside the frame.** `MeshLoadPanel::Draw` calls `OnLoaded` + from the UI callback, which `Renderer::DrawFrame` invokes *after* + `vkAcquireNextImageKHR`; the handler calls `SetMeshCpu` → `vkDeviceWaitIdle`. + Legal, and command recording has not started yet (`vkBeginCommandBuffer` comes + later) — but it waits for device idle while holding an acquired swapchain image. 3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and - `Renderer`'s separate `hasMesh` flag must be kept consistent by hand. + `Renderer`'s separate `hasMesh` flag must be kept consistent by hand; + `MeshRenderer::Draw` then re-checks `IsValid()` a third time. 4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category `MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category immediately after; the category throttles at 0.5 s and both are `Info`, so the - second message is always dropped. Post-weld counts have therefore never appeared - in any log. + second message is always dropped (only Warn/Error/Critical bypass the throttle). + Post-weld counts have therefore never appeared in any log. -Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** (~200 lines). -`GridRenderer`, `GpuMesh` and the depth attachment all call VMA directly. +Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** — 362 lines +across four files, still compiled into `simv_vk`. `GridRenderer`, `GpuMesh` and the +depth attachment all call VMA directly. `Image::Desc` even defaults to +`VK_IMAGE_TYPE_3D` and an `R32G32B32A32_SFLOAT` storage image, a leftover from some +compute/CFD design. Note `README.md:53-54` still lists both as part of the working +Vulkan stack — that line is wrong. ### Shaders GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders` -compiles each to `build//spirv/editor/.spv` targeting `vulkan1.3`, -with `shaders/` as the `-I` root (so `#include "common/foo.glsl"` would resolve). -Adding a shader under that globbed dir is picked up automatically -(`CONFIGURE_DEPENDS`). `mesh.frag` reconstructs a flat normal from screen-space -derivatives, so the vertex stream carries only positions (tight `float3`, one -binding, one attribute). **The Rust port on branch `rust` compiles this same -directory** — a shader change affects both versions. +compiles each to `build//spirv/editor/..spv` — the stage suffix +is **kept**, so the real artefacts are `mesh.vert.spv`, `mesh.frag.spv` and so on — +targeting `vulkan1.3`, with `shaders/` as the `-I` root (so `#include +"common/foo.glsl"` would resolve). The glob is `CONFIGURE_DEPENDS`, so a new shader +is picked up automatically, but it matches **only `editor/*.vert` and +`editor/*.frag`**; a `.comp` or `.geom` dropped there is silently ignored. + +`mesh.frag` reconstructs a flat normal from screen-space derivatives, so the mesh +vertex stream carries only positions (tight `float3`, one binding, one attribute). +The *grid* pipeline is different — two attributes, position + colour. +**The Rust port on branch `rust` compiles this same directory** (`build.rs` reaches +`../../../shaders`) with the same `glslangValidator` flags — a shader change affects +both versions. ### Logging `simv::core::Logger` (spdlog-backed, singleton) with category-based throttling. Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`, -`Vulkan`, `MeshIO`, `UI`, `Test`. Per-category level, throttle interval and -on/off are settable at runtime (`SetMinLevel` / `SetThrottle` / `SetEnabled`). -Some low-level Vulkan code still calls `spdlog::` directly — which is the only -reason the GPU name survives startup (it bypasses category throttling; see -defect 4 above). +`Vulkan`, `MeshIO`, `UI`, `Test` (printed as `core`, `vk`, `mesh`, `ui`, `test`). + +Two things the API surface does not tell you: + +- **Only `MeshIO` is ever used.** The entire tree contains three `LogFmt` call + sites — two in `ObjLoader.cpp`, one in `main.cpp`. `Core`, `Vulkan`, `UI` and + `Test` have zero. +- **The runtime knobs are never turned.** `SetMinLevel` / `SetThrottle` / + `SetEnabled` exist but have no callers, and no flag, env var or UI control is + wired to them. They are settable programmatically, not configurable. + +And the whole of `vk/` logs through **raw `spdlog::`**, never through `LogFmt` — +as do `core/App.cpp` and `main.cpp`'s fatal handler. That bypasses not just the +category throttle but the `"simv"` logger entirely (different pattern, unaffected +by `Logger::Shutdown()`), which is the only reason the GPU name survives startup; +see defect 4 above. --- @@ -279,42 +378,67 @@ defect 4 above). A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions sit side by side and build independently. Same Vulkan 1.3, dynamic rendering, -synchronization2, and the same `shaders/` and `assets/` directories. Needs **Rust -1.82+** (tested on 1.97.1) and the Vulkan SDK for `glslangValidator` only. +synchronization2, and the same `shaders/` directory. + +**The declared MSRV is wrong.** `rust/Cargo.toml:13` says `rust-version = "1.82"` +(and `rust/README.md` repeats it), but the locked `egui`/`egui-winit`/`epaint` +0.36.1 each declare `rust-version = "1.95"`, so Cargo refuses to resolve on 1.82. +**The real minimum is 1.95**; the tree is tested on 1.97.1. The Vulkan SDK is needed +for `glslangValidator` only. ```sh cd rust cargo build --release -./target/release/SimVulcan # runnable from any directory -cargo test # mesh bounds, welding, Cube.obj, camera clip space +./target/release/SimVulcan # SimVulcan.exe on Windows; package is simv-app +cargo test # 13 tests + 1 #[ignore]d diagnostic ``` **Work on the code with plain `cargo build`, not `--release`.** The release profile -sets `lto = "thin"`, which makes a one-line edit re-optimise and relink the whole -binary: 52 s versus ~5 s in debug. +sets `lto = "thin"` with `codegen-units = 1`, which makes a one-line edit +re-optimise and relink the whole binary: 52 s versus ~5 s in debug. + +The tests cover mesh bounds, welding (three cases), `Cube.obj` loading and its error +path, camera clip space and zoom/pitch clamping, and logger category mapping. They +need no GPU but are **not** pure math: the loader tests read real files from +`assets/meshes` via `$CARGO_MANIFEST_DIR`, so they require the repo checkout. The +ignored `weld_counts_for_bundled_assets` is the diagnostic that produced the loader +parity table below. Five crates mirror the CMake target map, and the "all Vulkan in one place" invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have -neither `ash` nor `simv-vk` among their dependencies: +neither `ash` nor `simv-vk` among their dependencies. Third-party deps in full +(each crate also depends on the workspace crates shown): ``` -crates/simv-core/ logger, window → log, chrono, winit -crates/simv-mesh/ .obj reading, welding, bounds → glam, tobj -crates/simv-vk/ all Vulkan + build.rs (GLSL→SPIR-V) → ash, gpu-allocator, - egui-ash-renderer -crates/simv-editor/ camera, input, panels → egui -crates/simv-app/ App and entry point → all of the above +crates/simv-core/ logger, window → log, chrono, winit +crates/simv-mesh/ .obj, welding, bounds → glam, tobj, log, thiserror + (+ simv-core) +crates/simv-vk/ all Vulkan + build.rs → ash, ash-window, raw-window-handle, + (GLSL→SPIR-V) gpu-allocator, egui, egui-winit, + egui-ash-renderer, winit, glam, + bytemuck, log, thiserror + (+ simv-core, simv-mesh) +crates/simv-editor/ camera, input, panels → egui, glam, log + (+ simv-core, simv-mesh) +crates/simv-app/ App and entry point → the four crates above + + egui, winit, glam, log, anyhow ``` Structural differences from C++, each with a reason (full list in `docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the target-graph cycle); the UI closure **returns** a `FrameState` instead of calling -setters (it is invoked from a renderer method, so it cannot borrow the renderer); -the loaded mesh is uploaded *after* the frame, not from mid-recording; `Buffer` and -`Image` are actually used; no pImpl (private module fields give the same -isolation); winit owns the event loop, so `ShouldClose`/`PollEvents` are gone. +setters (it is invoked from a renderer method, so it cannot borrow the renderer) — +`Renderer`'s public API has no `Set*` methods at all; the loaded mesh is uploaded +*after* `draw_frame` returns, not from inside it; `Buffer` and `Image` are actually +used; no pImpl (private module fields give the same isolation); winit owns the +event loop, so `ShouldClose`/`PollEvents` are gone. + SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime -shader lookup and no requirement to run from a particular directory. +shader lookup. Two runtime path dependencies remain, though, so "runs from any +directory" is not quite true: `assets/meshes` is searched from the executable's +ancestors **and then the CWD's**, and `pipeline_cache.bin` is written next to the +exe. Unlike the C++ build, nothing copies `assets/` — the port finds the repo-root +copy by walking up from `rust/target//`. The UI toolkit differs **by design**: egui instead of Dear ImGui, because the `imgui` crate is bindings and would compile ~40k lines of C++, defeating the point @@ -322,30 +446,46 @@ of comparing ecosystems. ## Comparison result (`docs/rust_vs_cpp.md`) -The document deliberately picks no winner; **no measurement differs by an order of -magnitude.** Machine: i5-1135G7 / Iris Xe / Windows 11, both builds release. +The document deliberately picks no winner. Machine: i5-1135G7 / Iris Xe / +Windows 11, both builds release. | | C++ | Rust | |---|---|---| -| lines of code | 2434 | 2962 (+22 %) | +| lines of code | 2434 | 2976 (+22 %) | | release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) | | release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) | | debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** | | release exe | **1.18 MB** | 5.80 MB | -| dependency sources | 775 MB **per build dir** | **80 MB** shared registry | -| startup to renderer ready | 1068 ms | **754 ms** | +| dependency sources | ~400 MB **per build dir** | **80 MB** shared registry | +| startup to renderer ready (best of 3) | 1068 ms | **754 ms** | | CPU at 60 Hz | **11.2 %** | 12.6 % | Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap -is egui rebuilding its layout every frame, not the language. The +22 % line count -is almost entirely the **missing vk-bootstrap replacement**: `Context` + `Swapchain` -go from 301 to 606 lines, while the rest of the Vulkan layer matches nearly line -for line. The one sharp build cell (52 s) is the price of thin LTO, not of Rust — -with LTO off it is 5.5 s against 4.4 s. +is egui rebuilding its layout every frame, not the language. The one sharp build +cell (52 s) is the price of thin LTO, not of Rust — with LTO off it is 5.5 s +against 4.4 s. Note the document's summary claim that *no* measurement differs by +an order of magnitude is false on its own numbers: 52.3 / 4.4 = **11.9×**. Every +other cell stays under 10×. + +The +22 % line count is **mostly, not almost entirely**, the missing vk-bootstrap +replacement: `Context` + `Swapchain` go from 301 to 606 lines, which is 305 of the +542-line overhang (56 %); the whole Vulkan layer accounts for 346 (64 %). The rest +is real — mesh +117, editor +104, app +90, core −64, and 176 in-module test lines +against C++'s 51. Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj` -differs by 0.13 %, entirely due to the reading libraries (tobj drops vertices no -face references; fan triangulation of n-gons versus tinyobjloader's earcut). +differs by 0.13 %, entirely due to the reading libraries (tobj drops the 35 +vertices no face references; fan triangulation of n-gons versus tinyobjloader's +earcut). + +**Known errors in `docs/rust_vs_cpp.md`** (numbers below are the measured truth): +line-count total is 2976 / 494 comments, not 2962 / 492 — the `mesh` row is stale by +the 17 lines of the ignored diagnostic test; "775 MB per build dir" is the foreign +`build/` tree (see the Build section) and so is the matching "10 packages in the +C++ graph" — there are nine; the Rust package graph is 82, not 76; `plane.obj` has +119 faces with five or more vertices, not 121; "three Catch2 tests ported verbatim" +should be *all three ported, two more added*; and `Renderer::Impl`'s destructor is +22 lines, not thirty. --- @@ -354,14 +494,17 @@ face references; fan triangulation of n-gons versus tinyobjloader's earcut). The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC collision operator**, written from scratch in Rust (not a port): flow past bodies in a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall -models, and GIF output locked to *physical* flow time. Needs **Rust 1.75+**. -Physics is anchored to the method authors' papers in `docs/origins/`; formula -references in the code follow the 2D paper (arXiv:1507.02509). +models, and GIF output locked to *physical* flow time. `Cargo.toml` declares +**Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the +method authors' papers in `docs/origins/`; formula references in the code follow the +2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models +citing Dorschner et al. *JFM* 801 (2016) for `grad` and Malaspinas 2015 / +Coreixas et al. *PRE* 96 for `hrr`. -`README.md` there is the authoritative status/validation log (~600 lines: what is +`README.md` there is the authoritative status/validation log (609 lines: what is verified, against which equation, with what measured numbers) and `bench/README.md` -covers the validation campaign. **Read them before changing physics or defaults** — -most defaults are the outcome of a documented measurement, not a guess. +(248 lines) covers the validation campaign. **Read them before changing physics or +defaults** — most defaults are the outcome of a documented measurement, not a guess. ## Build / test / run @@ -370,10 +513,14 @@ cd docs/theory/2d_solver cargo build --release # with the GPU backend (default feature `gpu`) cargo build --release --no-default-features # CPU only, no wgpu cargo test --release # 29 fast tests -cargo test --release -- --include-ignored # + 4 long paper benchmarks (~18 s) +cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s) cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic ``` +33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the +4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two +diagnostics, not four benchmarks. + **Never build `--no-default-features` last.** Both builds write the same `target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and `--backend gpu` then refuses to run. Order: no-default-features first, normal @@ -385,20 +532,27 @@ build second. --gif wake.gif --gif-field vorticity --verbose full ``` -`--help` groups every key by role (Physics, Grid, Body, Time, Scheme, Animation, -Output). Note `--time ` as an alternative to `--steps`: a step is not a -fixed slice of time (δt = u_lat·δx/u_phys), so refining the cell silently shortens -a fixed step count. +`--help` groups every key by role, under **Russian** headings — +`Физика, Сетка, Тело, Время, Схема, Анимация, Вывод`. Note `--time ` as an +alternative to `--steps` (the two conflict): a step is not a fixed slice of time +(`dt = u_lat * dx / u_phys`, literally `math.rs`'s `Units::new`), so refining the +cell silently shortens a fixed step count. + +Defaults worth knowing beyond the three discussed below: `--backend` is **`cpu`**, +`--collision kbc`, `--outlet extrapolate`, `--case channel`, `--refine 2`. ## File roles — one concern each +`src/` holds exactly these five files — no `lib.rs`, no `benches/`, no integration +tests. + | file | owns | |---|---| -| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, unit conversion. Tests against the papers' formulas live here. `pub type R = f64`. | -| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. | +| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. | +| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. | | `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. | | `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. | -| `src/gif.rs` | encoding, palettes, normalisation, and the physical-time frame timing. | +| `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. | ## Invariants (read before editing) @@ -407,47 +561,63 @@ a fixed step count. `cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the GPU silently running Bouzidi for `--wall staircase`, caught only because two models produced bit-identical output where they had to differ. Backend parity is - a hard requirement (CPU f64 vs GPU f32 agree to 4–5 significant digits; all four - wall models agree to 0.007 % on Cd) and `bench/parity.py` checks it. + a hard requirement and `bench/parity.py` checks it: CPU f64 vs GPU f32 agree to + 4–5 significant digits, and for **each** wall model the two backends agree to + 0.007 % on Cd. (That number is *backend* parity, not agreement between wall + models — those differ from one another by ~1 %.) - **The backends have different step contracts.** `Backend` in `main.rs` is an enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` / - `flush(out)` push a *batch* — results accumulate in a 128-slot ring and sync once - per batch. Adding a per-step GPU readback outside `StepRec` reintroduces a - `map_async` + `poll(Wait)` per step, which cost ~8× throughput before batching. -- **`GREL = 1e-8` is a relative threshold** — a fraction of ⟨Δh|Δh⟩, not an - absolute one. ⟨Δh|Δh⟩ is quadratic in non-equilibrium and physically tiny - (~1e-7…1e-9), so an absolute threshold fires almost everywhere and silently - substitutes γ = 2, i.e. plain LBGK instead of KBC. The report prints the degenerate- - node fraction; on a healthy threshold it must be ~0 (except the very first step). + `flush(out)` push a *batch* — results accumulate in a 128-slot ring (`HIST`, + declared identically in Rust and WGSL) and sync once per batch. Adding a per-step + GPU readback outside `StepRec` reintroduces a `map_async` + `poll(Wait)` per step, + which cost ~8× throughput before batching (measured 863 → 6715 steps/s). +- **`GREL = 1e-8` is a relative threshold.** The test is `den > GREL * nrm`, where + `den = ⟨Δh|Δh⟩` and `nrm = ⟨Δ|Δ⟩` — so GREL is a fraction of the **full** + non-equilibrium norm, not of ⟨Δh|Δh⟩ itself. ⟨Δh|Δh⟩ is quadratic in + non-equilibrium and physically tiny (~1e-7…1e-9), so an absolute threshold fires + almost everywhere and silently substitutes γ = 2, i.e. plain LBGK instead of KBC. + The report prints the degenerate-node fraction; on a healthy threshold it must be + ~0 (except the very first step). - **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no hardware helps. Convergence studies therefore run on CPU, everything else on GPU; below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude. - Reductions and force sums use compensated (Kahan–Neumaier) summation — that is - where f32 was losing most of its precision. + The **statistics reductions** (ρ, ⟨γ⟩, node and degenerate counts) use compensated + Kahan–Neumaier summation via `kadd`. The **force/torque kernel does not** — it + accumulates in plain f32 and reduces with a naive tree. Both `README.md:366` and + earlier revisions of this file claim otherwise; the code is the authority. - **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` / `wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The workgroup → node mapping stays exactly linear, which is what lets the reductions keep working; verified to 4096×4096. -- **Population storage has two binding layouts.** One combined binding normally; - **nine bindings, one per direction**, when the adapter's max binding size is small - (dzn/WSL2 caps a binding at 128 MiB while allowing a 2047 MiB buffer). Chosen from - adapter limits, no `switch` in the accessors; both give identical numbers, split - costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS=1` to compare on one - card. +- **Population storage has two binding layouts.** Normally two bindings (`f` and + `post`). When the adapter's max binding size is small (dzn/WSL2 caps a binding at + 128 MiB while allowing a 2047 MiB buffer), each array is split across nine + bindings — one per direction, **18 in total** — raising the ceiling from 3.73 M to + 33.5 M nodes. The choice comes from adapter limits. The split layout is the one + built out of `switch` statements in `fget`/`fset`/`pget`/`pset`; the combined + layout has none, which is the whole reason it is kept. Both give identical numbers; + the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to + anything other than `0`. - **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but - per-body force is bucketed — bodies beyond the fourth have their forces merged. + per-body force is bucketed — bodies beyond the fourth have their forces merged + (`min(body, MAXB-1)` in three places), and the CLI warns when that happens. - **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge cannot remove it, since a standing mode has its pressure node exactly where the - sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but higher bulk viscosity, - which damps longitudinal acoustics); `--wall hrr` (**provisional** — chosen on - scheme structure, pending campaign group C; `bouzidi` resolves geometry 4× better). + sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but ~4× higher bulk + viscosity, which damps longitudinal acoustics); `--wall hrr` (**provisional** — + chosen on scheme structure, pending campaign group C; `bouzidi` resolves geometry + ~4× better). Accepted values: `--init uniform|rest`, `--kbc-model n1|n2`, + `--wall hrr|bouzidi|grad|staircase`. ## Validation campaign (`bench/`) -115 runs, ≈90 GPU-hours in nine groups; each run writes logs, series, a -machine-readable summary and a GIF into its own folder under `out/` (untracked — -**those are results, never clean them**). +115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90 +machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs +of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and +`summary.json` into its own folder under `out/` (untracked — **those are results, +never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no +animation at all. ```sh cd bench @@ -457,18 +627,28 @@ python preflight.py # start every scenario for two steps — catches ./run_campaign.sh --resume # run, skipping what is already done ``` -Run `preflight.py` for real — it caught an entire group failing on a GPU limit and -nine runs passing `--body-x` twice, both of which would otherwise have surfaced -twenty hours into a server campaign. +`run_campaign.sh` also takes `--smoke`, `--group`, `--only` and `--budget-hours`; +`run_campaign.ps1` is the Windows twin. Run `preflight.py` for real — it caught +group I failing on a GPU limit and group F's nine runs passing `--body-x` twice, +both of which would otherwise have surfaced twenty hours into a server campaign. + +Note the campaign has barely started: `out/` currently holds four finished runs +(`A01`, `A02`, `A05`, `C09`) plus `campaign.log` and `summary.csv`. ## Deployment -Published image **`notbigghost/kbc2d:1.2.0`** (`linux/amd64`); the server needs -only `docker-compose.server.yml`, not the sources. `ENTRYPOINT` is the campaign -driver and `CMD` defaults to `--dry-run`, so a stray `docker run` prints an -estimate instead of starting a 100-hour job. Check the card first -(`--profile check run --rm vulkan`) — a missing GPU is better discovered in a -minute than in an hour. +Published image **`notbigghost/kbc2d:1.2.0`**; the server needs only +`docker-compose.server.yml` (the only compose file with no `build:` section), not +the sources. `ENTRYPOINT` is the campaign driver and `CMD` defaults to `--dry-run`, +so a stray `docker run` prints an estimate instead of starting a 100-hour job. Check +the card first (`--profile check run --rm vulkan`) — a missing GPU is better +discovered in a minute than in an hour. The `check` profile also carries +`preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file). + +Two version caveats: the `1.2.0` tag lives only in the compose files and prose — +`Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults +`ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose +file sets `platform:` and the Dockerfile sets no `--platform`. Two environment traps, both already handled in the image but overridable from outside: NVIDIA Container Toolkit only injects the Vulkan ICD when @@ -478,7 +658,8 @@ arrives over `/dev/dxg`. That path instead uses **dzn** (Mesa's Vulkan→D3D12 translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build dzn), needs no NVIDIA runtime, and requires `WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports -`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default. Use +`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default (the variable +is read only because `gpu.rs` builds the instance with `.with_env()`). Use `docker-compose.wsl.yml` there. --- @@ -486,16 +667,63 @@ dzn), needs no NVIDIA runtime, and requires # `docs/theory/` — Python prototype (all branches) The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow -with a ×2 nested AMR patch and SDF+Bouzidi boundaries, **GPU/CuPy only, no CPU -fallback**. `kbc2d` is a deliberate rewrite of this, not a port, and the two differ -in two documented places (no collision inside the body; restriction skips fine -source nodes inside the body). Still live: it owns the notebook's figures. +with a nested AMR patch and SDF+Bouzidi boundaries. `kbc2d` is a deliberate rewrite +of this, not a port. Do not fold any of it into the CMake build, and do not treat it +as dead code — but do not treat it as one program either. -- `solver_2x_sdf/README.md` — maps every file to the physics it owns; entry points - `run.py`, `run_blockage.py`, `run_factors.py`. -- `demos_gpu/` — demo runs producing the notebook's figures/GIFs; they import - physics from `solver_2x_sdf` and never duplicate it. -- `kbc_lbm.ipynb` — the write-up; embeds pre-rendered artefacts, executes nothing. -- `docs/origins/` — the source PDFs behind both implementations. +**This directory holds four generations of the same solver.** CLAUDE.md used to name +only the newest, which made the rest look like clutter. Oldest to newest: -Do not fold any of this into the CMake build, and do not treat it as dead code. +| generation | files | status | +|---|---|---| +| the original write-up | `_legacy/kbc_lbm.ipynb` (64 cells, executed, outputs stored) + `_legacy/kbc_lbm.py` | superseded | +| CPU reference prototypes | `_amr_anims.py`, `_amr_anims_sdf.py` (pure NumPy) | superseded, but these are the "verified CPU prototypes" the GPU core is checked against | +| GPU core + runners | `amr_gpu_core.py` (**has a NumPy CPU fallback**) driven by `amr_cylinder_gpu.py` (no SDF) and `amr_cylinder_sdf_gpu.py` (SDF+Bouzidi); both self-titled "ОСНОВНОЙ", three modes: none / 2× / nested 2×+4× | superseded as a model, **still the producer of committed figures 11e–11g and four `*_cyl*.gif` animations** — deleting it orphans notebook images | +| optimised rewrite | `amr_opt/` (`*_opt.py`: no NumPy fallback, `cp.fuse`, precomputed link masks, CUDA Graphs) | dead intermediate — produced nothing that is committed | +| **the current model** | `solver_2x_sdf/` (17 modules + README), demos in `demos_gpu/` | live | + +So "**GPU/CuPy only, no CPU fallback**" is true of `solver_2x_sdf/` (`backend.py` +raises outright), `demos_gpu/` and `amr_opt/` — but **not** of `amr_gpu_core.py`, +which silently falls back to NumPy, nor of `_amr_anims*.py` and `_legacy/`, which are +CPU-only. + +Also committed and easy to mistake for junk: `figures/` (14 PNG) and `anim/` +(10 GIF, 55 MB) are tracked **deliberately** so the notebook renders on a fresh +clone; `_amr_probe_data*.npz` are probe scalars one runner writes and the other +reads; `factors.csv` is a hand-preserved snapshot of the Experiment-№1 registry +(nothing writes that path — `run_factors.py` writes `solver_2x_sdf/out/factors.csv`, +which is gitignored). + +- `solver_2x_sdf/README.md` — maps every file to the physics it owns (17 rows, 17 + modules, exact); entry points `run.py`, `run_blockage.py`, `run_factors.py`. + **Two defects in it:** it suggests switching `cfg.collision="bgk"` or registering + a TRT operator in a `get` registry — neither the config field nor the registry + exists, and the same README says two paragraphs later that the operator is fixed; + and it declares **all of its own quantitative results stale** pending a re-run + after the `GREL` fix, which also invalidates notebook §22–24. +- `demos_gpu/` — 8 scripts + README producing 17 of the notebook's artefacts. They + import the model's operators from `solver_2x_sdf` and add only what the model + deliberately lacks: `_common.py` reimplements a polynomial equilibrium, BGK, the + H-function and a γ-field for visualisation. `figures_static.py` is pure matplotlib + and imports no physics at all. +- `kbc_lbm.ipynb` — the write-up. It does not merely "execute nothing": it has **20 + cells, all markdown, zero code cells**. Artefacts are relative-path links, not + embedded data, so it is only readable in place — and **4 of its 26 image links are + dead**, all pointing into the gitignored `solver_2x_sdf/out/`. +- `docs/origins/` — the six source PDFs behind both implementations (seven on + `research`, which adds Grad's approximation). + +## How kbc2d differs from this prototype + +Documented in one paragraph at `docs/theory/2d_solver/README.md:390` — **on branch +`research` only**; nothing under `docs/theory/` mentions the divergence. There are +**three** differences, not two: + +1. kbc2d does not collide inside the body (those populations are fictitious; γ + statistics are then gathered strictly over fluid). The Python solver *does* + collide there — `kbc_collide` takes no fluid mask. +2. kbc2d's restriction additionally skips coarse nodes whose *fine source* node lies + inside the body. The Python restriction already gates on the **coarse + destination** mask. +3. kbc2d's CPU backend is f64 and its GPU backend f32; the Python side is fp32 by + default, f64 only via `AMR_FP64`.