CLAUDE.md: сверка всех разделов с кодом, исправлены расхождения
Сплошная проверка файла против дерева проекта: каждое фактическое утверждение сверено с исходниками, CMake, манифестами Cargo и git. Ветки и заголовок: - research не имеет неотправленных коммитов (все пять веток совпадают с CFDManager), main/dev стоят на коммит дальше начального импорта; - research добавляет ещё и docs/origins/Grad's_aproximation.pdf; - pipeline_cache.bin и imgui.ini лежат не в корне, а рядом с исполняемым файлом; добавлены прочие незаметные остатки; - сказано, что вне research в docs/theory/2d_solver/ остаются только артефакты, и как читать решатель через git без переключения ветки; - build/ в рабочем дереве собран из чужого каталога и содержит 262 МБ nlohmann_json, который проект не объявляет. C++-редактор: - панель называется Mesh, а не Mesh load; - версия Vulkan SDK и версии компиляторов нигде не проверяются; - 775 МБ зависимостей — измерение по чужому build/, на деле ~400 МБ; - VK_NO_PROTOTYPES задан на двух целях, а не на проекте целиком; - ASan без UBSan на MSVC; уточнён список предупреждений; - карта целей дополнена опущенными библиотеками и видимостью связей; - RecreateSwapDependent пересоздаёт ещё и буфер глубины, первым; - дефект 1: цикл существует на уровне исходников, граф CMake ацикличен — именно поэтому CMake и молчит; - дефект 2: ожидание простоя происходит до vkBeginCommandBuffer, а не посреди записи команд; - мёртвый код Buffer/Image — 362 строки, а не ~200; - glob шейдеров ловит только *.vert и *.frag, суффикс стадии остаётся; - из категорий журнала используется лишь MeshIO, рантайм-настройки не вызываются ниоткуда, весь vk/ пишет через сырой spdlog. Rust-порт: - объявленный минимум 1.82 недостижим: залоченный egui 0.36.1 требует 1.95, на 1.82 обрывается резолвинг; - перечни зависимостей крейтов приведены полностью; - строк кода 2976, а не 2962 (устарела строка mesh на 17 строк); - утверждение «ни одно измерение не отличается на порядок» неверно: 52.3/4.4 = 11.9x; - рост +22 % на 56 % объясняется заменой vk-bootstrap, а не «почти весь»; - assets/ не копируется, откат на текущий каталог сохранён; - перечислены ошибки самого docs/rust_vs_cpp.md. kbc2d: - GREL — доля от полной неравновесной нормы, den > GREL*nrm; - 0.007 % по Cd — это сходимость бэкендов для каждой модели стенки, а не согласие четырёх моделей между собой; - компенсированное суммирование только в статистике: ядро сил считает обычным f32; - раскладок привязок две и восемнадцать, switch — признак расщеплённой; - заголовки --help русские; - гифку пишут 59 прогонов из 115; ≈87 GPU-часов плюс ≈3 CPU-часа; - тег образа 1.2.0 не отражён ни в Cargo.toml, ни в Dockerfile. Python-прототип: - в каталоге четыре поколения решателя, а не одно; добавлена таблица, какое живое, какое опорное, какое мёртвое; - «только CuPy без отката» верно не для всего дерева; - ноутбук состоит из 20 markdown-ячеек без единой ячейки кода, ссылки относительные, четыре из них битые; - отличий от kbc2d три, а не два, и второе сформулировано наоборот; - README solver_2x_sdf описывает несуществующий переключатель collision и сам объявляет свои числа устаревшими. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -49,18 +49,40 @@ single branch contains all of it.** Before assuming a directory exists, check
|
|||||||
|
|
||||||
| branch | what it adds | trees present |
|
| branch | what it adds | trees present |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `main`, `dev` | baseline, nothing beyond the initial import (both at the same commit) | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) |
|
| `main`, `dev` | baseline; both at the same commit, one past the initial import — and that commit only untracks `CLAUDE.md` | `src/`, `shaders/`, `tests/`, `docs/theory/` (Python) |
|
||||||
| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images. Has unpushed commits. | `+ docs/theory/2d_solver/` |
|
| `research` | the CFD work — **`kbc2d`**, a Rust LBM solver, its validation campaign and Docker images | `+ docs/theory/2d_solver/`, `+ docs/origins/Grad's_aproximation.pdf` |
|
||||||
| `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` |
|
| `rust` | a **Rust port of the C++ editor** plus a measured C++/Rust comparison | `+ rust/`, `docs/rust_vs_cpp.md` |
|
||||||
|
|
||||||
|
All five branches (including `notes`) are fully pushed to the `CFDManager` remote;
|
||||||
|
verify with `git ls-remote CFDManager` rather than trusting a stale tracking ref.
|
||||||
|
|
||||||
Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD
|
Nothing is shared between the two Rust efforts: `docs/theory/2d_solver/` (a CFD
|
||||||
solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are
|
solver, branch `research`) and `rust/` (a port of the 3D editor, branch `rust`) are
|
||||||
unrelated projects that merely both happen to be Rust. Neither is part of the CMake
|
unrelated projects that merely both happen to be Rust. Neither is part of the CMake
|
||||||
build.
|
build — `CMakeLists.txt` never mentions either tree.
|
||||||
|
|
||||||
Untracked leftovers you will see in the working tree regardless of branch:
|
**A checkout away from `research` does not remove `docs/theory/2d_solver/`** — the
|
||||||
`build/`, `rust/target/`, `docs/theory/2d_solver/out/` (campaign results —
|
tracked sources vanish but the untracked artefacts stay, so on `rust` that directory
|
||||||
**the user's data, do not clean**), `pipeline_cache.bin`, `imgui.ini`.
|
contains only `out/`, `target/`, `wake.gif` and `demo_wake.gif` and looks like a
|
||||||
|
mutilated copy of the solver. To read the solver from another branch, go through
|
||||||
|
git rather than the filesystem:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
git show research:docs/theory/2d_solver/src/math.rs
|
||||||
|
git archive research docs/theory/2d_solver | tar -x -C /tmp/kbc2d # to grep it
|
||||||
|
```
|
||||||
|
|
||||||
|
Untracked leftovers you will see regardless of branch: `build/`, `rust/target/`,
|
||||||
|
`docs/theory/2d_solver/target/`, the `__pycache__` dirs under `docs/theory/`, and
|
||||||
|
`docs/theory/2d_solver/out/` plus the two ~9 MB GIFs beside it (campaign results —
|
||||||
|
**the user's data, do not clean**). `pipeline_cache.bin` and `imgui.ini` are written
|
||||||
|
next to whichever executable ran, i.e. in `build/<preset>/src/app/` and
|
||||||
|
`rust/pipeline_cache.bin` — not at the repo root.
|
||||||
|
|
||||||
|
**`build/` in this working tree is foreign and stale.** Its
|
||||||
|
`CMakeCache.txt` records `CMAKE_HOME_DIRECTORY=C:/Users/ivan/UAS/SimV4(Not_KBC)/SimVulcan`,
|
||||||
|
and `_deps/` carries a 262 MB `nlohmann_json` clone that this project never declares.
|
||||||
|
Do not measure anything from it without re-configuring first.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -69,15 +91,25 @@ Untracked leftovers you will see in the working tree regardless of branch:
|
|||||||
A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY,
|
A minimal **Vulkan 1.3 / C++20** 3D model editor: three reference grid planes (XY,
|
||||||
XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj`
|
XZ, YZ) through the origin plus coloured X/Y/Z axes, an orbit camera, and `.obj`
|
||||||
loading with selectable display modes (solid, wireframe, solid + wireframe). It
|
loading with selectable display modes (solid, wireframe, solid + wireframe). It
|
||||||
runs **no simulation** — it is a rendering skeleton with an ImGui interface (a
|
runs **no simulation** — no compute pipeline or `.comp` shader exists anywhere, and
|
||||||
Viewport control panel and a Mesh load panel). Unchanged since the initial commit
|
`Renderer.cpp` disables async compute explicitly. It is a rendering skeleton with
|
||||||
on every branch.
|
two ImGui windows, titled **`Viewport`** and **`Mesh`** (the class behind the second
|
||||||
|
is `MeshLoadPanel`, which is why the docs keep calling it the "mesh load panel").
|
||||||
|
Unchanged since the initial commit on every branch — `git log --all -- src` returns
|
||||||
|
exactly one commit.
|
||||||
|
|
||||||
## Build
|
## Build
|
||||||
|
|
||||||
Prerequisites: **Vulkan SDK 1.3.290+** (provides `glslangValidator` for offline
|
Prerequisites: **Vulkan SDK** (for `glslangValidator`), **CMake 3.26+**, **Ninja**,
|
||||||
shader compilation), **CMake 3.26+**, **Ninja**, a C++20 compiler (MSVC 19.36+,
|
a C++20 compiler. Only CMake 3.26 and C++20 are actually enforced
|
||||||
gcc 11+, clang 14+). On macOS, Vulkan is via MoltenVK (Apple Silicon only).
|
(`CMakeLists.txt:1`, `:9-11`). `cmake/VulkanSetup.cmake:3` is
|
||||||
|
`find_package(Vulkan REQUIRED COMPONENTS glslangValidator)` with **no version
|
||||||
|
argument**, so the "SDK 1.3.290+" in `README.md` is a comment, not a check; the
|
||||||
|
de-facto floor comes from the pinned volk `1.3.295` and vk-bootstrap `v1.3.295`.
|
||||||
|
The compiler minimums (MSVC 19.36+, gcc 11+, clang 14+) are likewise README prose,
|
||||||
|
enforced nowhere. On macOS the only hard constraint is
|
||||||
|
`CMAKE_OSX_ARCHITECTURES: arm64` in the preset; MoltenVK is mentioned in a
|
||||||
|
`message(STATUS)`.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
cmake --preset windows-msvc-release
|
cmake --preset windows-msvc-release
|
||||||
@@ -86,42 +118,67 @@ cmake --build --preset windows-msvc-release
|
|||||||
|
|
||||||
Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`,
|
Presets: `windows-msvc-debug`, `windows-msvc-release`, `linux-gcc-release`,
|
||||||
`linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under
|
`linux-clang-release`, `macos-arm64-release` (all Ninja, one dir per preset under
|
||||||
`build/<presetName>/`). The first configure fetches dependencies via FetchContent
|
`build/<presetName>/`). None of them pins a compiler — they inherit whatever the
|
||||||
(GLFW, GLM, volk, vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog,
|
environment provides, so `README.md`'s "MSVC / clang-cl" is aspirational.
|
||||||
tinyobjloader, Catch2) and needs network access — it clones nine repositories with
|
|
||||||
full history (**775 MB**, ~15 min on a slow link) and **has no shared cache**, so
|
|
||||||
every new build directory downloads all of it again. `VK_NO_PROTOTYPES` is set
|
|
||||||
project-wide; Vulkan entry points load through **volk**.
|
|
||||||
|
|
||||||
Warnings come from `simv_set_warnings` (`/W4 /permissive-`, or
|
The first configure fetches nine repositories via FetchContent (GLFW, GLM, volk,
|
||||||
`-Wall -Wextra -Wpedantic -Wshadow -Wold-style-cast …`); they are **not** errors.
|
vk-bootstrap, VulkanMemoryAllocator, Dear ImGui, spdlog, tinyobjloader, Catch2) and
|
||||||
`cmake/Sanitizers.cmake` defines `simv_enable_sanitizers` (Debug-only ASan/UBSan)
|
needs network access. There is no `GIT_SHALLOW`, so each is a clone with full
|
||||||
but no target currently calls it — wire it in manually when chasing memory bugs.
|
history, and **no shared cache** is configured (the only `FETCHCONTENT_*` variable
|
||||||
|
set is `FETCHCONTENT_QUIET OFF`) — every new build directory downloads all of it
|
||||||
|
again. Budget **~400 MB of sources / ~510 MB of `_deps` once built**, not the
|
||||||
|
775 MB quoted in `docs/rust_vs_cpp.md`: that figure was measured on the foreign
|
||||||
|
`build/` tree described above and includes 262 MB of `nlohmann_json` that nothing
|
||||||
|
here declares.
|
||||||
|
|
||||||
|
`VK_NO_PROTOTYPES` is **not** project-wide — there is no `add_compile_definitions`
|
||||||
|
at project scope. It is set `PUBLIC` on exactly two targets, `imgui`
|
||||||
|
(`CMakeLists.txt:114`) and `simv_vk` (`src/vk/CMakeLists.txt:29`). Because
|
||||||
|
`simv_core` links `simv_vk` **PRIVATE**, the definition never reaches `simv_mesh`,
|
||||||
|
`simv_editor` or `simv_tests`. Vulkan entry points load through **volk**.
|
||||||
|
|
||||||
|
Warnings come from `simv_set_warnings` (`cmake/CompilerWarnings.cmake`), called by
|
||||||
|
all six targets: `/W4 /permissive- /Zc:preprocessor /Zc:__cplusplus /wd4100` plus
|
||||||
|
the defines `_CRT_SECURE_NO_WARNINGS`, `NOMINMAX`, `WIN32_LEAN_AND_MEAN` on MSVC;
|
||||||
|
`-Wall -Wextra -Wpedantic -Wno-unused-parameter -Wshadow -Wnon-virtual-dtor
|
||||||
|
-Wold-style-cast -Wcast-align -Wunused -Woverloaded-virtual` elsewhere. They are
|
||||||
|
**not** errors — no `/WX` or `-Werror` anywhere. `cmake/Sanitizers.cmake` defines
|
||||||
|
`simv_enable_sanitizers` (Debug-only; ASan **only** on MSVC, ASan+UBSan elsewhere)
|
||||||
|
but no target calls it — wire it in manually when chasing memory bugs.
|
||||||
|
|
||||||
## Running
|
## Running
|
||||||
|
|
||||||
The executable resolves SPIR-V relative to the working directory (`FindSpvPath`
|
The executable resolves SPIR-V relative to the working directory (`FindSpvPath`,
|
||||||
probes `spirv/<rel>` and `current_path()/spirv/<rel>`). The `SimVulcan` POST_BUILD
|
`src/vk/Shader.cpp:24-33`, probes `spirv/<rel>` then `current_path()/spirv/<rel>` —
|
||||||
step copies the compiled `spirv/` tree and `assets/` next to the executable, so
|
the second probe is a no-op, since the first is already CWD-relative). Two
|
||||||
**run from the executable's own directory** (`build/<preset>/src/app/`). `main.cpp`
|
`SimVulcan` POST_BUILD steps copy `assets/` and the compiled `spirv/` tree next to
|
||||||
also probes several `../` ancestors for `assets/meshes`. Meshes load at runtime
|
the executable, so **run from the executable's own directory**
|
||||||
through the ImGui Mesh panel.
|
(`build/<preset>/src/app/`). `main.cpp:40-50` additionally probes four `../`
|
||||||
|
ancestors for `assets/meshes`. Meshes load at runtime through the `Mesh` panel.
|
||||||
|
|
||||||
Running writes two files into the CWD: `pipeline_cache.bin` (serialised
|
Running writes two files into the CWD: `pipeline_cache.bin` (serialised
|
||||||
`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini`. Both are
|
`VkPipelineCache`, reloaded on the next start) and ImGui's `imgui.ini` (no
|
||||||
disposable — delete them if pipeline creation or the panel layout misbehaves.
|
`IniFilename` is ever set, so ImGui's default applies). Both are disposable —
|
||||||
|
delete them if pipeline creation or the panel layout misbehaves.
|
||||||
|
|
||||||
`ContextOptions::enableValidation` / `enableDebugUtils` default to **true** in
|
`ContextOptions::enableValidation` / `enableDebugUtils` default to **true**
|
||||||
every build config, so the Khronos validation layer is requested even in Release;
|
(`src/vk/Context.h:15-16`) and nothing overrides them, so the Khronos validation
|
||||||
messages (error + warning severity) go through spdlog.
|
layer is requested even in Release. The severity mask is error + warning only
|
||||||
|
(`Context.cpp:56-58`), which makes the `INFO` branch of the callback dead code.
|
||||||
|
Those messages go to **raw `spdlog::`**, i.e. spdlog's default logger — not the
|
||||||
|
`"simv"` logger that `core::Logger` creates.
|
||||||
|
|
||||||
`run.bat` at the repo root is **stale**: it launches
|
`run.bat` at the repo root is **stale**: it launches
|
||||||
`build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never
|
`build/vs2022/src/app/Release/SimVulcan.exe`, a path the Ninja presets never
|
||||||
produce. Do not point users at it without fixing the path first.
|
produce — while its own error message tells the user to run those presets. Do not
|
||||||
|
point users at it without fixing the path first.
|
||||||
|
|
||||||
## Tests
|
## Tests
|
||||||
|
|
||||||
Catch2 unit tests, pure CPU/math (mesh bounds + welding). No GPU required.
|
Catch2 unit tests, mesh bounds + welding. No GPU is touched at runtime, but the
|
||||||
|
binary is **not** free of Vulkan: static-lib propagation through
|
||||||
|
`simv_mesh → simv_core → (PRIVATE) simv_vk` makes `simv_tests.exe` link
|
||||||
|
`simv_vk`, `vk-bootstrap`, `imgui`, `volk` and `glfw3`.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
cmake --build --preset windows-msvc-debug --target simv_tests
|
cmake --build --preset windows-msvc-debug --target simv_tests
|
||||||
@@ -131,16 +188,14 @@ ctest --preset windows-msvc-debug
|
|||||||
`windows-msvc-debug` is the only preset with a `testPreset` — for the other
|
`windows-msvc-debug` is the only preset with a `testPreset` — for the other
|
||||||
configs, invoke `ctest` in `build/<preset>/` directly.
|
configs, invoke `ctest` in `build/<preset>/` directly.
|
||||||
|
|
||||||
Single test / subset — either through CTest (each `TEST_CASE` is registered
|
There are exactly three `TEST_CASE`s, all in `tests/unit/test_mesh.cpp`:
|
||||||
individually by `catch_discover_tests`):
|
`Mesh::RecalculateBounds finds AABB` `[mesh]`, `WeldVertices collapses
|
||||||
|
near-duplicate vertices` and `WeldVertices drops degenerate triangles`
|
||||||
```sh
|
(both `[mesh][decimator]`). Each is registered individually by
|
||||||
ctest --preset windows-msvc-debug -R "WeldVertices" -V
|
`catch_discover_tests`, so either filter works:
|
||||||
```
|
|
||||||
|
|
||||||
or by running the binary with a Catch2 name or tag filter:
|
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
|
ctest --preset windows-msvc-debug -R "WeldVertices" -V # matches the last two
|
||||||
./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]"
|
./build/windows-msvc-debug/tests/simv_tests.exe "[decimator]"
|
||||||
./build/windows-msvc-debug/tests/simv_tests.exe --list-tests
|
./build/windows-msvc-debug/tests/simv_tests.exe --list-tests
|
||||||
```
|
```
|
||||||
@@ -149,15 +204,23 @@ or by running the binary with a Catch2 name or tag filter:
|
|||||||
|
|
||||||
### Library / target map
|
### Library / target map
|
||||||
|
|
||||||
|
Arrows below show the **project** edges; each target also links third-party
|
||||||
|
libraries, and the visibility keywords matter more than they look.
|
||||||
|
|
||||||
```
|
```
|
||||||
SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor
|
SimVulcan (exe) → simv_core, simv_vk, simv_mesh, simv_editor (all PRIVATE)
|
||||||
simv_core → simv_vk (App owns Window + vk::Renderer; no Vulkan calls)
|
+ glm, spdlog; add_dependencies(… simv_shaders)
|
||||||
simv_editor → simv_core, simv_mesh (Camera, input, ImGui panels; no Vulkan)
|
simv_core → PUBLIC spdlog · PRIVATE simv_vk, glfw (App owns Window + Renderer)
|
||||||
simv_mesh → simv_core (CPU mesh only; no Vulkan)
|
simv_editor → PUBLIC simv_core, simv_mesh, imgui, glm (no Vulkan)
|
||||||
simv_vk → third-party (volk, vk-bootstrap, VMA, GLFW, GLM, spdlog, imgui)
|
simv_mesh → PUBLIC simv_core, glm, tinyobjloader (no Vulkan)
|
||||||
simv_shaders → glslangValidator (GLSL → SPIR-V)
|
simv_vk → PUBLIC Vulkan::Headers, volk, vk-bootstrap, VMA, glfw, glm, spdlog
|
||||||
|
PRIVATE imgui
|
||||||
|
simv_shaders → glslangValidator (GLSL → SPIR-V), an ALL custom target
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The `simv_core → simv_vk` edge being **PRIVATE** is what keeps `VK_NO_PROTOTYPES`
|
||||||
|
(and volk) from reaching the rest of the tree — see the Build section.
|
||||||
|
|
||||||
Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`,
|
Namespaces follow directories: `simv::core`, `simv::vk`, `simv::mesh`,
|
||||||
`simv::editor`. Each library exports `src/` as its include root, so includes are
|
`simv::editor`. Each library exports `src/` as its include root, so includes are
|
||||||
written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
|
written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
|
||||||
@@ -165,66 +228,77 @@ written module-qualified (`#include "mesh/Mesh.h"`, `#include "vk/Renderer.h"`).
|
|||||||
### Frame loop
|
### Frame loop
|
||||||
|
|
||||||
`main.cpp` creates `core::App`, which owns the `core::Window` and a
|
`main.cpp` creates `core::App`, which owns the `core::Window` and a
|
||||||
`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` calls the UI
|
`vk::Renderer`, then runs the loop. Each frame `Renderer::DrawFrame` acquires a
|
||||||
callback (between ImGui NewFrame/Render), then records the scene: grid + mesh into
|
swapchain image, calls the UI callback (between ImGui NewFrame/Render), *then*
|
||||||
one dynamic-rendering pass with a depth attachment, followed by ImGui, then
|
begins the command buffer and records the scene: grid + mesh into one
|
||||||
presents (2 frames in flight, sync2 submits).
|
dynamic-rendering pass with a depth attachment, followed by ImGui, then presents
|
||||||
|
(2 frames in flight, `vkQueueSubmit2` / `VkImageMemoryBarrier2`).
|
||||||
|
|
||||||
|
The acquire-before-callback ordering is the root of defect 2 below — keep it in
|
||||||
|
mind before moving work into the callback.
|
||||||
|
|
||||||
`main.cpp` wires the editor via the UI callback: it draws the panels, applies
|
`main.cpp` wires the editor via the UI callback: it draws the panels, applies
|
||||||
mouse input to the `editor::Camera`, and pushes the resulting state into the
|
mouse input to the `editor::Camera`, and pushes the resulting state into the
|
||||||
renderer (`SetViewProj`, `SetRenderMode`, `SetGridVisible`). Mesh loads go through
|
renderer (`SetRenderMode`, `SetGridVisible`, `SetViewProj`). Mesh loads go through
|
||||||
`MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`.
|
`MeshLoadPanel`'s callback → `Renderer::SetMeshCpu`.
|
||||||
|
|
||||||
### Mesh load path
|
### Mesh load path
|
||||||
|
|
||||||
`MeshLoadPanel` lists `*.obj` in the mesh directory and kicks off
|
`MeshLoadPanel` lists `*.obj` in the mesh directory (extension match is
|
||||||
`mesh::LoadObjAsync` (worker thread, `std::future`). The future is **drained on
|
**case-insensitive**, so `.OBJ` shows up too) and kicks off `mesh::LoadObjAsync`
|
||||||
the main thread** at the top of `MeshLoadPanel::Draw`, so the `OnLoaded` callback
|
(`std::async(std::launch::async, …)`, `std::future`). The future is **drained on
|
||||||
— and therefore the GPU upload — always runs on the render thread. Loader
|
the main thread** at the top of `MeshLoadPanel::Draw`, before the first
|
||||||
exceptions surface as the panel's status string.
|
`ImGui::Begin`, so the `OnLoaded` callback — and therefore the GPU upload — always
|
||||||
|
runs on the render thread. Loader exceptions surface as the panel's status string.
|
||||||
|
|
||||||
`main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs
|
`main.cpp`'s `OnLoaded` welds the mesh (`WeldVertices`, tolerance `1e-4`), logs
|
||||||
the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius.
|
the counts, uploads via `SetMeshCpu`, and reframes the camera to the bbox radius.
|
||||||
Despite the file name, `mesh/MeshDecimator.h` implements **only** spatial-hash
|
Despite the file name, `mesh/MeshDecimator.h` declares **one** function and
|
||||||
vertex welding (collapse near-duplicates, drop degenerate triangles, recompute
|
implements **only** spatial-hash vertex welding (round-to-nearest cell hash,
|
||||||
bounds) — there is no LOD/decimation.
|
collapse near-duplicates, drop degenerate triangles, recompute bounds) — there is
|
||||||
|
no LOD/decimation.
|
||||||
|
|
||||||
### Non-obvious invariants (read before editing)
|
### Non-obvious invariants (read before editing)
|
||||||
|
|
||||||
- **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no
|
- **All Vulkan lives in `src/vk/`.** `core/`, `mesh/`, `editor/` and `app/` make no
|
||||||
Vulkan API calls. `vk::Renderer`'s public header is deliberately Vulkan-free
|
Vulkan API calls — a grep for `volk|vulkan|Vk[A-Z]` across them returns zero hits.
|
||||||
(pImpl + glm/`RenderMode` only) so `App` can own it without pulling in volk. Keep
|
`vk::Renderer`'s public header is deliberately Vulkan-free (pImpl + glm/`RenderMode`
|
||||||
it that way — do not leak `Vk*` types into the public interfaces of those modules.
|
only) so `App` can own it without pulling in volk. Keep it that way — do not leak
|
||||||
Note this is **convention only**: every header is visible, so nothing stops a
|
`Vk*` types into the public interfaces of those modules. Note this is **convention
|
||||||
|
only**: every library exports `src/` publicly, so nothing stops a
|
||||||
`#include <volk.h>` in `mesh/`; the compiler will not catch it.
|
`#include <volk.h>` in `mesh/`; the compiler will not catch it.
|
||||||
- **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one
|
- **Single render pass with depth.** Scene (grid + mesh) and ImGui draw into one
|
||||||
`vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth
|
`vkCmdBeginRendering` pass that has both a colour and a `D32_SFLOAT` depth
|
||||||
attachment. The ImGui backend is initialised with `depthAttachmentFormat` set so
|
attachment (depth `storeOp` is `DONT_CARE`). The ImGui backend is initialised with
|
||||||
its pipeline matches the pass; the depth buffer is recreated with the swapchain.
|
`depthAttachmentFormat` set so its pipeline matches the pass.
|
||||||
- **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and
|
- **Wireframe needs `fillModeNonSolid`.** `MeshRenderer` builds a fill pipeline and
|
||||||
a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small depth bias so
|
a `VK_POLYGON_MODE_LINE` pipeline; the line pipeline uses a small *static* depth
|
||||||
the overlay sits on top of the fill. The device feature is requested in `Context`.
|
bias (`constantFactor` and `slopeFactor` both `-1.0`) so the overlay sits on top
|
||||||
Culling is off (`VK_CULL_MODE_NONE`) — loaded models may have mixed winding.
|
of the fill. The device feature is requested in `Context`. Culling is off
|
||||||
|
(`VK_CULL_MODE_NONE`) — loaded models may have mixed winding.
|
||||||
- **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance)
|
- **Camera is fixed on the origin.** `editor::Camera` orbits (yaw/pitch/distance)
|
||||||
the world origin; loaded models are recentred there via a translate-only model
|
the world origin; loaded models are recentred there via a translate-only model
|
||||||
matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip,
|
matrix. Projection uses Vulkan clip space (`GLM_FORCE_DEPTH_ZERO_TO_ONE` + Y flip,
|
||||||
isolated to `Camera.cpp`).
|
isolated to `Camera.cpp` — the macro appears nowhere else in the tree).
|
||||||
- **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on
|
- **Buffers.** `vk::GpuMesh` owns the model's vertex/index buffers (staged upload on
|
||||||
the transfer queue). `GridRenderer` builds a static host-visible line buffer once.
|
the dedicated transfer queue). `GridRenderer` builds a static host-visible mapped
|
||||||
Both scene renderers use only a push constant — no descriptor sets.
|
line buffer once in its constructor. Both scene renderers create pipeline layouts
|
||||||
|
with one push-constant range and **no** `pSetLayouts` — no descriptor sets at all.
|
||||||
- **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous
|
- **Push-constant layout is a cross-file contract.** `MeshRenderer.cpp`'s anonymous
|
||||||
`MeshPC { mat4 mvp; vec4 color; }` must stay byte-identical to the
|
`MeshPC { mat4 mvp; vec4 color; }` (range `VERTEX|FRAGMENT`) must stay
|
||||||
`push_constant` block in `mesh.vert`/`mesh.frag`; `color.a` is a *flag*, not
|
byte-identical to the `push_constant` block in `mesh.vert`/`mesh.frag`; `color.a`
|
||||||
alpha (1 = flat-shaded, 0 = constant colour for wireframe). `GridRenderer` pushes
|
is a *flag*, not alpha — `mesh.frag` does `mix(base, base*shade, pc.color.a)`, so
|
||||||
a bare `mat4` (vertex stage only). Change either side and you must change both.
|
1 = flat-shaded and 0 = constant colour for wireframe. `GridRenderer` pushes a
|
||||||
- **Swapchain recreation rebuilds sync objects.** `RecreateSwapDependent` recreates
|
bare `mat4` (vertex stage only). Change either side and you must change both.
|
||||||
the per-frame `imageAvailable` semaphores (a failed acquire can leave one
|
- **Swapchain recreation rebuilds depth *and* sync objects.** `RecreateSwapDependent`
|
||||||
signalled) *and* the per-swapchain-image `renderFinished` semaphores (the image
|
runs `WaitIdle` → `swap->Recreate()` → **`CreateDepth`** → per-frame
|
||||||
count may change), then re-ensures both pipelines. Keep that ordering if you
|
`imageAvailable` semaphores (a failed acquire can leave one signalled) → per-
|
||||||
touch resize handling.
|
swapchain-image `renderFinished` semaphores (the image count may change) → re-ensure
|
||||||
|
both pipelines. Keep that ordering if you touch resize handling.
|
||||||
- **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle`
|
- **Mesh upload stalls the device.** `SetMeshCpu`/`ClearMesh` call `WaitIdle`
|
||||||
before touching `GpuMesh` — acceptable because loads are rare; do not copy that
|
before touching `GpuMesh` — acceptable because loads are rare; do not copy that
|
||||||
pattern into per-frame paths.
|
pattern into per-frame paths. `SetMeshCpu` early-returns on an empty mesh *before*
|
||||||
|
the wait.
|
||||||
|
|
||||||
### Known defects (found by the Rust port, all still present here)
|
### Known defects (found by the Rust port, all still present here)
|
||||||
|
|
||||||
@@ -232,46 +306,71 @@ Porting the editor surfaced four real bugs that nobody was looking for. They are
|
|||||||
unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md`
|
unfixed in the C++ tree on every branch; full write-up in `docs/rust_vs_cpp.md`
|
||||||
(branch `rust`).
|
(branch `rust`).
|
||||||
|
|
||||||
1. **The CMake target graph has a cycle.** `core::App` owns `vk::Renderer`, and
|
1. **A source-level dependency cycle that the CMake graph hides.** `core::App` owns
|
||||||
`vk::Renderer` takes a `core::Window&`. It links only because `simv_vk` does
|
`vk::Renderer`, and `vk::Renderer` takes a `core::Window&` — so the *sources*
|
||||||
**not** link `simv_core` — it sees the headers through
|
are mutually dependent. The **CMake graph is acyclic**, which is exactly why
|
||||||
`target_include_directories(simv_vk PUBLIC ..)` and everything resolves at
|
nothing complains: `simv_vk` links no project target at all, seeing
|
||||||
executable link time. CMake never complains.
|
`core/Window.h` through `target_include_directories(simv_vk PUBLIC ..)` and
|
||||||
2. **Mesh upload runs mid-frame.** `MeshLoadPanel::Draw` calls `OnLoaded` from
|
getting `Window`'s symbols at executable link time. Cargo rejects the same shape
|
||||||
inside `Renderer::DrawFrame`, i.e. *after* `vkAcquireNextImageKHR`; the handler
|
outright, which is what surfaced it.
|
||||||
calls `SetMeshCpu` → `vkDeviceWaitIdle`. Legal, but it waits for device idle in
|
2. **Mesh upload runs inside the frame.** `MeshLoadPanel::Draw` calls `OnLoaded`
|
||||||
the middle of command recording.
|
from the UI callback, which `Renderer::DrawFrame` invokes *after*
|
||||||
|
`vkAcquireNextImageKHR`; the handler calls `SetMeshCpu` → `vkDeviceWaitIdle`.
|
||||||
|
Legal, and command recording has not started yet (`vkBeginCommandBuffer` comes
|
||||||
|
later) — but it waits for device idle while holding an acquired swapchain image.
|
||||||
3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and
|
3. **Two sources of truth for "is there a mesh".** `GpuMesh::IsValid()` and
|
||||||
`Renderer`'s separate `hasMesh` flag must be kept consistent by hand.
|
`Renderer`'s separate `hasMesh` flag must be kept consistent by hand;
|
||||||
|
`MeshRenderer::Draw` then re-checks `IsValid()` a third time.
|
||||||
4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category
|
4. **`Mesh ready` is never printed.** `ObjLoader::LoadObj` logs to category
|
||||||
`MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category
|
`MeshIO`, and `main.cpp`'s handler logs `Mesh ready: …` to the same category
|
||||||
immediately after; the category throttles at 0.5 s and both are `Info`, so the
|
immediately after; the category throttles at 0.5 s and both are `Info`, so the
|
||||||
second message is always dropped. Post-weld counts have therefore never appeared
|
second message is always dropped (only Warn/Error/Critical bypass the throttle).
|
||||||
in any log.
|
Post-weld counts have therefore never appeared in any log.
|
||||||
|
|
||||||
Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** (~200 lines).
|
Also dead weight: **`vk::Buffer` and `vk::Image` are used by nobody** — 362 lines
|
||||||
`GridRenderer`, `GpuMesh` and the depth attachment all call VMA directly.
|
across four files, still compiled into `simv_vk`. `GridRenderer`, `GpuMesh` and the
|
||||||
|
depth attachment all call VMA directly. `Image::Desc` even defaults to
|
||||||
|
`VK_IMAGE_TYPE_3D` and an `R32G32B32A32_SFLOAT` storage image, a leftover from some
|
||||||
|
compute/CFD design. Note `README.md:53-54` still lists both as part of the working
|
||||||
|
Vulkan stack — that line is wrong.
|
||||||
|
|
||||||
### Shaders
|
### Shaders
|
||||||
|
|
||||||
GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders`
|
GLSL under `shaders/editor/` (`mesh.{vert,frag}`, `grid.{vert,frag}`). `simv_shaders`
|
||||||
compiles each to `build/<preset>/spirv/editor/<name>.spv` targeting `vulkan1.3`,
|
compiles each to `build/<preset>/spirv/editor/<name>.<stage>.spv` — the stage suffix
|
||||||
with `shaders/` as the `-I` root (so `#include "common/foo.glsl"` would resolve).
|
is **kept**, so the real artefacts are `mesh.vert.spv`, `mesh.frag.spv` and so on —
|
||||||
Adding a shader under that globbed dir is picked up automatically
|
targeting `vulkan1.3`, with `shaders/` as the `-I` root (so `#include
|
||||||
(`CONFIGURE_DEPENDS`). `mesh.frag` reconstructs a flat normal from screen-space
|
"common/foo.glsl"` would resolve). The glob is `CONFIGURE_DEPENDS`, so a new shader
|
||||||
derivatives, so the vertex stream carries only positions (tight `float3`, one
|
is picked up automatically, but it matches **only `editor/*.vert` and
|
||||||
binding, one attribute). **The Rust port on branch `rust` compiles this same
|
`editor/*.frag`**; a `.comp` or `.geom` dropped there is silently ignored.
|
||||||
directory** — a shader change affects both versions.
|
|
||||||
|
`mesh.frag` reconstructs a flat normal from screen-space derivatives, so the mesh
|
||||||
|
vertex stream carries only positions (tight `float3`, one binding, one attribute).
|
||||||
|
The *grid* pipeline is different — two attributes, position + colour.
|
||||||
|
**The Rust port on branch `rust` compiles this same directory** (`build.rs` reaches
|
||||||
|
`../../../shaders`) with the same `glslangValidator` flags — a shader change affects
|
||||||
|
both versions.
|
||||||
|
|
||||||
### Logging
|
### Logging
|
||||||
|
|
||||||
`simv::core::Logger` (spdlog-backed, singleton) with category-based throttling.
|
`simv::core::Logger` (spdlog-backed, singleton) with category-based throttling.
|
||||||
Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`,
|
Log via `LogFmt(LogCategory, LogLevel, fmt, args...)`. Categories: `Core`,
|
||||||
`Vulkan`, `MeshIO`, `UI`, `Test`. Per-category level, throttle interval and
|
`Vulkan`, `MeshIO`, `UI`, `Test` (printed as `core`, `vk`, `mesh`, `ui`, `test`).
|
||||||
on/off are settable at runtime (`SetMinLevel` / `SetThrottle` / `SetEnabled`).
|
|
||||||
Some low-level Vulkan code still calls `spdlog::` directly — which is the only
|
Two things the API surface does not tell you:
|
||||||
reason the GPU name survives startup (it bypasses category throttling; see
|
|
||||||
defect 4 above).
|
- **Only `MeshIO` is ever used.** The entire tree contains three `LogFmt` call
|
||||||
|
sites — two in `ObjLoader.cpp`, one in `main.cpp`. `Core`, `Vulkan`, `UI` and
|
||||||
|
`Test` have zero.
|
||||||
|
- **The runtime knobs are never turned.** `SetMinLevel` / `SetThrottle` /
|
||||||
|
`SetEnabled` exist but have no callers, and no flag, env var or UI control is
|
||||||
|
wired to them. They are settable programmatically, not configurable.
|
||||||
|
|
||||||
|
And the whole of `vk/` logs through **raw `spdlog::`**, never through `LogFmt` —
|
||||||
|
as do `core/App.cpp` and `main.cpp`'s fatal handler. That bypasses not just the
|
||||||
|
category throttle but the `"simv"` logger entirely (different pattern, unaffected
|
||||||
|
by `Logger::Shutdown()`), which is the only reason the GPU name survives startup;
|
||||||
|
see defect 4 above.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -279,42 +378,67 @@ defect 4 above).
|
|||||||
|
|
||||||
A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions
|
A port of SimVulcan from C++20 to Rust, existing **for comparison**: both versions
|
||||||
sit side by side and build independently. Same Vulkan 1.3, dynamic rendering,
|
sit side by side and build independently. Same Vulkan 1.3, dynamic rendering,
|
||||||
synchronization2, and the same `shaders/` and `assets/` directories. Needs **Rust
|
synchronization2, and the same `shaders/` directory.
|
||||||
1.82+** (tested on 1.97.1) and the Vulkan SDK for `glslangValidator` only.
|
|
||||||
|
**The declared MSRV is wrong.** `rust/Cargo.toml:13` says `rust-version = "1.82"`
|
||||||
|
(and `rust/README.md` repeats it), but the locked `egui`/`egui-winit`/`epaint`
|
||||||
|
0.36.1 each declare `rust-version = "1.95"`, so Cargo refuses to resolve on 1.82.
|
||||||
|
**The real minimum is 1.95**; the tree is tested on 1.97.1. The Vulkan SDK is needed
|
||||||
|
for `glslangValidator` only.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
cd rust
|
cd rust
|
||||||
cargo build --release
|
cargo build --release
|
||||||
./target/release/SimVulcan # runnable from any directory
|
./target/release/SimVulcan # SimVulcan.exe on Windows; package is simv-app
|
||||||
cargo test # mesh bounds, welding, Cube.obj, camera clip space
|
cargo test # 13 tests + 1 #[ignore]d diagnostic
|
||||||
```
|
```
|
||||||
|
|
||||||
**Work on the code with plain `cargo build`, not `--release`.** The release profile
|
**Work on the code with plain `cargo build`, not `--release`.** The release profile
|
||||||
sets `lto = "thin"`, which makes a one-line edit re-optimise and relink the whole
|
sets `lto = "thin"` with `codegen-units = 1`, which makes a one-line edit
|
||||||
binary: 52 s versus ~5 s in debug.
|
re-optimise and relink the whole binary: 52 s versus ~5 s in debug.
|
||||||
|
|
||||||
|
The tests cover mesh bounds, welding (three cases), `Cube.obj` loading and its error
|
||||||
|
path, camera clip space and zoom/pitch clamping, and logger category mapping. They
|
||||||
|
need no GPU but are **not** pure math: the loader tests read real files from
|
||||||
|
`assets/meshes` via `$CARGO_MANIFEST_DIR`, so they require the repo checkout. The
|
||||||
|
ignored `weld_counts_for_bundled_assets` is the diagnostic that produced the loader
|
||||||
|
parity table below.
|
||||||
|
|
||||||
Five crates mirror the CMake target map, and the "all Vulkan in one place"
|
Five crates mirror the CMake target map, and the "all Vulkan in one place"
|
||||||
invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have
|
invariant is **compiler-enforced** here — `simv-mesh` and `simv-editor` have
|
||||||
neither `ash` nor `simv-vk` among their dependencies:
|
neither `ash` nor `simv-vk` among their dependencies. Third-party deps in full
|
||||||
|
(each crate also depends on the workspace crates shown):
|
||||||
|
|
||||||
```
|
```
|
||||||
crates/simv-core/ logger, window → log, chrono, winit
|
crates/simv-core/ logger, window → log, chrono, winit
|
||||||
crates/simv-mesh/ .obj reading, welding, bounds → glam, tobj
|
crates/simv-mesh/ .obj, welding, bounds → glam, tobj, log, thiserror
|
||||||
crates/simv-vk/ all Vulkan + build.rs (GLSL→SPIR-V) → ash, gpu-allocator,
|
(+ simv-core)
|
||||||
egui-ash-renderer
|
crates/simv-vk/ all Vulkan + build.rs → ash, ash-window, raw-window-handle,
|
||||||
crates/simv-editor/ camera, input, panels → egui
|
(GLSL→SPIR-V) gpu-allocator, egui, egui-winit,
|
||||||
crates/simv-app/ App and entry point → all of the above
|
egui-ash-renderer, winit, glam,
|
||||||
|
bytemuck, log, thiserror
|
||||||
|
(+ simv-core, simv-mesh)
|
||||||
|
crates/simv-editor/ camera, input, panels → egui, glam, log
|
||||||
|
(+ simv-core, simv-mesh)
|
||||||
|
crates/simv-app/ App and entry point → the four crates above
|
||||||
|
+ egui, winit, glam, log, anyhow
|
||||||
```
|
```
|
||||||
|
|
||||||
Structural differences from C++, each with a reason (full list in
|
Structural differences from C++, each with a reason (full list in
|
||||||
`docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the
|
`docs/rust_vs_cpp.md`): `App` lives in the executable crate (Cargo rejects the
|
||||||
target-graph cycle); the UI closure **returns** a `FrameState` instead of calling
|
target-graph cycle); the UI closure **returns** a `FrameState` instead of calling
|
||||||
setters (it is invoked from a renderer method, so it cannot borrow the renderer);
|
setters (it is invoked from a renderer method, so it cannot borrow the renderer) —
|
||||||
the loaded mesh is uploaded *after* the frame, not from mid-recording; `Buffer` and
|
`Renderer`'s public API has no `Set*` methods at all; the loaded mesh is uploaded
|
||||||
`Image` are actually used; no pImpl (private module fields give the same
|
*after* `draw_frame` returns, not from inside it; `Buffer` and `Image` are actually
|
||||||
isolation); winit owns the event loop, so `ShouldClose`/`PollEvents` are gone.
|
used; no pImpl (private module fields give the same isolation); winit owns the
|
||||||
|
event loop, so `ShouldClose`/`PollEvents` are gone.
|
||||||
|
|
||||||
SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime
|
SPIR-V is embedded via `include_bytes!` from `build.rs`, so there is no runtime
|
||||||
shader lookup and no requirement to run from a particular directory.
|
shader lookup. Two runtime path dependencies remain, though, so "runs from any
|
||||||
|
directory" is not quite true: `assets/meshes` is searched from the executable's
|
||||||
|
ancestors **and then the CWD's**, and `pipeline_cache.bin` is written next to the
|
||||||
|
exe. Unlike the C++ build, nothing copies `assets/` — the port finds the repo-root
|
||||||
|
copy by walking up from `rust/target/<profile>/`.
|
||||||
|
|
||||||
The UI toolkit differs **by design**: egui instead of Dear ImGui, because the
|
The UI toolkit differs **by design**: egui instead of Dear ImGui, because the
|
||||||
`imgui` crate is bindings and would compile ~40k lines of C++, defeating the point
|
`imgui` crate is bindings and would compile ~40k lines of C++, defeating the point
|
||||||
@@ -322,30 +446,46 @@ of comparing ecosystems.
|
|||||||
|
|
||||||
## Comparison result (`docs/rust_vs_cpp.md`)
|
## Comparison result (`docs/rust_vs_cpp.md`)
|
||||||
|
|
||||||
The document deliberately picks no winner; **no measurement differs by an order of
|
The document deliberately picks no winner. Machine: i5-1135G7 / Iris Xe /
|
||||||
magnitude.** Machine: i5-1135G7 / Iris Xe / Windows 11, both builds release.
|
Windows 11, both builds release.
|
||||||
|
|
||||||
| | C++ | Rust |
|
| | C++ | Rust |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| lines of code | 2434 | 2962 (+22 %) |
|
| lines of code | 2434 | 2976 (+22 %) |
|
||||||
| release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) |
|
| release from scratch | **114 s** | 194 s (thin LTO) · 182 s (no LTO) |
|
||||||
| release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) |
|
| release, one-file edit | **4.4 s** | 52.3 s (LTO) · **5.5 s** (no LTO) |
|
||||||
| debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** |
|
| debug from scratch / one-file edit | 112 s / 8.0 s | **83 s** / **5.2 s** |
|
||||||
| release exe | **1.18 MB** | 5.80 MB |
|
| release exe | **1.18 MB** | 5.80 MB |
|
||||||
| dependency sources | 775 MB **per build dir** | **80 MB** shared registry |
|
| dependency sources | ~400 MB **per build dir** | **80 MB** shared registry |
|
||||||
| startup to renderer ready | 1068 ms | **754 ms** |
|
| startup to renderer ready (best of 3) | 1068 ms | **754 ms** |
|
||||||
| CPU at 60 Hz | **11.2 %** | 12.6 % |
|
| CPU at 60 Hz | **11.2 %** | 12.6 % |
|
||||||
|
|
||||||
Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap
|
Frame rate measures nothing — both are FIFO on a 60 Hz screen. The 1.4 pp CPU gap
|
||||||
is egui rebuilding its layout every frame, not the language. The +22 % line count
|
is egui rebuilding its layout every frame, not the language. The one sharp build
|
||||||
is almost entirely the **missing vk-bootstrap replacement**: `Context` + `Swapchain`
|
cell (52 s) is the price of thin LTO, not of Rust — with LTO off it is 5.5 s
|
||||||
go from 301 to 606 lines, while the rest of the Vulkan layer matches nearly line
|
against 4.4 s. Note the document's summary claim that *no* measurement differs by
|
||||||
for line. The one sharp build cell (52 s) is the price of thin LTO, not of Rust —
|
an order of magnitude is false on its own numbers: 52.3 / 4.4 = **11.9×**. Every
|
||||||
with LTO off it is 5.5 s against 4.4 s.
|
other cell stays under 10×.
|
||||||
|
|
||||||
|
The +22 % line count is **mostly, not almost entirely**, the missing vk-bootstrap
|
||||||
|
replacement: `Context` + `Swapchain` go from 301 to 606 lines, which is 305 of the
|
||||||
|
542-line overhang (56 %); the whole Vulkan layer accounts for 346 (64 %). The rest
|
||||||
|
is real — mesh +117, editor +104, app +90, core −64, and 176 in-module test lines
|
||||||
|
against C++'s 51.
|
||||||
|
|
||||||
Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj`
|
Loader parity: `cow.obj` matches exactly (2451 verts / 4898 tris); `plane.obj`
|
||||||
differs by 0.13 %, entirely due to the reading libraries (tobj drops vertices no
|
differs by 0.13 %, entirely due to the reading libraries (tobj drops the 35
|
||||||
face references; fan triangulation of n-gons versus tinyobjloader's earcut).
|
vertices no face references; fan triangulation of n-gons versus tinyobjloader's
|
||||||
|
earcut).
|
||||||
|
|
||||||
|
**Known errors in `docs/rust_vs_cpp.md`** (numbers below are the measured truth):
|
||||||
|
line-count total is 2976 / 494 comments, not 2962 / 492 — the `mesh` row is stale by
|
||||||
|
the 17 lines of the ignored diagnostic test; "775 MB per build dir" is the foreign
|
||||||
|
`build/` tree (see the Build section) and so is the matching "10 packages in the
|
||||||
|
C++ graph" — there are nine; the Rust package graph is 82, not 76; `plane.obj` has
|
||||||
|
119 faces with five or more vertices, not 121; "three Catch2 tests ported verbatim"
|
||||||
|
should be *all three ported, two more added*; and `Renderer::Impl`'s destructor is
|
||||||
|
22 lines, not thirty.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -354,14 +494,17 @@ face references; fan triangulation of n-gons versus tinyobjloader's earcut).
|
|||||||
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
|
The active research work. A **D2Q9 Lattice Boltzmann solver with the entropic KBC
|
||||||
collision operator**, written from scratch in Rust (not a port): flow past bodies in
|
collision operator**, written from scratch in Rust (not a port): flow past bodies in
|
||||||
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall
|
a channel, two interchangeable backends, a ×2 nested AMR patch, sub-grid wall
|
||||||
models, and GIF output locked to *physical* flow time. Needs **Rust 1.75+**.
|
models, and GIF output locked to *physical* flow time. `Cargo.toml` declares
|
||||||
Physics is anchored to the method authors' papers in `docs/origins/`; formula
|
**Rust 1.75+**; the Dockerfile pins `rust:1.97-bookworm`. Physics is anchored to the
|
||||||
references in the code follow the 2D paper (arXiv:1507.02509).
|
method authors' papers in `docs/origins/`; formula references in the code follow the
|
||||||
|
2D paper (arXiv:1507.02509), the only arXiv id in the tree, with the wall models
|
||||||
|
citing Dorschner et al. *JFM* 801 (2016) for `grad` and Malaspinas 2015 /
|
||||||
|
Coreixas et al. *PRE* 96 for `hrr`.
|
||||||
|
|
||||||
`README.md` there is the authoritative status/validation log (~600 lines: what is
|
`README.md` there is the authoritative status/validation log (609 lines: what is
|
||||||
verified, against which equation, with what measured numbers) and `bench/README.md`
|
verified, against which equation, with what measured numbers) and `bench/README.md`
|
||||||
covers the validation campaign. **Read them before changing physics or defaults** —
|
(248 lines) covers the validation campaign. **Read them before changing physics or
|
||||||
most defaults are the outcome of a documented measurement, not a guess.
|
defaults** — most defaults are the outcome of a documented measurement, not a guess.
|
||||||
|
|
||||||
## Build / test / run
|
## Build / test / run
|
||||||
|
|
||||||
@@ -370,10 +513,14 @@ cd docs/theory/2d_solver
|
|||||||
cargo build --release # with the GPU backend (default feature `gpu`)
|
cargo build --release # with the GPU backend (default feature `gpu`)
|
||||||
cargo build --release --no-default-features # CPU only, no wgpu
|
cargo build --release --no-default-features # CPU only, no wgpu
|
||||||
cargo test --release # 29 fast tests
|
cargo test --release # 29 fast tests
|
||||||
cargo test --release -- --include-ignored # + 4 long paper benchmarks (~18 s)
|
cargo test --release -- --include-ignored # + 4 long CPU runs (~18 s)
|
||||||
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
|
cargo test --release -- --ignored taylor_green_kbc_vs_bgk --nocapture # one diagnostic
|
||||||
```
|
```
|
||||||
|
|
||||||
|
33 `#[test]` functions exist (18 in `math.rs`, 9 in `cpu.rs`, 6 in `gif.rs`); the
|
||||||
|
4 `#[ignore]`d ones are all in `cpu.rs` and are two reference benchmarks plus two
|
||||||
|
diagnostics, not four benchmarks.
|
||||||
|
|
||||||
**Never build `--no-default-features` last.** Both builds write the same
|
**Never build `--no-default-features` last.** Both builds write the same
|
||||||
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
|
`target/release/kbc2d`, so a CPU-only build silently overwrites the wgpu one and
|
||||||
`--backend gpu` then refuses to run. Order: no-default-features first, normal
|
`--backend gpu` then refuses to run. Order: no-default-features first, normal
|
||||||
@@ -385,20 +532,27 @@ build second.
|
|||||||
--gif wake.gif --gif-field vorticity --verbose full
|
--gif wake.gif --gif-field vorticity --verbose full
|
||||||
```
|
```
|
||||||
|
|
||||||
`--help` groups every key by role (Physics, Grid, Body, Time, Scheme, Animation,
|
`--help` groups every key by role, under **Russian** headings —
|
||||||
Output). Note `--time <seconds>` as an alternative to `--steps`: a step is not a
|
`Физика, Сетка, Тело, Время, Схема, Анимация, Вывод`. Note `--time <seconds>` as an
|
||||||
fixed slice of time (δt = u_lat·δx/u_phys), so refining the cell silently shortens
|
alternative to `--steps` (the two conflict): a step is not a fixed slice of time
|
||||||
a fixed step count.
|
(`dt = u_lat * dx / u_phys`, literally `math.rs`'s `Units::new`), so refining the
|
||||||
|
cell silently shortens a fixed step count.
|
||||||
|
|
||||||
|
Defaults worth knowing beyond the three discussed below: `--backend` is **`cpu`**,
|
||||||
|
`--collision kbc`, `--outlet extrapolate`, `--case channel`, `--refine 2`.
|
||||||
|
|
||||||
## File roles — one concern each
|
## File roles — one concern each
|
||||||
|
|
||||||
|
`src/` holds exactly these five files — no `lib.rs`, no `benches/`, no integration
|
||||||
|
tests.
|
||||||
|
|
||||||
| file | owns |
|
| file | owns |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, unit conversion. Tests against the papers' formulas live here. `pub type R = f64`. |
|
| `src/math.rs` | **all solver mathematics**, nodewise and pure: D2Q9 lattice, product-form entropic equilibrium, shear projector, γ stabiliser, collision, Zou–He, body SDFs, wall-model closures, unit conversion, spectral diagnostics (`strouhal`). Tests against the papers' formulas live here. `pub type R = f64`. |
|
||||||
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. |
|
| `src/cpu.rs` | CPU backend: AoS layout, rayon, **and the shared topology** — `Geom::build` (masks), Bouzidi link assembly, `Patch` (AMR level coupling), `initial_field`. Also holds the four ignored reference runs. |
|
||||||
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
|
| `src/gpu.rs` | GPU backend: wgpu + WGSL (Vulkan/DX12/Metal), SoA layout, f32. Re-implements the physics line-for-line in WGSL but **imports topology from `cpu`** rather than duplicating it. |
|
||||||
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. |
|
| `src/main.rs` | CLI, problem assembly, step loop, live output, report, CSV. Owns the `Spec` / `StepRec` / `FieldKind` contract shared by both backends. |
|
||||||
| `src/gif.rs` | encoding, palettes, normalisation, and the physical-time frame timing. |
|
| `src/gif.rs` | encoding, palettes, normalisation, the HUD, delay dithering, and the physical-time frame timing. |
|
||||||
|
|
||||||
## Invariants (read before editing)
|
## Invariants (read before editing)
|
||||||
|
|
||||||
@@ -407,47 +561,63 @@ a fixed step count.
|
|||||||
`cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the
|
`cpu::Patch` / `cpu::initial_field`. Keep it that way — a past regression had the
|
||||||
GPU silently running Bouzidi for `--wall staircase`, caught only because two
|
GPU silently running Bouzidi for `--wall staircase`, caught only because two
|
||||||
models produced bit-identical output where they had to differ. Backend parity is
|
models produced bit-identical output where they had to differ. Backend parity is
|
||||||
a hard requirement (CPU f64 vs GPU f32 agree to 4–5 significant digits; all four
|
a hard requirement and `bench/parity.py` checks it: CPU f64 vs GPU f32 agree to
|
||||||
wall models agree to 0.007 % on Cd) and `bench/parity.py` checks it.
|
4–5 significant digits, and for **each** wall model the two backends agree to
|
||||||
|
0.007 % on Cd. (That number is *backend* parity, not agreement between wall
|
||||||
|
models — those differ from one another by ~1 %.)
|
||||||
- **The backends have different step contracts.** `Backend` in `main.rs` is an
|
- **The backends have different step contracts.** `Backend` in `main.rs` is an
|
||||||
enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` /
|
enum, not a trait: CPU `step()` returns one `StepRec`, GPU `advance(out)` /
|
||||||
`flush(out)` push a *batch* — results accumulate in a 128-slot ring and sync once
|
`flush(out)` push a *batch* — results accumulate in a 128-slot ring (`HIST`,
|
||||||
per batch. Adding a per-step GPU readback outside `StepRec` reintroduces a
|
declared identically in Rust and WGSL) and sync once per batch. Adding a per-step
|
||||||
`map_async` + `poll(Wait)` per step, which cost ~8× throughput before batching.
|
GPU readback outside `StepRec` reintroduces a `map_async` + `poll(Wait)` per step,
|
||||||
- **`GREL = 1e-8` is a relative threshold** — a fraction of ⟨Δh|Δh⟩, not an
|
which cost ~8× throughput before batching (measured 863 → 6715 steps/s).
|
||||||
absolute one. ⟨Δh|Δh⟩ is quadratic in non-equilibrium and physically tiny
|
- **`GREL = 1e-8` is a relative threshold.** The test is `den > GREL * nrm`, where
|
||||||
(~1e-7…1e-9), so an absolute threshold fires almost everywhere and silently
|
`den = ⟨Δh|Δh⟩` and `nrm = ⟨Δ|Δ⟩` — so GREL is a fraction of the **full**
|
||||||
substitutes γ = 2, i.e. plain LBGK instead of KBC. The report prints the degenerate-
|
non-equilibrium norm, not of ⟨Δh|Δh⟩ itself. ⟨Δh|Δh⟩ is quadratic in
|
||||||
node fraction; on a healthy threshold it must be ~0 (except the very first step).
|
non-equilibrium and physically tiny (~1e-7…1e-9), so an absolute threshold fires
|
||||||
|
almost everywhere and silently substitutes γ = 2, i.e. plain LBGK instead of KBC.
|
||||||
|
The report prints the degenerate-node fraction; on a healthy threshold it must be
|
||||||
|
~0 (except the very first step).
|
||||||
- **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no
|
- **GPU is f32 and cannot be otherwise** — WGSL has no `f64` type at all, so no
|
||||||
hardware helps. Convergence studies therefore run on CPU, everything else on GPU;
|
hardware helps. Convergence studies therefore run on CPU, everything else on GPU;
|
||||||
below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude.
|
below a true error of ~1e-3 f32 diverges from f64 by an order of magnitude.
|
||||||
Reductions and force sums use compensated (Kahan–Neumaier) summation — that is
|
The **statistics reductions** (ρ, ⟨γ⟩, node and degenerate counts) use compensated
|
||||||
where f32 was losing most of its precision.
|
Kahan–Neumaier summation via `kadd`. The **force/torque kernel does not** — it
|
||||||
|
accumulates in plain f32 and reduces with a naive tree. Both `README.md:366` and
|
||||||
|
earlier revisions of this file claim otherwise; the code is the authority.
|
||||||
- **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` /
|
- **GPU dispatch is 2-D with a linear index rebuilt in the shader** (`lin()` /
|
||||||
`wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The
|
`wlin()`), lifting the 65535-workgroup limit that capped grids at ~2048². The
|
||||||
workgroup → node mapping stays exactly linear, which is what lets the reductions
|
workgroup → node mapping stays exactly linear, which is what lets the reductions
|
||||||
keep working; verified to 4096×4096.
|
keep working; verified to 4096×4096.
|
||||||
- **Population storage has two binding layouts.** One combined binding normally;
|
- **Population storage has two binding layouts.** Normally two bindings (`f` and
|
||||||
**nine bindings, one per direction**, when the adapter's max binding size is small
|
`post`). When the adapter's max binding size is small (dzn/WSL2 caps a binding at
|
||||||
(dzn/WSL2 caps a binding at 128 MiB while allowing a 2047 MiB buffer). Chosen from
|
128 MiB while allowing a 2047 MiB buffer), each array is split across nine
|
||||||
adapter limits, no `switch` in the accessors; both give identical numbers, split
|
bindings — one per direction, **18 in total** — raising the ceiling from 3.73 M to
|
||||||
costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS=1` to compare on one
|
33.5 M nodes. The choice comes from adapter limits. The split layout is the one
|
||||||
card.
|
built out of `switch` statements in `fget`/`fset`/`pget`/`pset`; the combined
|
||||||
|
layout has none, which is the whole reason it is kept. Both give identical numbers;
|
||||||
|
the split costs 5.7 % bandwidth. Force it with `KBC2D_SPLIT_POPULATIONS` set to
|
||||||
|
anything other than `0`.
|
||||||
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
|
- **`MAX_BODY_BUCKETS = 4`.** Geometry is assembled from any number of bodies, but
|
||||||
per-body force is bucketed — bodies beyond the fourth have their forces merged.
|
per-body force is bucketed — bodies beyond the fourth have their forces merged
|
||||||
|
(`min(body, MAXB-1)` in three places), and the CLI warns when that happens.
|
||||||
- **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps
|
- **Defaults encode measurements, not taste**: `--init uniform` (a rest start pumps
|
||||||
a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge
|
a quarter-wave channel resonance to 75 % of U and wrecks St/Cd/Cl — and the sponge
|
||||||
cannot remove it, since a standing mode has its pressure node exactly where the
|
cannot remove it, since a standing mode has its pressure node exactly where the
|
||||||
sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but higher bulk viscosity,
|
sponge sits); `--kbc-model n1` (equal accuracy to `n2`, but ~4× higher bulk
|
||||||
which damps longitudinal acoustics); `--wall hrr` (**provisional** — chosen on
|
viscosity, which damps longitudinal acoustics); `--wall hrr` (**provisional** —
|
||||||
scheme structure, pending campaign group C; `bouzidi` resolves geometry 4× better).
|
chosen on scheme structure, pending campaign group C; `bouzidi` resolves geometry
|
||||||
|
~4× better). Accepted values: `--init uniform|rest`, `--kbc-model n1|n2`,
|
||||||
|
`--wall hrr|bouzidi|grad|staircase`.
|
||||||
|
|
||||||
## Validation campaign (`bench/`)
|
## Validation campaign (`bench/`)
|
||||||
|
|
||||||
115 runs, ≈90 GPU-hours in nine groups; each run writes logs, series, a
|
115 runs in nine groups (A 12, B 16, C 20, D 14, E 12, F 9, G 14, H 9, I 9), ≈90
|
||||||
machine-readable summary and a GIF into its own folder under `out/` (untracked —
|
machine-hours — ≈87 GPU-hours over 108 runs plus ≈3 CPU-hours over the 7 CPU runs
|
||||||
**those are results, never clean them**).
|
of group A. Each run writes `cmd.txt`, `log.txt`, `report.txt`, `series.csv` and
|
||||||
|
`summary.json` into its own folder under `out/` (untracked — **those are results,
|
||||||
|
never clean them**). Only **59 of the 115** pass `--gif`; the rest produce no
|
||||||
|
animation at all.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
cd bench
|
cd bench
|
||||||
@@ -457,18 +627,28 @@ python preflight.py # start every scenario for two steps — catches
|
|||||||
./run_campaign.sh --resume # run, skipping what is already done
|
./run_campaign.sh --resume # run, skipping what is already done
|
||||||
```
|
```
|
||||||
|
|
||||||
Run `preflight.py` for real — it caught an entire group failing on a GPU limit and
|
`run_campaign.sh` also takes `--smoke`, `--group`, `--only` and `--budget-hours`;
|
||||||
nine runs passing `--body-x` twice, both of which would otherwise have surfaced
|
`run_campaign.ps1` is the Windows twin. Run `preflight.py` for real — it caught
|
||||||
twenty hours into a server campaign.
|
group I failing on a GPU limit and group F's nine runs passing `--body-x` twice,
|
||||||
|
both of which would otherwise have surfaced twenty hours into a server campaign.
|
||||||
|
|
||||||
|
Note the campaign has barely started: `out/` currently holds four finished runs
|
||||||
|
(`A01`, `A02`, `A05`, `C09`) plus `campaign.log` and `summary.csv`.
|
||||||
|
|
||||||
## Deployment
|
## Deployment
|
||||||
|
|
||||||
Published image **`notbigghost/kbc2d:1.2.0`** (`linux/amd64`); the server needs
|
Published image **`notbigghost/kbc2d:1.2.0`**; the server needs only
|
||||||
only `docker-compose.server.yml`, not the sources. `ENTRYPOINT` is the campaign
|
`docker-compose.server.yml` (the only compose file with no `build:` section), not
|
||||||
driver and `CMD` defaults to `--dry-run`, so a stray `docker run` prints an
|
the sources. `ENTRYPOINT` is the campaign driver and `CMD` defaults to `--dry-run`,
|
||||||
estimate instead of starting a 100-hour job. Check the card first
|
so a stray `docker run` prints an estimate instead of starting a 100-hour job. Check
|
||||||
(`--profile check run --rm vulkan`) — a missing GPU is better discovered in a
|
the card first (`--profile check run --rm vulkan`) — a missing GPU is better
|
||||||
minute than in an hour.
|
discovered in a minute than in an hour. The `check` profile also carries
|
||||||
|
`preflight`, `calibrate` and `plan` services (plus `parity` in the WSL file).
|
||||||
|
|
||||||
|
Two version caveats: the `1.2.0` tag lives only in the compose files and prose —
|
||||||
|
`Cargo.toml` still says `version = "0.1.0"` and the Dockerfile defaults
|
||||||
|
`ARG VERSION=dev`. And `linux/amd64` is asserted in `README.md` only; no compose
|
||||||
|
file sets `platform:` and the Dockerfile sets no `--platform`.
|
||||||
|
|
||||||
Two environment traps, both already handled in the image but overridable from
|
Two environment traps, both already handled in the image but overridable from
|
||||||
outside: NVIDIA Container Toolkit only injects the Vulkan ICD when
|
outside: NVIDIA Container Toolkit only injects the Vulkan ICD when
|
||||||
@@ -478,7 +658,8 @@ arrives over `/dev/dxg`. That path instead uses **dzn** (Mesa's Vulkan→D3D12
|
|||||||
translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build
|
translation, why the base image is `archlinux:base` — Debian/Ubuntu do not build
|
||||||
dzn), needs no NVIDIA runtime, and requires
|
dzn), needs no NVIDIA runtime, and requires
|
||||||
`WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports
|
`WGPU_ALLOW_UNDERLYING_NONCOMPLIANT_ADAPTER=1` because dzn reports
|
||||||
`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default. Use
|
`conformanceVersion = 0.0.0.0` and wgpu hides such adapters by default (the variable
|
||||||
|
is read only because `gpu.rs` builds the instance with `.with_env()`). Use
|
||||||
`docker-compose.wsl.yml` there.
|
`docker-compose.wsl.yml` there.
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -486,16 +667,63 @@ dzn), needs no NVIDIA runtime, and requires
|
|||||||
# `docs/theory/` — Python prototype (all branches)
|
# `docs/theory/` — Python prototype (all branches)
|
||||||
|
|
||||||
The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow
|
The earlier **Python/CuPy** implementation of the same LBM physics — cylinder flow
|
||||||
with a ×2 nested AMR patch and SDF+Bouzidi boundaries, **GPU/CuPy only, no CPU
|
with a nested AMR patch and SDF+Bouzidi boundaries. `kbc2d` is a deliberate rewrite
|
||||||
fallback**. `kbc2d` is a deliberate rewrite of this, not a port, and the two differ
|
of this, not a port. Do not fold any of it into the CMake build, and do not treat it
|
||||||
in two documented places (no collision inside the body; restriction skips fine
|
as dead code — but do not treat it as one program either.
|
||||||
source nodes inside the body). Still live: it owns the notebook's figures.
|
|
||||||
|
|
||||||
- `solver_2x_sdf/README.md` — maps every file to the physics it owns; entry points
|
**This directory holds four generations of the same solver.** CLAUDE.md used to name
|
||||||
`run.py`, `run_blockage.py`, `run_factors.py`.
|
only the newest, which made the rest look like clutter. Oldest to newest:
|
||||||
- `demos_gpu/` — demo runs producing the notebook's figures/GIFs; they import
|
|
||||||
physics from `solver_2x_sdf` and never duplicate it.
|
|
||||||
- `kbc_lbm.ipynb` — the write-up; embeds pre-rendered artefacts, executes nothing.
|
|
||||||
- `docs/origins/` — the source PDFs behind both implementations.
|
|
||||||
|
|
||||||
Do not fold any of this into the CMake build, and do not treat it as dead code.
|
| generation | files | status |
|
||||||
|
|---|---|---|
|
||||||
|
| the original write-up | `_legacy/kbc_lbm.ipynb` (64 cells, executed, outputs stored) + `_legacy/kbc_lbm.py` | superseded |
|
||||||
|
| CPU reference prototypes | `_amr_anims.py`, `_amr_anims_sdf.py` (pure NumPy) | superseded, but these are the "verified CPU prototypes" the GPU core is checked against |
|
||||||
|
| GPU core + runners | `amr_gpu_core.py` (**has a NumPy CPU fallback**) driven by `amr_cylinder_gpu.py` (no SDF) and `amr_cylinder_sdf_gpu.py` (SDF+Bouzidi); both self-titled "ОСНОВНОЙ", three modes: none / 2× / nested 2×+4× | superseded as a model, **still the producer of committed figures 11e–11g and four `*_cyl*.gif` animations** — deleting it orphans notebook images |
|
||||||
|
| optimised rewrite | `amr_opt/` (`*_opt.py`: no NumPy fallback, `cp.fuse`, precomputed link masks, CUDA Graphs) | dead intermediate — produced nothing that is committed |
|
||||||
|
| **the current model** | `solver_2x_sdf/` (17 modules + README), demos in `demos_gpu/` | live |
|
||||||
|
|
||||||
|
So "**GPU/CuPy only, no CPU fallback**" is true of `solver_2x_sdf/` (`backend.py`
|
||||||
|
raises outright), `demos_gpu/` and `amr_opt/` — but **not** of `amr_gpu_core.py`,
|
||||||
|
which silently falls back to NumPy, nor of `_amr_anims*.py` and `_legacy/`, which are
|
||||||
|
CPU-only.
|
||||||
|
|
||||||
|
Also committed and easy to mistake for junk: `figures/` (14 PNG) and `anim/`
|
||||||
|
(10 GIF, 55 MB) are tracked **deliberately** so the notebook renders on a fresh
|
||||||
|
clone; `_amr_probe_data*.npz` are probe scalars one runner writes and the other
|
||||||
|
reads; `factors.csv` is a hand-preserved snapshot of the Experiment-№1 registry
|
||||||
|
(nothing writes that path — `run_factors.py` writes `solver_2x_sdf/out/factors.csv`,
|
||||||
|
which is gitignored).
|
||||||
|
|
||||||
|
- `solver_2x_sdf/README.md` — maps every file to the physics it owns (17 rows, 17
|
||||||
|
modules, exact); entry points `run.py`, `run_blockage.py`, `run_factors.py`.
|
||||||
|
**Two defects in it:** it suggests switching `cfg.collision="bgk"` or registering
|
||||||
|
a TRT operator in a `get` registry — neither the config field nor the registry
|
||||||
|
exists, and the same README says two paragraphs later that the operator is fixed;
|
||||||
|
and it declares **all of its own quantitative results stale** pending a re-run
|
||||||
|
after the `GREL` fix, which also invalidates notebook §22–24.
|
||||||
|
- `demos_gpu/` — 8 scripts + README producing 17 of the notebook's artefacts. They
|
||||||
|
import the model's operators from `solver_2x_sdf` and add only what the model
|
||||||
|
deliberately lacks: `_common.py` reimplements a polynomial equilibrium, BGK, the
|
||||||
|
H-function and a γ-field for visualisation. `figures_static.py` is pure matplotlib
|
||||||
|
and imports no physics at all.
|
||||||
|
- `kbc_lbm.ipynb` — the write-up. It does not merely "execute nothing": it has **20
|
||||||
|
cells, all markdown, zero code cells**. Artefacts are relative-path links, not
|
||||||
|
embedded data, so it is only readable in place — and **4 of its 26 image links are
|
||||||
|
dead**, all pointing into the gitignored `solver_2x_sdf/out/`.
|
||||||
|
- `docs/origins/` — the six source PDFs behind both implementations (seven on
|
||||||
|
`research`, which adds Grad's approximation).
|
||||||
|
|
||||||
|
## How kbc2d differs from this prototype
|
||||||
|
|
||||||
|
Documented in one paragraph at `docs/theory/2d_solver/README.md:390` — **on branch
|
||||||
|
`research` only**; nothing under `docs/theory/` mentions the divergence. There are
|
||||||
|
**three** differences, not two:
|
||||||
|
|
||||||
|
1. kbc2d does not collide inside the body (those populations are fictitious; γ
|
||||||
|
statistics are then gathered strictly over fluid). The Python solver *does*
|
||||||
|
collide there — `kbc_collide` takes no fluid mask.
|
||||||
|
2. kbc2d's restriction additionally skips coarse nodes whose *fine source* node lies
|
||||||
|
inside the body. The Python restriction already gates on the **coarse
|
||||||
|
destination** mask.
|
||||||
|
3. kbc2d's CPU backend is f64 and its GPU backend f32; the Python side is fp32 by
|
||||||
|
default, f64 only via `AMR_FP64`.
|
||||||
|
|||||||
Reference in New Issue
Block a user