From 385be2eca72eb3833efa4be0a0088b34e764788a Mon Sep 17 00:00:00 2001
From: shuishen <1109946754@qq.com>
Date: Mon, 31 Aug 2026 09:04:45 +0800
Subject: [PATCH] feat(pointcloud): complete annotation and result lifecycle workflows

---
 capabilities/05-3d-pointcloud/README.md |  318 +++++++++++++++++++++++++++++++++++++++++++++++++++-
 1 files changed, 311 insertions(+), 7 deletions(-)

diff --git a/capabilities/05-3d-pointcloud/README.md b/capabilities/05-3d-pointcloud/README.md
index 49a5fe9..c33876b 100644
--- a/capabilities/05-3d-pointcloud/README.md
+++ b/capabilities/05-3d-pointcloud/README.md
@@ -21,6 +21,22 @@
 export, or image sequence is copied unchanged into a raw run directory before a
 console run is processed.
 
+The local console accepts 3-1000 JPG/JPEG photos per reconstruction run. The
+same-flight and same-camera requirement still applies. CPU reconstruction is an
+asynchronous job: a large sequence can take hours or longer, depending on image
+resolution, overlap, and available storage; uploads and source images remain
+byte-for-byte preserved.
+
+The console shows completed reconstruction milestones from the fixed sparse and
+dense scripts, including feature extraction, matching, sparse mapping, RGB
+preparation, undistortion, dense fusion, meshing, texture export, and final
+artifact packaging. Its percentage is a weighted workflow-stage milestone, not
+a remaining-time estimate, because COLMAP and OpenMVS do not provide reliable
+live progress inside their individual native processing steps.
+
+The console permits this asynchronous CPU job to run for up to 24 hours before
+reporting a timeout. This is a safety ceiling, not a runtime estimate.
+
 ```powershell
 $py = .\.venvs\05-3d-pointcloud\Scripts\python.exe
 & $py .\capabilities\05-3d-pointcloud\run_pointcloud_understanding.py `
@@ -50,6 +66,45 @@
   --threads 12 --max-resolution 2400 --dense-resolution-level 0 `
   --dense-number-views 8 --dense-number-views-fuse 2 --target-faces 800000
 ```
+
+### CPU and GPU compatibility
+
+Photo reconstruction does **not** require a GPU. The CPU route above is the
+complete, verified baseline: image preparation, sparse SfM, dense MVS, meshing,
+texturing, GLB export, and metadata all run locally on CPU. It is slower for
+large or high-resolution photo sets, but it is the default route for any
+computer and remains the console's only selectable reconstruction route.
+
+CUDA is an optional acceleration experiment for the OpenMVS dense depth/fusion
+stage only. It does not move EXIF processing, image conversion, sparse SfM,
+meshing, texturing, exports, or the local console to the GPU. Use it only with a
+separately installed, compatible CUDA OpenMVS binary and an empty output
+directory:
+
+```powershell
+& $py .\capabilities\05-3d-pointcloud\run_cpu_dense_reconstruction.py `
+  --input <coherent-photo-directory> `
+  --sparse-model <verified-sparse-model-directory> `
+  --output <new-output-directory> `
+  --openmvs-bin <cuda-openmvs-binary-directory> `
+  --dense-device cuda
+```
+
+CUDA is accepted only when the native OpenMVS log explicitly reports CUDA/GPU
+execution; the run never silently falls back to CPU. `run_metadata.json` records
+the requested and actual dense-MVS device. On this workstation, PyTorch CUDA is
+available, but both tested official CUDA MVS binaries failed on the six-photo
+validation set before depth-map output (OpenMVS: CUDA error 801; COLMAP: CUDA
+device/context failure). Consequently CUDA MVS is not exposed in the console or
+claimed as usable here. CPU MVS remains the supported path until a compatible
+driver/toolchain combination or a validated server environment is available.
+
+The separate supervised point-cloud semantic-model training and inference flow
+does support `auto`, `cpu`, and `cuda` in the local console. This applies only
+to PyTorch model computation, not to photo reconstruction, point-cloud file I/O,
+DSM/vector generation, or export. `auto` falls back to CPU with a recorded
+reason; an explicit CUDA request fails when the fixed CUDA environment cannot
+be verified.
 
 WebODM/ODX-compatible RTK input preparation (metadata only; it neither starts
 WebODM nor reconstructs a model):
@@ -254,6 +309,135 @@
 metric-accurate. Obtain registered source data and manually reviewed truth before
 operational or commercial use.
 
+## Photo, pose, and point-cloud multimodal preflight
+
+The planned `same-scene photo features + RGB/XYZ point cloud + Point Transformer
+V3` route begins with a bounded, CPU-only compatibility check. It is B work:
+Laspy, PyProj, OpenCV, Pillow, and NumPy verify ODM/OpenSfM camera data and
+point-to-image projection. `geoai-py` has no built-in API for this workflow.
+
+The preflight requires a coloured LAS/LAZ with a declared CRS, original photos,
+ODM/OpenSfM `cameras.json`, and `shots.geojson`. It validates exact filename
+matching, normalises ODM's optional `v2` camera-name prefix, transforms every
+shot GeoJSON WGS84 point into the LAS CRS, then compares it to the reported
+camera-centre translation. It samples the LAS with bounded memory and projects
+those points through the Brown camera model and OpenSfM Rodrigues pose. Source
+photos, LAS data, and camera files remain unchanged.
+
+```powershell
+$py = .\.venvs\05-3d-pointcloud\Scripts\python.exe
+& $py .\capabilities\05-3d-pointcloud\prepare_multimodal_pointcloud_dataset.py `
+  --images .\baseData\img `
+  --las .\baseData\las\Task-of-2026-08-22T084705607Z-georeferenced_model.las `
+  --cameras .\baseData\camera\Task-of-2026-08-22T084705607Z-cameras.json `
+  --shots .\baseData\camera\Task-of-2026-08-22T084705607Z-shots.geojson `
+  --output .\shared\outputs\05-3d-pointcloud\multimodal-preflight-<new-run-id>
+```
+
+The output contains `input_inventory.json`, `pose_projection_report.json`,
+12 representative `projection-overlays/*.jpg`, and `run_metadata.json`. The
+acceptance gate requires no missing photo/shot pairs, a maximum WGS84-to-LAS
+camera-centre XY residual at or below 0.25 m, and at least 50 sampled LAS points
+inside every representative calibrated image frame. It proves data availability
+only. A sparse red-point overlay is not an alignment acceptance artifact: it can
+be in-frame while remaining visually unreadable or offset.
+
+Real preflight on 2026-08-25 used 550 original 4032 x 3024 DJI M4TD images,
+the 75,870,130-point RGB `EPSG:32650` LAS export, one Brown calibration, and
+550 shots. All filenames matched. The maximum 550-shot WGS84-to-UTM camera
+centre XY residual was 0.0 m. From a deterministic 120,000-point bounded LAS
+sample, all 12 distributed diagnostic photos passed the in-frame check (minimum
+17,319; median 34,403 sampled points). The artefacts are in
+`shared/outputs/05-3d-pointcloud/multimodal-preflight-20260825-550-photos/`.
+This is not yet a feature fusion, semantic label, occlusion, or accuracy result.
+
+The second diagnostic uses the same pose data but renders a deterministic
+one-million-point RGB LAS sample through a closest-point z-buffer. It compares
+the current ODM pose convention with an inverse-rotation negative control,
+checks the derived viewing axis against the DJI XMP gimbal yaw/pitch, and writes
+three inspectable PNGs per representative image: point-only RGB,
+photo/point-RGB blend, and point/edge overlay. It automatically detects whether
+LAS RGB values are effectively 8-bit or 16-bit; LAS dimensions alone are not a
+reliable colour-scale indicator.
+
+```powershell
+& $py .\capabilities\05-3d-pointcloud\diagnose_multimodal_projection.py `
+  --images .\baseData\img `
+  --las .\baseData\las\Task-of-2026-08-22T084705607Z-georeferenced_model.las `
+  --cameras .\baseData\camera\Task-of-2026-08-22T084705607Z-cameras.json `
+  --shots .\baseData\camera\Task-of-2026-08-22T084705607Z-shots.geojson `
+  --output .\shared\outputs\05-3d-pointcloud\multimodal-projection-diagnostic-<new-run-id>
+```
+
+On 2026-08-26, three distributed images passed this stronger check. The current
+rotation has median RGB MAE 30.501, versus 60.612 for inverse rotation, and a
+48.5% median improvement over shuffled pixel locations. Its median DJI gimbal
+view-axis difference is 0.647 degrees. The screenshot image `0450` has RGB MAE
+23.641, correlation 0.7529, and 60.53% improvement over shuffled locations;
+its inverse rotation correlation is 0.0306. The result is stored at
+`shared/outputs/05-3d-pointcloud/multimodal-projection-diagnostic-20260826-v2/`.
+This accepts the current pose convention for photo-feature projection, but does
+not claim sub-pixel calibration, complete visibility/occlusion filtering, or
+semantic model accuracy.
+
+## Multi-view visibility and feature fusion
+
+`prepare_multiview_point_features.py` is the next CPU-only data-preparation
+step for the planned multimodal semantic workflow. It selects a bounded LAS
+sample, projects each point into a selected set of verified ODM photos, and
+uses a two-pixel closest-depth z-buffer to reject points hidden behind another
+point in the same image cell. For each remaining observation it gathers eight
+deterministic visual descriptors: photo RGB, hue sine/cosine, saturation,
+gradient magnitude, and local-intensity standard deviation. Each point stores
+the multi-view mean and standard deviation separately from its original
+`XYZ + LAS RGB`, plus its valid-view count.
+
+```powershell
+& $py .\capabilities\05-3d-pointcloud\prepare_multiview_point_features.py `
+  --images .\baseData\img `
+  --las .\baseData\las\Task-of-2026-08-22T084705607Z-georeferenced_model.las `
+  --cameras .\baseData\camera\Task-of-2026-08-22T084705607Z-cameras.json `
+  --shots .\baseData\camera\Task-of-2026-08-22T084705607Z-shots.geojson `
+  --output .\shared\outputs\05-3d-pointcloud\multiview-feature-<new-run-id> `
+  --max-points 50000 --max-shots 550 --z-buffer-cell-size 2
+```
+
+The run creates `multiview-point-features.npz`, a coverage-coloured PLY, an
+ordered original-RGB `multiview-annotation-source.ply`, per-photo visibility
+CSV, coverage JSON, and metadata. The compressed NPZ has
+`xyz`, `las_rgb`, `photo_feature_mean`, `photo_feature_stddev`,
+`visible_view_count`, and ordered `feature_names` arrays. It intentionally has
+no semantic class field or pseudo-label.
+
+The annotation PLY never uses the coverage pseudo-colours. Its vertex `i` is
+the same `xyz[i]` / `las_rgb[i]` row in the NPZ, and `run_metadata.json` records
+both SHA-256 values and the point count. The local console exposes it only when
+the PLY header, both NPZ array lengths, declared file names, and checksums all
+match. This prevents a display preview or an incomplete fusion directory from
+becoming a label source. To add this source once to an older completed fusion
+run without recomputing photo features:
+
+```powershell
+& $py .\capabilities\05-3d-pointcloud\prepare_multiview_point_features.py `
+  --output .\shared\outputs\05-3d-pointcloud\multiview-feature-<completed-run-id> `
+  --annotation-source-only
+```
+
+Real CPU verification on 2026-08-26 processed 50,000 deterministic LAS sample
+points against all 550 photos in 899.880 seconds. Every point had an occlusion-
+filtered observation; all had at least two and at least three valid views, and
+99.978% had at least ten. The mean and median valid-view counts were 147.2697
+and 154 respectively, with 235 maximum. All visual-feature arrays were finite;
+the feature standard-deviation mean was 0.140754 and p95 was 0.320565. The
+result is in `shared/outputs/05-3d-pointcloud/multiview-feature-550-photo-sample-20260826/`.
+
+This validates the fusion format and observation coverage only. The z-buffer
+is not a full physical visibility model, and these descriptors are not a
+pretrained photo encoder. A chunked full 75,870,130-point preparation should
+only start after reviewed semantic labels and the spatial training/validation/
+test split are agreed; then the deterministic descriptors can be replaced or
+augmented with a GPU photo encoder before Point Transformer V3 training.
+
 ## Semantic point-cloud classification baseline
 
 The existing-point-cloud route now writes a separate CPU semantic-classification
@@ -265,8 +449,8 @@
 Classes are `ground`, `vegetation`, `building_structure`, `power_line`,
 `pole_tower`, and `other_unknown`. The last three names describe review outputs:
 the first version is deliberately a transparent rule baseline rather than a
-trained semantic model. It exports a colour-coded preview PLY (capped at 400,000
-points for the browser), semantic LAS class codes, class GeoTIFF/PNG, candidate
+trained semantic model. It exports a colour-coded preview PLY containing the
+complete processed point set, semantic LAS class codes, class GeoTIFF/PNG, candidate
 GeoJSON, class-count CSV, and `run_metadata.json`.
 
 Real LAS check: the unchanged user input `baseData/las/part_01.las` contains
@@ -351,24 +535,132 @@
 ## Human annotation and supervised GPU/CPU training
 
 The workbench now provides a separate human annotation and supervised training
-area. A semantic run writes a 400,000-point `*.semantic-annotation-source.ply`
+area. A semantic run writes a complete processed `*.semantic-annotation-source.ply`
 that preserves observed RGB and coordinates. Rule-coloured previews are never
  used as model inputs or ground truth. Brush and rectangle selections are saved
  as separate revisions under `shared/outputs/05-3d-pointcloud/annotations/<annotation-id>/`.
- The annotation viewer renders observed RGB through an sRGB-correct, anti-aliased
- circular point sprite. It separates navigation from annotation: browse mode
- uses left-drag rotation, while brush/rectangle mode reserves left-drag for
- labels and retains right-drag rotation, middle-drag panning and wheel zoom.
+The annotation viewer renders observed RGB through an sRGB-correct, anti-aliased
+circular point sprite. It separates navigation from annotation: browse mode
+uses left-drag rotation, while brush/rectangle mode reserves left-drag for
+labels and retains right-drag rotation, middle-drag panning and wheel zoom.
+The console displays every point in the selected annotation source directly.
+It does not replace the view with a zoom-dependent local layer, so point indices
+remain stable for browsing, brush/rectangle labels, saved revisions, and
+training. Large sources require correspondingly more browser memory and GPU
+resources.
+
+### Automatic annotation review loop
+
+After a completed local model is selected, **Automatic annotation current
+source** runs that fixed local model against the selected source without a second
+upload. It writes an immutable candidate record with the source/model SHA-256,
+per-point confidence, per-class candidate counts and a user-selected confidence threshold. The console renders a separate class-coloured prediction viewer for review alongside the RGB annotation source. In that viewer, brush/rectangle review can reject a candidate or correct it to another class; its saved correction draft remains separate from human truth. **Confirm merge candidates** applies candidate labels, then the saved review draft, then existing human labels (highest priority), and creates a new
+annotation revision only after explicit confirmation; existing human labels win
+when they overlap model candidates. That new revision can be used for the next
+training run. This is an assisted-labelling workflow, not a claim that the model
+has created ground truth. The candidate preview and confidence JSON must be
+reviewed before acceptance, especially for classes with weak or absent spatial
+test evidence.
+New annotation uploads use a separate CPU preview preparation step rather than
+the DSM/footprint pipeline: it only reads XYZ and available vertex RGB, then
+writes every readable finite XYZ point and available vertex RGB value. It therefore does
+not create rasters, vectors, or meshes and is unaffected by a source cloud's
+spatial extent. A textured-mesh PLY without readable per-vertex RGB is shown
+with neutral points for geometry review, but is explicitly ineligible for the
+current RGB semantic-model trainer.
+
+### Textured mesh RGB baking
+
+Some PLY files are textured triangle meshes rather than RGB point clouds. They
+store only vertex XYZ plus per-face UV coordinates and `TextureFile` image
+references, so a point viewer cannot recover colour from the PLY alone. The
+CPU-only `bake_textured_mesh_pointcloud.py` reads one ZIP bundle containing the
+PLY and all referenced JPG/PNG texture tiles, samples each selected triangle's
+UV centroid with a vertical texture-axis correction, and writes a binary
+`XYZ + RGB` PLY. Its points represent sampled triangle-face centres, not the
+original vertices. This is B work using NumPy/Pillow/PLY parsing, not a native
+`geoai-py` point-cloud API.
+
+```powershell
+& .\.venvs\05-3d-pointcloud\Scripts\python.exe `
+  .\capabilities\05-3d-pointcloud\bake_textured_mesh_pointcloud.py `
+  --input E:\DJI\PLY\南航\PLY\Model_0.zip `
+  --output .\shared\outputs\05-3d-pointcloud\texture-baked-<run-id> `
+```
+
+The representative `Model_0.zip` check on 2026-08-28 contained one 7,430,655
+vertex / 14,820,237 face PLY and 141 texture images. The 50,000-face sample
+used 140 referenced texture images and completed CPU baking in 174.921 seconds.
+Its RGB standard deviation was `[62.5, 58.0, 65.0]` and 80.17% of points had
+materially non-grey channel differences, confirming that texture RGB rather
+than placeholder grey was written. Visual UV alignment still needs user review
+before a complete annotation/training source is generated.
  Middle-button auto-scroll is suppressed within the viewer. The rotate-centre
  tool sets the camera target to a clicked visible point without creating a label.
  The model-orientation control independently rotates the displayed model about
  its centre on local X/Y/Z axes from 0 through 360 degrees. It is a viewer-only
  transform: source coordinates, semantic outputs and saved point indices remain
  unchanged, while point picking and rectangle selection use the transformed view.
+
+The sparse SfM, semantic-classification, supervised-prediction, and mesh
+preview panels share these display-only X/Y/Z orientation controls, a click-to-
+set orbit centre, top-down view, and view reset. They do not alter source PLY,
+GLB, LAS, semantic output, annotation index, or downloaded artifact bytes.
  A saved revision is never edited in place. If it contains an error, the console
  can remove that selected revision after confirmation; it cannot delete source
  LAS/PLY data, semantic results, or training outputs. Each revision records
  source run/checksum, point indices, class codes and time.
+
+The **Annotation, training, and model application** workspace also provides
+**Add annotation source** and **Manage annotation classes**. Add source accepts
+one local PLY/PCD/XYZ/LAS/LAZ through the bounded binary upload endpoint. It
+preserves the exact original bytes under the point-cloud raw-data layout, then
+runs a background CPU job to create a separate RGB/XYZ annotation preview; no
+existing source is overwritten. The local taxonomy starts with ground,
+vegetation, building, pole/tower, power line, and other/unknown. A user can add
+a Chinese display name, lowercase English key, and RGB colour; the console
+assigns a distinct LAS-compatible class code from 1-255. Each saved annotation
+snapshots its complete class definitions. A custom class used by any saved
+revision cannot be deleted, so historical labels and trained model metadata
+remain interpretable. Adding a source or category creates no semantic truth by
+itself: reviewed spatially distributed labels are still required for training.
+
+The selected annotation source also has **Remove complete data chain**. Before
+the irreversible confirmation, the local server calculates and displays the
+actual number of generated output directories, raw-upload copies, processed
+copies, annotation revisions, dependent training results, and model-inference
+results. Confirming removes those local workbench artifacts in dependency order.
+It never accepts a browser-supplied filesystem path and never removes
+`baseData` imagery, LAS, camera files, or any other external input. When an
+older derived source has no console-owned raw copy, only its generated output
+chain is removed and the confirmation says that the external input is retained.
+
+The model workspace also discovers a completed `multiview-feature-*` fusion
+directory as `多视角照片特征融合样本` when its ordered original-RGB PLY, NPZ and
+checksum contract pass. Its 50,000-point current sample is a human-review entry
+until it has reviewed labels. Its automatic training entry selects
+`train_multiview_point_transformer.py`, which reads only the checksum-linked
+same-order NPZ and uses 23 inputs per point: normalized XYZ, LAS RGB, eight
+photo-feature means, eight photo-feature standard deviations, and normalized
+visible-view count. It constructs 16 local geometric neighbours and trains a
+two-stage local point-attention baseline with complete XY-block train,
+validation, and test separation. It writes `model.pt`, `metrics.json`,
+`run_metadata.json`, and a colour-coded `predicted-semantic-preview.ply`.
+
+This model is explicitly **not** official Point Transformer V3 or Pointcept.
+The Pointcept/PTv3 dependency stack (`spconv` and related CUDA extensions) is
+not installed or validated on this machine, so the workbench does not make that
+claim. The baseline is CPU-compatible and can use CUDA only after the fixed
+PyTorch CUDA environment passes its probe. It cannot classify an arbitrary
+uploaded point cloud because that input would lack matching camera poses and
+multi-view photo descriptors.
+
+Training requires at least two reviewed classes with 500 points per class, and
+each class must appear in training, validation, and test XY blocks. For the
+current sample, label vegetation, pole/tower, and power line in multiple
+separated areas before clicking `多视角特征训练`. The retained checkpoint is
+selected by validation macro F1; metrics apply only to reviewed source blocks
+and are not field-wide accuracy or inspection evidence.
 
 `train_pointcloud_semantic_model.py` trains a compact PointNet-style shared MLP
 from user-confirmed labels only. It accepts `--device auto`, `cpu`, or `cuda`;
@@ -435,3 +727,15 @@
 1.554 seconds and records `processing.device: "cuda"` in `run_metadata.json`.
 This is a local labelled-block comparison only; it does not establish field-wide
 accuracy or authorise asset/inspection conclusions.
+
+## Console result lifecycle
+
+The local console treats every console-owned point-cloud case as a complete
+result chain. Before removal, it obtains a server-calculated plan and requires a
+second confirmation. A removable case is limited to
+`shared/outputs/05-3d-pointcloud/runs/<run-id>/` and its fixed console raw and
+processed copies; `baseData`, external inputs, validation/baseline assets, and
+unrelated runs are preserved. The point-cloud training-model selector separately
+removes the chosen local training directory and only its discovered model
+inference and automatic-annotation outputs. It preserves annotation revisions
+so reviewed labels are never discarded merely because a model is removed.

--
Gitblit v1.9.3