geomcheck: a mesh validity check in Python for AI-generated models

By Miles Carter · I work on modelfy.art. The six test models below were generated with Modelfy. Logged 3 October 2026.

Real output: geomcheck's per-face flags on the six test models. Blue = thin wall, purple = hidden, orange = small separate piece, red = self-intersection.

What it is, in one paragraph

geomcheck answers concrete questions about mesh geometry: does it have holes, extra pieces, crossing faces, surface nobody can see, or walls too thin to print? Everything is computed on a copy scaled so the bounding-box diagonal is 1, so thresholds work at any model size. Random sampling uses a fixed seed. There is no learned model, no training and no GPU. It does not tell you whether a mesh looks good. The study it came from was built to measure exactly that gap (see "What it can't measure").

Log 1: install from a clean environment

The docs show pip install geomcheck, but when I checked on 3 October 2026 the package was not on PyPI yet (the PyPI simple index returned 404). Installing straight from GitHub works:

python3 -m venv .venv
.venv/bin/pip install "geomcheck[viz] @ git+https://github.com/Stark-Will/geometric-probes-3d"

In a fresh Python 3.13.5 venv this pulled in trimesh 5.1.1, PyMeshLab 2025.7.post1 and embreex 4.4.0, plus NumPy 2.5.3, SciPy 1.18.1 and Matplotlib (for the [viz] renderer). On a clone of the repository, pytest -q tests/ passed all 16 synthetic-mesh tests.

One Linux gotcha: PyMeshLab's decimation filter needs the system library libOpenGL.so.0 (sudo apt-get install libopengl0 on Debian/Ubuntu). No display or GPU is needed. geomcheck.decimation_available() returned True on my box. Without that library every probe still runs, but the optional decimation step raises a clear error.

Log 2: run it on six AI-generated models

Test set: six showcase models generated with Modelfy (its backend is Tencent Hunyuan 3D). All six are public on Sketchfab, so you can download the same files: the chest, the dragon, the workshop, the guardian, the radio and the marble bust. I used the Sketchfab-compatible copies, which have meshopt compression removed and PNG textures. Each is about 150k triangles.

The whole script:

import sys, time
from geomcheck import compute_all

for p in sys.argv[1:]:
    t = time.perf_counter()
    r = compute_all(p)                      # full resolution, default thresholds
    nm = round(r["nonmanifold_edge_frac"] * r["n_edges"])
    print(p, r["n_faces"], r["watertight"], nm, r["n_components"], r["n_small_components"],
          f'{100*r["self_intersect_face_frac"]:.3f}', f'{100*r["hidden_surface_frac"]:.2f}',
          f'{100*r["thin_frac_0.005"]:.2f}', f"{time.perf_counter()-t:.1f}s")

Real output from my run (formatted version of the script above), 8-core x86_64, CPU only.

FileClosed solid?Edges with 3+ facesParts / under 1% areaCrossing faces, full res → decimatedNever-visible surfaceSurface under 0.005 diag
chestyes01 / 00.001% → 0.000%0.63%0.00%
dragonyes01 / 00.009% → 0.000%0.01%0.04%
workshopyes07 / 50.005% → 1.170%11.67%30.63%
guardianyes01 / 00.012% → 0.030%0.06%0.66%
radiono11 / 00.011% → 0.020%0.00%0.22%
marble bustno22 / 10.041% → 0.190%2.72%1.62%

"Decimated" is the same run with decimate_to=10000 (Log 4). None had an open boundary loop. compute_all took 1.5–1.9 s per model in this first pass (Log 7 has cleaner timings). Running it a second time on the bust returned an identical dictionary.

Log 3: what it actually caught

A buried set of blocks in the workshop. "7 pieces, 5 small" sounds like floating debris, so I coloured each connected component. Piece #2 is the anvil: separate but intentional, 4,740 faces. Pieces #3–#7 are five small pad blocks at floor level under the wall corners. Then I crossed the component labels with the per-face visibility flags from geomcheck.visualize.face_flags. All 2,412 faces of those five blocks are hidden, and none of them self-intersects. They sit completely inside the walls. You could delete them and nobody would see a difference, but a slicer would get five extra solids nested inside the walls.

Real render (geomcheck.visualize). Orange = the five small pieces, purple = the anvil, grey = main shell.

Non-manifold edges you can't see. The radio and the bust are the two "not watertight" results, with zero holes. The cause is 1 and 2 edges where four faces meet: one near the radio's lower knob, two in the bust's hair. I located them by listing edges used by more than two faces and projecting their midpoints onto the render.

Real render with projected markers. The rings mark positions; the edges themselves are too small to see at this size.

A "floater" that is a tree. The tiny extra part on the bust, a 180-face shell covering a quarter of a percent of the surface, turned out to be a cypress inside the open head. A blanket "delete small components" step would remove something the model is supposed to have. geomcheck counts pieces; deciding what they are is still your job.

Thin walls with a scale problem. 30.63% of the workshop's surface is thinner than 0.005 of its diagonal. Scaled so its longest side is 100 mm, that threshold is 0.83 mm. Thickness is a ratio until you pick a print size.

Log 4: the numbers move if you decimate first

compute_all has a decimate_to option (the paper used 10,000 faces for most of its datasets). I ran all six both ways:

Real data from my run.

Thin-wall and hidden shares barely moved (workshop 30.63% → 29.98% thin, 11.67% → 11.53% hidden). Self-intersection did move: workshop 0.005% → 1.17%, bust 0.041% → 0.19%. Decimating created crossing faces that the original mesh did not have. Pick one setting and keep it, and never compare a decimated number with a full-resolution one. The docs warn about this too, and now I have a concrete case.

Log 5: how much do the thresholds matter?

Every flag in geomcheck is a number compared with a cut-off, and the cut-offs live in one module-level dict, geomcheck.CONFIG. The defaults are reasonable guesses, not laws: thin below 0.005 of the diagonal, "small" below 1% of area, 64 view directions for visibility. I wanted to see how far each verdict slides if you pick a different cut-off, so I re-ran the three ray and component steps with the package's own internal functions at six settings each. Same welded mesh, same 20,000 samples, same seed. Only the threshold changes.

Real data from my sweep. Dotted lines mark the defaults.

Thin-wall threshold. Share of surface samples flagged thin:

Threshold (× diagonal)0.0010.00250.005 (default)0.010.020.04
chest0000.32%0.43%2.12%
dragon0.04%0.04%0.04%0.78%5.76%18.39%
workshop2.09%21.42%30.63%35.81%50.75%71.41%
guardian0.06%0.31%0.67%1.10%1.57%7.35%
radio000.22%2.71%6.45%16.08%
marble bust0.05%0.41%1.62%4.29%11.18%18.54%

The workshop jumps from 2% to 21% between 0.001 and 0.0025 of the diagonal. That cliff is the wall plates. Among the samples where the ray found the far side of a wall, the workshop's 25th-percentile thickness is 0.0027 of the diagonal. Scaled to 100 mm on the longest side, the diagonal is 165.5 mm, so that is 0.45 mm. When I later cut the same model with a plane and measured the back wall directly, I got 0.46–0.47 mm. So the probe's number is a real wall, not a sampling artefact.

The other models behave differently. Dragon and radio sit near zero until 0.01 and then climb fast, because they have fins, knobs and grille bars that are thin only at a coarse cut-off. The ordering of the six also changes with the threshold. At the default, the bust (1.62%) is second; at 0.04, the bust, dragon and radio are within half a point of each other (18.54%, 18.39%, 16.08%). If you sort a batch by "thinness", the threshold you pick decides the order.

Ray direction count. Share of surface samples flagged as never visible:

Directions8163264 (default)128256
chest2.00%0.97%0.74%0.63%0.53%0.52%
dragon1.18%0.10%00.02%00
workshop34.65%21.63%14.44%11.67%10.21%9.57%
guardian3.19%0.61%0.11%0.06%0.01%0.01%
radio0.97%0.44%0.13%0.01%00
marble bust9.95%5.71%3.91%2.72%2.00%1.64%

A sample counts as hidden when none of the tested directions escapes, so with nested direction sets more directions could only lower the figure. The dragon shows that geomcheck's sets are not nested: 0 at 32 directions, 0.015% at 64, 0 again at 128. The Fibonacci direction sets are generated fresh for each count, so 64 directions are not a superset of 32. A handful of samples in deep crevices escape along one of the 32 directions and none of the 64. Small effect, but it means the curve is not strictly monotonic.

Cost scales with the count, as expected. On the workshop the visibility step took 0.075 s at 8 directions, 0.35 s at 64 and 1.21 s at 256. Going from 64 to 256 directions cuts the workshop's figure by about 2 points for 3.4× the time. The default looks like a fair trade, but do not compare a hidden-surface figure across tools that use different counts.

Small-piece threshold. Count of pieces below the area cut-off:

Area cut-off0.1%0.25%0.5%1% (default)2%5%
workshop025556
marble bust001111
other four000000

This one is the most stable of the three. Between 0.5% and 2% nothing changes. At 5% the anvil (2.32% of the workshop's area) joins the "small" list, which is wrong for anyone treating "small" as "junk". Below 0.5% the bust's tree (0.25%) drops out. Neither threshold separates "buried junk" from "small but wanted". The visibility cross-reference from Log 3 does that, and size alone does not.

Log 6: cross-checking against other tools

A checker you cannot cross-check is a checker you have to trust. I installed four other tools on the same Debian box and ran each on the same six files: Blender 4.3.2 with the 3D-Print Toolbox extension 1.4.1, driven headless with blender --background --python; PyMeshLab's get_topological_measures; ADMesh 0.98.5; and the part count from PrusaSlicer 2.9.2's --info. For Blender I scaled each model to 100 mm and called the toolbox's own check functions from a script, with its default 0.1 mm zero-area threshold and 1 mm thickness.

The first lesson came before any comparison. Run on the file exactly as Blender's glTF importer delivers it, the toolbox reported 47,560–110,048 non-manifold edges, 1,544–9,396 shells and 10,499–22,460 intersecting faces per model. Those are the UV seams again (Log 2's welding note), the same numbers PyMeshLab gives on the raw mesh. After a Merge by Distance at 1e-6 m, the comparison became meaningful:

QuantitygeomcheckBlender toolbox (after merge)PyMeshLab (welded)ADMesh / PrusaSlicerAgree?
Pieces, chest / dragon / workshop / guardian1 / 1 / 7 / 11 / 1 / 7 / 11 / 1 / 7 / 11 / 1 / 7 / 1yes, all five
Pieces, radio / bust1 / 21 / 21 / 22 / 4STL tools split at non-manifold edges
Non-manifold, radio / bust1 / 2 edges3 / 64 / 8not reportedsame models, different counting
Non-manifold, other four000not reportedyes
Self-intersecting faces (6 models)2 / 13 / 7 / 18 / 17 / 612 / 10 / 7 / 15 / 10 / 34not runnot runsame ranking, toolbox lower
Degenerate faces0–34,979–34,694 "zero faces"not run0 degeneratetoolbox threshold is absolute
Thin, share of faces under 1 mm (toolbox) / of surface under 0.005 diag (geomcheck)0 / 0.04 / 30.6 / 0.7 / 0.2 / 1.6%0.2 / 0.7 / 33.9 / 3.3 / 18.7 / 11.9%——workshop yes, radio no

Where they agree: pieces and the location of the problems. Every tool that welds found the same component counts, and every tool that counts non-manifold edges pointed at the radio and the bust and nothing else. PyMeshLab's count is exactly four times geomcheck's: 4 vs 1 and 8 vs 2. I have not traced Blender's 3 and 6 back to the code, so I will not guess.

Where they disagree, three reasons came out of the data:

  1. Counting units. For non-manifold edges, geomcheck counts unique edges, while the other tools apparently count something per incident face or per half-edge. The PyMeshLab ratio of exactly 4 hints at that, but it is an inference. Compare "is it zero?" across tools, not the count.
  2. Absolute vs relative thresholds. The toolbox's degenerate check uses 0.1 mm as the zero threshold. At a 100 mm print size, many perfectly valid small triangles are under that, hence thousands of "zero faces" on meshes where geomcheck and ADMesh found 0–3 real degenerates. Scale the model to 1 m and the toolbox count would drop. geomcheck normalises everything to the diagonal, so it has the opposite problem: its thresholds mean nothing in millimetres until you supply a size.
  3. Different definitions of "thin". The radio is the clearest case: the toolbox flagged 28,121 faces (18.7%) at 1 mm, geomcheck flagged 0.22% of its surface below 0.69 mm and 2.71% below 1.37 mm. In the toolbox source (bmesh_check_thick_object in lib.py), the thickness test casts a ray backwards from points on each face and counts any hit within the distance. geomcheck only counts a hit if the ray leaves through a surface facing away from it, meaning a real exit through the opposite wall. Grille slots and stacked knob details produce near hits that are not walls. I would trust the slicer over either number for print decisions. The workshop, where both agree, really did slice to single-extrusion walls.

What only one tool covered: hidden surface. None of the other four reports never-visible area, so the workshop's 11.67% has no cross-check beyond my own per-component test in Log 3. The toolbox's thickness pass took about 5–6 s per model on its own, roughly 3–4× geomcheck's whole run.

Log 7: runtime vs face count

For batch use the question is how this scales. I decimated each model with geomcheck's own PyMeshLab path to 5k, 10k, 25k, 50k and 100k faces, kept the full ~150k file as the last point, and timed every probe family. One run per point on the same 8-core CPU, so read these as rough.

Real timings from my run, geomcheck 0.1.0, CPU only.

FacesAll probes (range over six models)Decimation step alone
5k0.14–0.30 s2.04–2.21 s
10k0.18–0.37 s2.01–2.38 s
25k0.30–0.51 s1.78–1.91 s
50k0.51–0.74 s1.51–1.63 s
100k0.94–1.19 s0.98–1.04 s
~150k (no decimation)1.40–1.69 s0

Two things surprised me.

First, decimating to save time does not save time at this size. Reaching 5k faces costs about 2.1 s, more than probing the full 150k mesh (1.4–1.7 s). Decimation gets slower the further you reduce, presumably because the quadric collapse has more edges to remove. The decimate_to option is there to make results comparable across datasets with very different face counts, which is why the paper used it. It is not a speed knob for files this size.

Second, the per-family split. On the workshop at full resolution, topology took 0.53 s, the ray probes 0.54 s, self-intersection 0.47 s, roughness 0.13 s and triangle quality 0.008 s. Topology and self-intersection grow roughly in line with faces. The ray probes use a fixed 20,000 samples, so their time grows slowly: 0.11–0.27 s at 5k faces and 0.28–0.54 s at full resolution. I assume the growth is Embree building and walking a bigger acceleration structure, but I did not profile it. Extrapolating the straight-ish line, a 1M-face scan would take on the order of ten seconds. That is an extrapolation; I have not run one.

Decimation also changes answers, as Log 4 showed. The full curve for the workshop's self-intersection share: 3.96% at 5k, 1.17% at 10k, 0.204% at 25k, 0.032% at 50k, 0.017% at 100k, 0.005% at full resolution. The bust showed 1.08% at 5k. Hidden and thin shares moved by at most about 3 points (workshop thin 27.8% at 5k vs 30.6% at full; hidden 10.7% vs 11.7%). If you must decimate, self-intersection is the number you can no longer trust.

Known limitations

Collected from the logs above and from reading the code. None of these is hypothetical; each one showed up on these six files or in the source.

Roadmap (what I would like to add, no dates)

These are directions, not commitments, roughly in the order the logs above made me want them:

  1. Publish the package to PyPI so the documented install command works.
  2. An optional scale_mm argument, so thin-wall output can be reported in millimetres at a given print size.
  3. Per-edge and per-face location output for non-manifold edges and self-intersections, so a report can say where, not just how many.
  4. A "buried piece" flag that combines the component and visibility probes, which is the check that separated the workshop's pads from the bust's tree.
  5. Nested direction sets for the visibility probe, so more directions can never raise the hidden share.
  6. Passing thresholds as function arguments instead of a global dict.
  7. A command-line entry point that prints the table from Log 2 and writes JSON.

If you try it on your own meshes and something disagrees with another tool, an issue on the repository with the file and both outputs is the most useful thing you can send.

What it can't measure

The probes come from my paper, Geometric validity is not perceptual quality: what deterministic probes can and cannot measure in generated 3D assets (doi:10.5281/zenodo.22995915). The parts that matter for anyone using geomcheck as a checker:

So: it is a gate (is this file printable, riggable, exportable?), not a score. It ignores textures, materials and whether the model matches its prompt. A perfect-looking result like the dragon tells you the geometry is clean, not that the dragon is good.

Illustration, not test output: the probes see defects in the lens. Whether the model looks good is a separate judgment.

How this was tested

Links