3D bone geometry

From two X-rays, an AI model reconstructs a 3D femur without CT scans.

Recovering a 3D femur from two X-ray views without CT or a neural network is a bold constraint, and this pipeline makes it work.

4 min readMachine Learning
From two X-rays, an AI model reconstructs a 3D femur without CT scans.
Reconstructing 3D bone geometry from 2 X-ray silhouettes using a statistical shape model + differentiable rendering [P]

The most honest thing about this reconstruction pipeline is that it nearly failed. The author of this reconstructing 3D bone geometry from 2 X-ray silhouettes post did not present a clean win. They presented a war story. And that is precisely why it is worth reading. The goal was simple: take two X-ray silhouettes of a distal femur, fit a statistical shape model to them, and produce a patient-specific 3D bone. No CT, no neural network, no massive training set. Just a PCA model built from 50 CT-derived meshes, a differentiable renderer, and a whole lot of suffering.

The suffering was real. The correspondence step, which maps points between the model and the target, broke every method they tried. KD-tree nearest neighbor gave a roughness score 50.7 times worse than the CT ground truth. CPD was better but still 28.2 times off. BCPD was worse. FilterReg would not even run. It was only when they got ShapeWorks working that the numbers dropped to 3.3 times, finally passing the 5x acceptance gate they had set before testing. That is a brutal sequence. But it is also the most instructive part of the whole post. The takeaway is not that ShapeWorks is the answer. The takeaway is that in reconstruction pipelines, the shape model is often the easy part. The correspondence is where ideas go to die.

The validation results tell a similar story. On five held-out femurs, the pipeline achieved 0.86 to 1.43 mm error on targets within the model's coverage. That is genuinely good. But the two extreme cases failed, not because the optimizer was weak, but because the model simply did not have the shape coefficients to represent those bones. Mode 1 did not extend far enough. The optimizer cannot recover a coefficient the model does not support. And here is where the author made a smart observation: the bridge ICP alignment was also poor on those cases, with an inlier fraction of 0.6. That means the alignment error was actually larger than the shape fitting error. The geometry was wrong before the fitting even began.

The most subtle finding, though, was the sigma annealing endpoint. The soft rasterizer in PyTorch3D has a blur parameter, and it turns out that hardcoding a constant tuned on one statistical shape model caused an 87x accuracy degradation on another. The fix was to tie the endpoint to the camera extent times 1e-4. That is the kind of detail that does not make it into papers, but it is exactly the kind of thing that determines whether a method works in practice. It also echoes a lesson from the whitetree library for Mahalanobis nearest-neighbour search that we covered earlier: the difference between a method that works and one that falls apart is often the sensitivity to a single hyperparameter.

What would we tell a reader who is considering this approach? First, expect the correspondence step to consume most of your time. Budget for it. Second, do not trust a shape model to generalize beyond its training distribution. The acceptance gate caught that, and it will catch you too. Third, and this is the concrete takeaway worth quoting: *the renderer's smoothing parameter must be tied to the camera geometry, not tuned as a free constant*. That single detail caused an 87x swing in accuracy. It is the difference between a pipeline that works on one dataset and one that silently breaks on another.

Work is still ongoing on real X-ray validation and automatic segmentation. That is the right next step. But the work here already tells you something important: the path to patient-specific 3D bone models from plain X-rays is not blocked by a lack of deep learning. It is blocked by the unglamorous, unglamorized work of making the model, the renderer, and the correspondence all agree on the same geometry. That is not a revolutionary claim. It is just the truth. And it is a truth worth acting on.

From Machine Learning

Working on a pipeline that recovers a patient specific 3D distal femur from two orthogonal X-ray views (PA + lateral). No CT, no neural network, no massive training set.

approach: build a PCA shape model from 50 CT-derived femur meshes (MedShapeNet), then fit it to two silhouettes using PyTorch3D's soft rasterizer with sigma annealing. 10 shape coefficients, Mahalanobis prior to keep things plausible, Adam optimizer, ~1000 iterations.

Read the original at Machine Learning