Preprint · 2026

Shape-Bayes

Bayesian Inference of Structured Shapes under Visual Ambiguity

♣,1 Tosh Brown2 Michel Valstar2,3
1Amsterdam, The Netherlands 2BlueSkeye AI 3University of Nottingham, U.K.

♣Independently conceived and developed by the first author during their tenure at BlueSkeye AI.

Trust visual evidence where it is reliable. Rely on the shape prior where it is ambiguous.
The core intuition of Shape-Bayes
VR headset: the deterministic prediction breaks where the face is hidden.
Eyes and brows hidden
Popcorn tub: the deterministic prediction breaks where the face is hidden.
Mouth and chin hidden
Burger: the deterministic prediction breaks where the face is hidden.
Mouth and jaw hidden
Fig. 1Point estimate versus posterior under occlusion. Gold: the prediction of HR-noSA, a state-of-the-art deterministic model, which distorts when the face is hidden. Cyan: the Shape-Bayes posterior — a distribution of plausible face shapes. The animation sweeps through the model's uncertainty, showing that only the hidden landmarks move. Ellipses indicate where the model is most uncertain. AI-generated faces.

Abstract

Perceiving structured shapes, such as human faces, from pixels is an inherently ambiguous task in real-world conditions. Yet shape inference is largely posed as a deterministic regression task predicting fixed spatial coordinates. We find that deterministic regression is brittle when visual evidence is ambiguous or incomplete: under severe occlusions, deterministic models exhibit structural collapse, predicting incoherent shapes or reverting to generic averages.

To address this, we introduce Shape-Bayes, a probabilistic framework that couples uncertainty-aware visual perception with Bayesian shape reasoning. Rather than forcing point estimates, Shape-Bayes dynamically weights visual evidence against geometric priors to infer a structurally valid shape posterior. By guaranteeing complete structural integrity, Shape-Bayes achieves 100% In-Distribution Rate (IDR) and establishes a new state of the art for robust 2D face shape regression under severe occlusion.

100%
In-distribution rate for all six base models on all three occluded benchmarks. HR-noSA alone reaches 66.6% on 300W.
7.56→1.97
Expected calibration error of HR-noSA on occluded 300W, after the Bayesian update.
−21.2%
Relative error on occluded landmarks for PiPNet on 300W, with failure rate down 13.7 points.
9.03→6.60
minNME over 100 posterior samples versus the deterministic prediction, on occluded landmarks of occluded 300W (oracle selection).
The problem

Structural Collapse Under Severe Occlusion

Deep regressors are highly accurate on unoccluded faces. But when a region has no meaningful visual evidence, a model that must output one point per landmark still has to guess. It either draws a topologically invalid shape or falls back to a generic average. Even HR-noSA, the state-of-the-art deterministic face shape regressor (Yang & Yeh, ICCV 2025), keeps only 66.6% of its 300W predictions inside the valid shape space once faces are occluded.

Even when its guess is plausible, a deterministic model gives one answer where the image supports several, and nothing tells the user that the answer was a guess.

Drag the dividers to compare the same images: the deterministic base model on the left, the Shape-Bayes posterior with its uncertainty on the right.

Deterministic landmarks on a face partly covered by a baseball glove; the mouth and nose contours are tangled.
Shape-Bayes posterior on the same face: a coherent face shape with uncertainty halos around hidden landmarks.
Deterministic Shape-Bayes
Deterministic landmarks on a face wearing a VR headset and partly covered by a hand; eyes and brows are distorted.
Shape-Bayes posterior on the same face: plausible eyes, brows and jaw, with wide uncertainty under the headset.
Deterministic Shape-Bayes
Deterministic landmarks on a face covered by a burger; the mouth contour is flattened and displaced.
Shape-Bayes posterior on the same face: a natural mouth and jaw inferred behind the burger.
Deterministic Shape-Bayes

Occluded 300W test images. Left of each divider: HR-noSA (Yang & Yeh, 2025). Right: the Shape-Bayes posterior mean with per-landmark uncertainty.

Left: a deterministic regressor's face landmarks collapse under a VR headset. Right: Shape-Bayes predicts uncertainties, infers a posterior mean with covariance, and samples several plausible shapes.
Fig. 4Robust Bayesian shape inference under severe occlusion. Deterministic regression collapses when forced to guess exact coordinates for hidden regions (left). Shape-Bayes balances the noisy visual evidence against a structural prior, guided by the model's own uncertainty. The result is a distribution of valid shapes, from which multiple plausible hypotheses can be sampled (right).
The idea

Adaptive Arbitration of Visual Evidence and Geometric Prior

Human vision has long been described as Bayesian inference: incomplete, noisy bottom-up observations are reconciled with robust top-down structural priors. Shape-Bayes makes this explicit for shape regression. A base model reports not just where each landmark is, but how sure it is. We measure this "aleatoric uncertainty" intuitively: if an image is rotated or blurred and the model's prediction scatters, the model is uncertain. This distilled uncertainty acts as a spatial gate in a closed-form Bayesian update, deciding how much each coordinate can move the shape.

Schematic: an input image yields an observed shape S′ with uncertainty σ′. This defines a likelihood with precision W and, through a learned prior, a precision Λ. A closed-form Bayesian update combines them into the posterior P(S).
Fig. 5Schematic overview. Given an initial noisy shape estimate, Shape-Bayes balances the image evidence (W) against a data-driven structural prior (Λ) through a mathematical Bayesian update, outputting a complete distribution of probable shapes P(S).
Deterministic regressionpoint estimate
Shape-Bayes posteriormean + 2σ ellipses
Relative evidence weightvisual evidenceshape prior
Hand over mouth21 of 68 landmarks with σ′ > σ₀
Hand over mouth: the deterministic shape under the hand.
Hand over mouth: the Shape-Bayes posterior with 2σ ellipses over the hidden mouth and jaw.
Hand over mouth: landmarks under the hand coloured olive, visible landmarks pink.
Hand over right eye23 of 68 landmarks with σ′ > σ₀
Hand over right eye: the deterministic shape under the hand.
Hand over right eye: the Shape-Bayes posterior with 2σ ellipses over the hidden eye, brow and jaw side.
Hand over right eye: the hidden eye, brow and jaw side coloured olive, visible landmarks pink.
Hand over left eye22 of 68 landmarks with σ′ > σ₀
Hand over left eye: the deterministic shape under the hand.
Hand over left eye: the Shape-Bayes posterior with 2σ ellipses over the hidden eye, brow and jaw side.
Hand over left eye: the hidden eye, brow and jaw side coloured olive, visible landmarks pink.
Fig. 2Evidence and prior, weighted per landmark. One subject covers her mouth, her right eye and her left eye. All panels are model outputs from HR-noSA, followed by the Shape-Bayes update. Left: the deterministic prediction — a single shape with no uncertainty estimate. Middle: the Shape-Bayes posterior mean with uncertainty ellipses (covariance scaled ×5 for visibility). Right: each landmark is coloured by how much its position is determined by visual evidence (pink) versus the shape prior(olive). Pink landmarks are mostly determined by what the camera sees; olive landmarks, hidden under the hand, are mostly determined by what a face should look like. AI-generated photographs.

Posterior precision is a sum of information

Shapes live on a PCA manifold, \(\mathbf{S} = \bar{\mathbf{S}} + \mathbf{P}\mathbf{b}\). With a Gaussian likelihood and a Gaussian prior over the coefficients \(\mathbf{b}\), the posterior is exactly Gaussian. Its precision is the information from the image plus the information from the prior.

$$\tilde{\boldsymbol{\Sigma}}_b^{-1} \;=\; \underbrace{\htmlClass{lik}{\mathbf{P}^{\top}\mathbf{W}\,\mathbf{P}}}_{\text{observation}} \;+\; \underbrace{\htmlClass{pri}{\boldsymbol{\Lambda}_{\text{prior}}}}_{\text{prior}}$$ $$\htmlClass{post}{\mathbf{b}_{\mu}} \;=\; \tilde{\boldsymbol{\Sigma}}_b\,\mathbf{P}^{\top}\htmlClass{lik}{\mathbf{W}}\,(\mathbf{s}' - \bar{\mathbf{S}})$$ $$\htmlClass{lik}{\mathbf{W}} = \operatorname{diag}\!\Big(\frac{1}{\boldsymbol{\sigma}'^{2} + \epsilon}\Big)$$
Evidence from the image Learned shape prior Posterior

Where the face is hidden

\(\sigma' \to \infty\), so \(\mathbf{W} \to 0\). The observation term contributes no information there, and those landmarks are determined by the prior together with the visible landmarks. The estimate remains on the PCA shape manifold by construction.

Where the face is visible

\(\sigma' \to 0\), so \(\mathbf{W}\) dominates. Shape-Bayes becomes almost an identity mapping, so accuracy is kept: 300W NME goes from 2.86 to 2.81 on clean images.

Posterior sampling

Probabilistic Posterior Sampling

When an occluder covers a facial region, the visual evidence is insufficient to resolve ambiguities (e.g., whether the mouth is open or closed). A deterministic regressor still outputs a single shape with no measure of its reliability. In contrast, Shape-Bayes infers a closed-form Gaussian posterior over the PCA shape coefficients. Its principal directions of variance align with the unobserved shape dimensions: moving along the leading direction opens and closes the hidden mouth while the visible eyes stay fixed.

Input. Burger hides the mouth and jaw; the box marks the zoomed region.
InputBurger hides the mouth and jaw; the box marks the zoomed region.
Deterministic prediction, zoomed: a tangled mouth behind the burger.
DeterministicSingle estimate, distorted under the occluder
Twelve Shape-Bayes posterior samples: they spread only where the face is hidden.
12 posterior samplesVariation concentrated in the occluded region
Two Shape-Bayes hypotheses overlaid, cyan and white. Mouth closed versus mouth open.
Two hypotheses · cyan vs whiteMouth closed vs open: inner-lip gap 1.0% vs 9.8% of the inter-ocular distance. The eyes move 0.1%.
Input. Headset hides the eyes and brows; the box marks the zoomed region.
InputHeadset hides the eyes and brows; the box marks the zoomed region.
Deterministic prediction, zoomed: crossed, broken eyes and brows behind the headset.
DeterministicSingle estimate, distorted under the occluder
Twelve Shape-Bayes posterior samples: they spread only where the face is hidden.
12 posterior samplesVariation concentrated in the occluded region
Two Shape-Bayes hypotheses overlaid, cyan and white. Brows and eyes in different positions.
Two hypotheses · cyan vs whiteBrows and eyes shift by 6.1% and 3.9% of the inter-ocular distance. The visible mouth moves 0.1%.
Input. The popcorn tub hides the mouth and chin; the box marks the zoomed region.
InputTub hides the mouth and chin; the box marks the zoomed region.
Deterministic prediction, zoomed: a tangled mouth behind the popcorn tub.
DeterministicSingle estimate, distorted under the occluder
Twelve Shape-Bayes posterior samples: they spread only where the face is hidden.
12 posterior samplesVariation concentrated in the occluded region
Two Shape-Bayes hypotheses overlaid: mouth open in cyan, mouth closed in white.
Two hypotheses · cyan vs whiteMouth open vs closed: inner-lip gap 11.7% vs 2.5% of the inter-ocular distance. The eyes move 0.1%.
Fig. 3Posterior samples and structural hypotheses. Each row is zoomed to the occluded region (box on the input). Shape-Bayes infers a distribution of possible shapes. From this, we can sample random plausible shapes, or we can extract the most prominent hypotheses—such as an open mouth versus a closed mouth—by finding the directions of greatest uncertainty. AI-generated faces; HR-noSA base predictions.

Posterior samples improve geometric accuracy

On occluded 300W, the posterior sample closest to the ground truth among 100 has an error of 6.60, against 9.03 for the deterministic prediction and 7.91 for the posterior mean. The minimum keeps decreasing with the number of samples, to below 6 at 1,000 (Fig. 9). Because selection uses the ground truth, this measures the quality of the hypothesis set, not of a deployable estimator. Realising the gain requires an external signal that discriminates between hypotheses, such as later video frames, a second view or a task-specific verifier.

Error on hidden landmarks (NME) ↓

Deterministic prediction9.03
Shape-Bayes posterior mean7.91
Best of 100 posterior samples6.60

Occluded 300W, HR-noSA base model. Best-of-N picks the sample using the ground truth, so it measures how good the hypothesis set is.

Method

Three components, one exact update

Inference flow: an uncertainty-aware base regressor outputs landmarks and uncertainties; these define a likelihood precision W and, through a Transformer, a prior precision Λ in PCA space; a closed-form Bayesian update yields the posterior N(b_μ, Σ_b), from which shapes are sampled and projected.
Fig. 6Inference flow. (1) A base regression model predicts landmarks S′ and uncertainties σ′, defining the likelihood precision W. (2) A Transformer maps these observations to an adaptive prior precision Λprior in PCA space. (3) A closed-form Bayesian update computes the shape posterior 𝒩(bμ, Σ̃b); probable shapes are obtained by sampling and projecting, S = S̄ + Pb.
Step 1
Distillation pipeline: a masked face is rotated from −45° to 45°, a frozen teacher predicts landmarks for each view, predictions are inverse-transformed back, and their spread gives per-landmark uncertainty, large around the mask.
Fig. 7Distilled uncertainty supervision. Teacher predictions on affine-augmented views are mapped back to the original frame; their spread is the aleatoric target.

Uncertainty from visual feature stability

A visible, salient region has robust local features: its predicted coordinates barely move when the image is rotated, scaled or blurred. An occluded region has no discriminative evidence, so the model extrapolates from weak global context, and those guesses scatter under the same transformations.

Shape-Bayes turns this into supervision without manual labels. A frozen teacher predicts shapes for M augmented views. The predictions are warped back with the inverse transforms, and their root-mean-square deviation from the unaugmented prediction becomes the target uncertainty:

$$\boldsymbol{\sigma}'_{\text{target}} = \Big(\tfrac{1}{M}\textstyle\sum_{i=1}^{M} \big(\mathbf{S}_i - \mathbf{S}'_{\text{target}}\big)^2\Big)^{1/2}$$

The base model learns to output coordinates and uncertainties together, so at inference it needs a single forward pass. The scheme works with any architecture: heatmap, coordinate or Transformer regressors. Models that already predict uncertainty, such as LUVLi, need no distillation at all.

M = 30 viewsrotation ±35°scale 1.1–1.4blur · flips
Step 2

Input-conditioned structural prior

A static PCA prior is often insufficient for large deformations such as head pose. Shape-Bayes employs a lightweight Transformer encoder conditioned on the observations: each of the N landmarks acts as a token, \(\mathbf{X} = [\mathbf{S}', \boldsymbol{\sigma}'] \in \mathbb{R}^{N\times 4}\).

Through self-attention, high-confidence visible regions inform the structural placement of occluded landmarks. The encoder outputs the diagonal of the prior precision matrix over the latent shape coefficients:

$$\htmlClass{pri}{\boldsymbol{\Lambda}_{\text{prior}}} = \operatorname{diag}\big(\exp(E_{\phi}(\mathbf{X}))\big) \in \mathbb{R}^{K\times K}$$

Prior \(\mathcal{N}(\mathbf{0},\, \boldsymbol{\Lambda}_{\text{prior}}^{-1})\) over \(\mathbf{b}\). The exponential keeps it strictly positive.

5 encoder layersd = 648 heads429K parametersPCA rank K captures >99% variance
Step 3

Differentiable closed-form Bayesian solver

The product of the likelihood and prior yields an exact Gaussian posterior over shape coefficients. Because the PCA manifold is low-dimensional, inverting the K×K precision matrix is computationally trivial and fully differentiable, enabling end-to-end training.

The posterior mean serves as the nominal shape prediction, while samples drawn from the posterior represent plausible alternative hypotheses. By construction, every sample lies on the valid PCA shape manifold.

$$p(\mathbf{b}\mid\mathbf{s}',\mathbf{X}) = \mathcal{N}\big(\htmlClass{post}{\mathbf{b}_{\mu}},\, \htmlClass{post}{\tilde{\boldsymbol{\Sigma}}_b}\big)$$ $$\mathbf{b} \sim \mathcal{N}\big(\mathbf{b}_{\mu}, \tilde{\boldsymbol{\Sigma}}_b\big) \;\Rightarrow\; \mathbf{S} = \bar{\mathbf{S}} + \mathbf{P}\mathbf{b}$$
Training

Decoupling spatial prediction and uncertainty

During training, the base model is frozen and the PCA basis is precomputed; only the prior Transformer is optimized using occlusion-augmented shapes. An MSE objective on the posterior mean anchors the spatial coordinates, while a negative log-likelihood (NLL) objective calibrates the posterior.

A stop-gradient applied to the mean within the NLL term decouples these objectives: MSE governs spatial localization, while NLL governs uncertainty estimation. This prevents the model from artificially reducing the NLL loss by shifting the mean rather than properly calibrating its variance.

$$\mathcal{L} = \lambda_{\text{MSE}}\,\lVert \mathbf{b}_{\mu} - \mathbf{b}_{\text{gt}} \rVert^2 + \lambda_{\text{NLL}}\,\mathcal{L}_{\text{NLL}}$$ $$\begin{aligned}\mathcal{L}_{\text{NLL}} = {}& \tfrac{1}{2}(\mathbf{b}_{\text{gt}} - \bar{\mathbf{b}}_{\mu})^{\top}\tilde{\boldsymbol{\Sigma}}_b^{-1}(\mathbf{b}_{\text{gt}} - \bar{\mathbf{b}}_{\mu}) \\ & - \tfrac{1}{2}\log\det\tilde{\boldsymbol{\Sigma}}_b^{-1}\end{aligned}$$

\(\bar{\mathbf{b}}_{\mu} = \operatorname{sg}[\mathbf{b}_{\mu}]\) is a stop-gradient on the mean.

One masked face through the pipeline: input image, base model shape prediction with errors, its uncertainty, the Bayesian posterior, two posterior samples with open and closed mouth, and an ensemble of five samples.
Fig. 8One image through the pipeline. Input, base model prediction and its aleatoric uncertainty, Bayesian posterior mean, and shapes sampled from the posterior. Under the mask, the posterior keeps both an open-mouth and a closed-mouth hypothesis, which the image cannot decide between.
Results

Structural Validity and Occlusion Robustness

Shape-Bayes is evaluated on the 300W, WFLW and COFW test sets under a dynamic occlusion protocol: synthetic wearable masks and a library of more than 80 real-world objects overlaid on the images, without changing the ground-truth shapes. It is attached to six state-of-the-art base models spanning heatmap, coordinate and Transformer paradigms.

Base model+ Shape-Bayes

In-distribution rate (IDR) ↑

% of predicted shapes inside the 95% hyperellipsoid of the shape manifold

Error on occluded landmarks (NMEocc) ↓

Normalized mean error, % of inter-ocular distance

Base model IDR ↑ NMEocc ↓ NMEall ↓ minNME@100 ↓ FR10 ↓ AUC10 ↑

Each cell shows base → +Shape-Bayes, with the relative change for error metrics. NATIVE marks models that already output uncertainty and need no distillation. A dash means the model was not evaluated on that benchmark. minNME@100 takes the posterior sample closest to the ground truth out of 100, so it measures the quality of the hypothesis set rather than of a single prediction.

Preserving structural integrity without over-smoothing

Prior-driven shape regularization typically trades accuracy on visible landmarks for global structural validity. Shape-Bayes avoids this trade-off. High-confidence visible landmarks strongly anchor the latent shape, while uncertain landmarks are down-weighted and inferred via the structural prior.

These gains are not traded between regions: on occluded 300W, error decreases across all facial parts.

Part-wise NMEocc on occluded 300W ↓

Base model: HR-noSA

Performance on unoccluded test sets

In the absence of occlusion, predicted uncertainties approach zero, causing the observation precision to dominate the prior. Shape-Bayes naturally reduces to an identity mapping, preserving or slightly improving baseline accuracy. Sampling 100 shapes from the posterior and selecting the oracle minimum yields the lowest error across all three benchmarks, confirming that the posterior assigns high probability mass to accurate geometric hypotheses.

Method 300W WFLW COFW
AWing 3.07 4.36 4.94
LUVLi 3.23 4.37 –
ViTPose 3.02 4.26 4.85
ADNet 2.93 4.14 4.68
HIH 3.09 4.08 4.63
SLPT 3.17 4.14 4.79
STAR 2.90 4.03 4.62
OccFace 2.83 3.77 4.53
HR-noSA 2.86 3.97 4.53
+ Shape-Bayes (mean) 2.81 3.94 4.50
+ Shape-Bayes (minNME@100) 2.79 3.72 4.23

NMEall ↓ on the original test sets. Bold marks the best single prediction per benchmark. The minNME row selects among samples using the ground truth and is reported separately.

Ablation study

Configuration NME ↓ minNME@100 ↓ FR ↓ AUC ↑

Unweighted least-squares PCA projection yields minimal improvement, as corrupted landmarks act as outliers that distort confident predictions. Weighting the projection by the distilled inverse variance—equivalent to Mahalanobis distance minimization—provides the majority of the performance gain. The adaptive prior accounts for the remaining improvement by accommodating input-specific variations such as head pose.

Because the output is a full probability distribution, drawing additional samples increasingly yields more accurate geometric hypotheses, significantly outperforming the deterministic baseline.

minNME as a function of the number of posterior samples, from about 8 at 5 samples to about 5.9 at 1000, below the deterministic baseline of 9 for all three variants; Shape-Bayes is lowest throughout.
Fig. 9minNME@N on occluded 300W versus the number of posterior samples N.

Occluded 300W, HR-noSA base model, metrics on occluded landmarks.

Calibration

Posterior calibration

Under out-of-distribution occlusion, baseline models are typically overconfident, predicting uncertainties that underestimate the true error. Propagating aleatoric uncertainty through the prior-constrained Bayesian update appropriately increases the posterior variance where visual evidence is degraded, consistently lowering the expected calibration error (ECE) across all evaluated architectures.

Three calibration plots on occluded 300W: (a) predicted uncertainty versus observed error with Spearman correlation 0.67; (b) reliability diagram close to the diagonal, ECE 1.966; (c) sparsification curve tracking the oracle.
Fig. 10Posterior calibration on occluded 300W. (a) Predicted uncertainty correlates with observed error (Spearman ρ = 0.67). (b) The reliability diagram stays close to perfect calibration (ECE = 1.97). (c) Rejecting the most uncertain predictions isolates accurate shapes, and the sparsification curve tracks the oracle.

Expected calibration error (ECE) ↓

Occluded 300W. *LUVLi uses its own native uncertainty.

Base model NLL ↓ ECE ↓
Cost

Computational cost

As the base network remains frozen, Shape-Bayes introduces only a lightweight Transformer and a K×K matrix inversion. Consequently, the backbone continues to dominate total inference time.

429K
parameters, 3.2% of the base network
0.10%
of base FLOPs (~30 MFLOPs vs 29.1 GFLOPs)
<0.5 ms
per image on an RTX 3090
<2 MB
inference memory, weights included
Scope

Generalization to other topologies

While evaluated on human faces—characterized by extreme rigid and non-rigid deformations, strict anatomical constraints, and abundant occlusion—the proposed Bayesian solver is topology-agnostic. It requires only a set of landmarks, their associated uncertainties, and a shape basis.

The framework naturally extends to other structured shape inference tasks under ambiguity, such as body and hand pose estimation, medical anatomical landmarking, and lane geometry extraction.

Citation

BibTeX

@article{tellamekala2026shapebayes,
  title   = {Shape-Bayes: Bayesian Inference of Structured Shapes
             under Visual Ambiguity},
  author  = {Tellamekala, Mani Kumar and Brown, Tosh and Valstar, Michel},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}