Contour · AI Fit Prediction for Fashion E-Commerce

The same body is a Size 29 in one pair of jeans and a Size 28 in a visually identical pair. A size chart cannot feel the difference. We built the engine that can.

Your returns problem is a fit problem, and fit is a physics problem. For a given body (measurements, no photo) and a garment (fabric mechanics plus finished dimensions), Contour computes the circumferential strain at each zone against that fabric's stretch limit, then recommends a size with confidence, or abstains honestly when no size is clean. Agents advise, code decides.

+40 pp

Accuracy lift over the waist-only size-chart incumbent, 100 percent versus 60 percent

Labeled 105-pair synthetic golden set

67.6%

Of pairs given a single high-confidence size, so the shopper need not bracket two

Same 105-pair golden set, confidence at least 0.80

39 of 105

Fit conflicts the waist chart would have shipped, caught before purchase

Same 105-pair golden set

A runnable demo of the mechanism on a synthetic cohort of 15 shoppers and a catalog of 8 garments. The deterministic strain engine and the benchmark run fully offline with no API key; the optional agents are the only part that can use a model.

Four numbers pretending to describe a body they have never measured for fabric

Apparel returns are dominated by fit, and the root cause is mechanical, not cosmetic.

Between 53 and 67 percent of apparel returns are fit-related, and about 63 percent of shoppers bracket, ordering two sizes to send one back (Veriprajna WP34 research, 2026, drawing on industry sources). The carrier eats the reverse logistics, and jeans run among the highest apparel return rates precisely because of stretch, rise, and waist mechanics.

The root cause is simple to state. A size chart gives four 1-D numbers to describe a 3-D body, and it has no idea whether the fabric is zero-stretch raw selvedge or 4-way ponte. So the same waist measurement lands the same shopper in the wrong size on a rigid denim and the right size on a stretch one, and the chart cannot tell those two garments apart. Bracketing is the rational shopper response to a tool that cannot see fabric.

The industry's reflex has been to solve returns with better pictures. A generative virtual try-on makes the garment look right on a body, but a picture cannot tell you whether the thigh exceeds that denim's stretch limit when the shopper sits down. Whether a size fits is a mechanical fact: the circumferential strain at each body zone against that fabric's elastic limit. That is the quantity a chart drops and a render never had, and it is the one Contour computes.

A typed agent crew normalizes the copy; a deterministic engine decides the size

One pipeline, shown a stage at a time. The load-bearing decision lives in plain, unit-tested Python, never in a model.

For a body and a garment, the pipeline runs in order: a fabric-mechanics agent extracts a structured spec from the vendor copy, an adversarial critic verifies it and can block, the strain engine decides the size in pure code, a policy gate picks the clean size or abstains, and a fit-note agent phrases the result. Every number that decides a recommendation is computed before any language model touches the wording.

product copy + body measurements → FabricMechanicsAgent (extract) → FabricMechanicsCritic (verify, can block) → strain engine (decides, pure code) → policy gate (pick clean size, else abstain) → FitNoteAgent (phrase) → audit receipt (JSON) + /api/fit endpoint

The deterministic trust core

The part outside any language model. For each candidate size the strain engine computes zone strain as the body circumference minus the garment's finished circumference over that finished circumference, then compares each zone to the fabric's elastic limit and a per-zone comfort band. The policy gate picks the size with the lowest total per-zone regret, and abstains when no size clears every zone within comfort. This is the load-bearing decision, it lives entirely in plain Python, and 8 unit tests pin the physics, so trust does not depend on the model.

The agents, bounded to advice

On top of the core, a typed crew reads and communicates. The FabricMechanicsAgent turns messy product copy, the title, fiber content, weight, and construction, into a structured fabric spec with a stretch class, an elastic strain limit, and the source phrases that justify each parameter. The FabricMechanicsCritic checks that spec against physics constraints and blocks contradictions. The FitNoteAgent phrases the computed result. The advisors are provider-swappable and default to claude-opus-4-8; with no key the fabric extraction falls back to a cached spec and the deterministic engine still runs live, so the demo records fully offline. Agents advise, code decides.

The three-agent crew, and who has the final say

Stage What it does Basis
FabricMechanicsAgent Reads vendor product copy into a structured fabric spec with stretch class, elastic strain limit, weight, and source phrases. LLM, typed
FabricMechanicsCritic Checks the spec against physics constraints and blocks contradictions, routing to human review with no recommendation. LLM plus deterministic rules
Strain engine Computes per-zone circumferential strain against the fabric's elastic limit for every candidate size. pure code, decides
Policy gate Picks the size with the lowest total per-zone regret, or abstains when no size clears every zone. pure code, abstains
FitNoteAgent Phrases the computed result, for example snug at hip, relaxed at thigh. It never computes the numbers. LLM, cosmetic
Audit receipt and API Writes a replayable JSON receipt and serves the same result at /api/fit for an AI shopping agent. deterministic output

The durable value is the physics plus the audit receipt, not a model that guesses sizing better. Even a perfect language model still needs the fabric mechanics, the body geometry, the per-zone strain calculation, the abstain policy, and a replayable receipt. That infrastructure is the product; the LLM is one swappable advisor inside it, which is why the approach does not age out as models improve.

Five moments, on screen

Every shopper and garment is synthetic: a cohort of 15 bodies and a catalog of 8 garments, with fictional names. Each on-screen number is produced by the engine at runtime, not hard-coded.

The catch: raw selvedge puts Riley in a 29, not the 28 the chart prints

Riley on the Ironside 14oz Raw Selvedge Straight Jean returns a recommended Size 29 at 95 percent confidence, comfortable at every zone. A waist-only chart would say 28, and it would be wrong: at 28 the thigh exceeds this zero-stretch denim's stretch limit. The per-zone strain bars make it legible, green where each zone sits inside the fabric's comfort band. The chart cannot feel the fabric, so it drops the one variable that decides the size.

Contour sizing Riley on the Ironside raw selvedge jean at Size 29, 95 percent confidence, with a note that a waist-only chart would say Size 28 and be wrong here, and per-zone strain bars for waist, hip, thigh, and inseam all marked comfortable.
Riley on raw selvedge: Size 29 at 95 percent confidence, with the waist-only chart's 28 struck through as wrong.

Same body, different fabric, different correct size

Switch the identical shopper to the Driftwood Stretch Slim Jean, a visually similar jean, and the correct size drops to 28 at 90 percent confidence, clean at every zone. Here the waist-only chart happens to also land on 28, because this fabric's stretch forgives the thigh that the raw selvedge punished. That is the whole thesis in one beat: 29 in one jean and 28 in a look-alike, on the same body, because the engine computes how each fabric's stretch limit meets that body.

The same shopper Riley on the Driftwood Stretch Slim Jean returning a recommended Size 28 at 90 percent confidence, with a note that the waist-only chart also lands on 28 here because the stretch denim forgives the other zones.
The same body on stretch denim: Size 28. The correct size changed with the fabric, not the shopper.

When no size is clean, it abstains instead of bluffing

The shopper Jordan on the Ironside raw selvedge jean produces a genuine fit conflict: small sizes fail the thigh, larger sizes run loose at the waist, and on this zero-stretch fabric no size clears every zone. Contour returns an honest abstain at 58 percent confidence, recommends the least-bad size and tells the shopper to bracket or speak to a stylist. The waist-only chart, by contrast, confidently prints 28. Across the labeled golden set the engine abstains on 33 of 105 pairs rather than guess.

Contour abstaining on Jordan and the Ironside raw selvedge jean at 58 percent confidence, an honest abstain banner citing a genuine fit conflict with waist and hip loose and thigh snug, recommending the least-bad size 31 and to bracket or speak to a stylist.
The abstain beat: a genuine fit conflict at 58 percent confidence, routed to bracketing or a stylist rather than a bluffed size.

An adversarial critic blocks physically contradictory copy

The Maverick Raw Selvedge Flex Jean ships vendor copy that is physically contradictory, 100 percent cotton raw selvedge denim with 4-way stretch. The FabricMechanicsCritic refuses the inference: raw selvedge and elastane are mechanically incompatible, a rigid woven cannot be 4-way stretch. No size recommendation is issued and the item routes to human review. This is honesty behavior, not a scored metric, and this garment is deliberately excluded from the scored golden set; it exists to show the block.

The critic blocking the Maverick Raw Selvedge Flex Jean, with a red Human review panel reading fabric inference rejected because raw selvedge and elastane are mechanically incompatible, a rigid woven cannot be 4-way stretch, so no size recommendation issued.
Contradictory copy, raw selvedge with 4-way stretch, is refused and routed to human review with no recommendation issued.

The benchmark, and the number to lead with

On the labeled 105-pair synthetic golden set, 15 bodies against 7 scored garments, Contour scores 100 percent against a waist-only baseline of 60 percent, a lift of 40 percentage points, with 39 conflicts caught and 67.6 percent of pairs given a single high-confidence size. State the scope every time: the golden-set labels use the same measurable fabric-stretch figures the engine uses, so the 100 percent is a property of this constructed set, not an open-world guarantee. The durable, independent number is the plus-40 point lift over the real incumbent, and it is largest on rigid fabrics: raw selvedge 100 versus 33, rigid chino 100 versus 40, wool suiting 100 versus 53, structured sateen 100 versus 60, stretch denim 100 versus 53, ponte knit 100 versus 87, ribbed knit 100 versus 93. On extreme-stretch knits the fabric forgives almost everything, so simple waist-matching nearly catches up.

The Contour benchmark view showing engine accuracy 100 percent, waist-only baseline 60 percent, a plus-40 percentage-point accuracy lift, 105 fits scored, 39 conflicts caught, and 67.6 percent single-size, with a per-shopper table for the raw selvedge jean marking the engine correct and the waist-only chart wrong on the abstain cases.
The +40 point lift on the labeled 105-pair golden set, with the engine's advantage concentrated in rigid denim and tailoring.

A replayable receipt, and the payload an AI agent consumes

Every recommendation writes a JSON audit receipt: the extracted fabric spec with the exact source phrases that drove each parameter, the critic verdict, the full per-size strain matrix, and the decision with its confidence. The same result is served at /api/fit as the machine-readable payload an AI shopping agent consumes, giving the recommended size, the confidence, an abstain flag, the per-zone fit array, and an audit id. As commerce moves to agents that transact on our behalf, the sizing signal they read has to be confidence-scored and auditable, a report rather than a picture.

The Contour audit receipt in JSON showing the fabric spec with fabric class raw selvedge denim, stretch 1 percent, elastic strain limit 0.02, extraction confidence 0.95, the source phrases 14oz raw selvedge denim, 100 percent cotton, rigid non-stretch, the critic verdict, and the per-size strain matrix.
The audit receipt: fabric provenance phrases, the critic verdict, and the full per-size strain matrix behind the decision.
The Contour agent API machine-readable response showing recommended_size 29, confidence 0.95, abstain false, the extracted fabric block, and a per_zone_fit array with strain and comfortable status for waist, hip, and thigh.
The /api/fit payload: recommended size, confidence, an abstain flag, and per-zone fit, all traceable to an audit id.

Where Contour sits, and where it does not

It is the fit-intelligence layer, the physics and the honesty policy, not a prettier picture or a smarter guesser.

Concern Size chart or virtual try-on Contour
The fabric's stretch limit Invisible; the same numbers for rigid and stretch Computed per zone against each fabric's elastic limit
Same body, two similar jeans One size for both, often wrong on the rigid one 29 in raw selvedge, 28 in stretch denim, each correct
When no size is clean Prints a confident size anyway Abstains and recommends bracketing or a stylist
Contradictory vendor copy Passed through unchecked Blocked by an adversarial critic, routed to human review
Who decides the size A lookup table, or an unaudited model A deterministic, unit-tested strain engine, not an LLM
What an AI shopping agent reads A picture, or nothing machine-readable A confidence-scored, abstain-aware /api/fit payload with an audit id

What this demo does not do

  • It does not promise perfect fit. The 100 percent is on a fixed, labeled 105-pair synthetic golden set whose labels share the engine's own measurable stretch figures, so it is not an open-world accuracy guarantee. The durable number is the plus-40 point lift over the waist-only incumbent.
  • The shoppers, bodies, garments, product copy, and measurements are synthetic. The cohort of 15 bodies and catalog of 8 garments, and names like Riley, Jordan, Ironside, Driftwood, and Maverick, are fictional.
  • It is a reduced-order analytical strain model, not a full FEA cloth simulation. We integrate with FEA tools like CLO3D, Browzwear, or Style3D at Tier-3; we do not rebuild them here.
  • It does not reconstruct a body from a photo. On-device photo-to-measurement capture is stubbed, and the demo accepts measurements directly, which is also the real privacy interface.
  • There are no live retail integrations. The Shopify, CLO-SET, and Browzwear connectors are a documented adapter interface behind a mock fixture, not live.
  • It is not a generative virtual try-on. The whole thesis is that pretty pixels are not fit data.
  • The abstain and the critic are shown as honesty behavior, not marketed as a scored percentage of errors caught.
  • There are no real customers, brands, deployments, testimonials, logos, or claimed ROI. The load-bearing decision is deterministic Python with 8 unit tests, and the LLM advisors are provider-swappable and skippable, defaulting to claude-opus-4-8.

Questions buyers ask

We already publish a size chart and added a virtual try-on. Why do we still get fit returns?

Because neither one can feel the fabric. A size chart gives four 1-D numbers to describe a 3-D body, and it has no idea whether the denim is zero-stretch raw selvedge or 4-way ponte. A generative try-on image makes the garment look right on a body but tells you nothing about whether the thigh exceeds that fabric's stretch limit when the shopper sits down. Fit is a mechanical fact, the circumferential strain at each zone against the fabric's elastic limit, and that is exactly what a chart and a pretty render both leave out. Contour computes it directly, which is why the same body can be a correct 29 in raw selvedge and a correct 28 in stretch denim.

Is this a virtual try-on that shows the clothes on a body?

No. That is the opposite of the thesis. A virtual try-on can show you the jeans on your body and still have no idea if they will actually fit, because pretty pixels are not fit data. Contour takes measurements in and returns a size, a confidence, and per-zone fit notes computed from the fabric's mechanics. It is a report, not a picture, which is also what makes it something an AI shopping agent can consume at checkout.

You say 100 percent on your benchmark. Is it actually always right?

No, and we do not claim that. The 100 percent is on a fixed, labeled 105-pair synthetic golden set whose labels use the same measurable fabric-stretch figures the engine uses, so that set is not a fully independent oracle and the 100 percent is a property of this constructed set, not an open-world accuracy guarantee. The number to lead with is the lift over the real incumbent method, waist-only size-chart matching, on the same labels: Contour beats it by 40 percentage points, 100 percent versus 60 percent, and the advantage is concentrated in rigid denim and tailoring where fit is hardest. Treat 100 percent as an internal ceiling on a labeled set, never as a promise of perfect fit in the wild.

Does it need a photo or a body scan of the shopper?

No. Contour is privacy-first: measurements in, no photo. The demo accepts centimetre measurements directly and checks them against the garment's fabric mechanics, which is also the real privacy-preserving interface we would ship. On-device photo-to-measurement capture is stubbed in this demo, not built, so nothing here depends on monocular body reconstruction.

Does it connect to Shopify or our PLM, CLO3D, or Browzwear today?

Not in this demo. The Shopify, CLO-SET, and Browzwear connectors are a documented adapter interface behind a mock fixture, not live integrations. The shoppers, garments, product copy, and measurements are synthetic test fixtures, not real catalog. And to be precise about scope, this is a reduced-order analytical strain model, the page's Tier-1 and Tier-2 physics, not a full FEA cloth simulation; we integrate with FEA tools like CLO3D or Browzwear at Tier-3, we do not rebuild them.

An AI is picking the size my checkout depends on. Why should I trust it?

Because the load-bearing decision is not made by an LLM. The size and the abstain are computed in plain, unit-tested Python: for each candidate size the engine computes the circumferential strain at each zone against the fabric's elastic limit, and 8 unit tests pin that physics. The language model only reads the messy vendor copy into a structured fabric spec and phrases the final note; it never computes the numbers or the decision. Agents advise, code decides, which is the property that makes it safe to expose as an API your checkout can call.

What does it do when it is not sure?

It abstains instead of bluffing. When no single size clears every zone within comfort, on a zero-stretch fabric where small sizes fail the thigh and larger sizes run loose at the waist, the policy gate recommends the least-bad size and flags bracket or speak to a stylist rather than confidently printing a wrong size the way the chart does. Across the labeled 105-pair golden set the engine abstains on 33 of 105 pairs. A separate honesty behavior, the adversarial critic, blocks physically contradictory vendor copy such as raw selvedge with 4-way stretch and routes it to human review with no recommendation issued.

Technical Research

The research behind this demo — the architecture, the verification design, and the enterprise blueprint.

Turn your returns problem into a fit problem you can compute

Physics that reads the fabric, and the discipline to abstain when no size is clean.

If your e-commerce, returns, or engineering team is working out how to give shoppers, and the AI agents that will soon shop for them, a sizing signal they can trust, we would genuinely like to hear how you are thinking about it. The problem is industry-wide and the answers will be too.

Fit and returns assessment

  • ✓ Map where fit returns and bracketing concentrate in your catalog
  • ✓ Model the fabric mechanics and per-zone comfort bands for your fits
  • ✓ Design the privacy-first measurement interface, no photo required
  • ✓ Define the abstain policy your merchandising and CX teams can stand behind

Build with us

  • ✓ A deterministic, unit-tested strain engine outside any model
  • ✓ A typed agent crew with an adversarial critic that can block bad copy
  • ✓ A confidence-scored /api/fit endpoint with replayable audit receipts
  • ✓ Adapter interfaces for Shopify, CLO-SET, Browzwear, and your PLM