Contour · AI Fit Prediction for Fashion E-Commerce
Your returns problem is a fit problem, and fit is a physics problem. For a given body (measurements, no photo) and a garment (fabric mechanics plus finished dimensions), Contour computes the circumferential strain at each zone against that fabric's stretch limit, then recommends a size with confidence, or abstains honestly when no size is clean. Agents advise, code decides.
+40 pp
Accuracy lift over the waist-only size-chart incumbent, 100 percent versus 60 percent
Labeled 105-pair synthetic golden set
67.6%
Of pairs given a single high-confidence size, so the shopper need not bracket two
Same 105-pair golden set, confidence at least 0.80
39 of 105
Fit conflicts the waist chart would have shipped, caught before purchase
Same 105-pair golden set
A runnable demo of the mechanism on a synthetic cohort of 15 shoppers and a catalog of 8 garments. The deterministic strain engine and the benchmark run fully offline with no API key; the optional agents are the only part that can use a model.
Apparel returns are dominated by fit, and the root cause is mechanical, not cosmetic.
Between 53 and 67 percent of apparel returns are fit-related, and about 63 percent of shoppers bracket, ordering two sizes to send one back (Veriprajna WP34 research, 2026, drawing on industry sources). The carrier eats the reverse logistics, and jeans run among the highest apparel return rates precisely because of stretch, rise, and waist mechanics.
The root cause is simple to state. A size chart gives four 1-D numbers to describe a 3-D body, and it has no idea whether the fabric is zero-stretch raw selvedge or 4-way ponte. So the same waist measurement lands the same shopper in the wrong size on a rigid denim and the right size on a stretch one, and the chart cannot tell those two garments apart. Bracketing is the rational shopper response to a tool that cannot see fabric.
The industry's reflex has been to solve returns with better pictures. A generative virtual try-on makes the garment look right on a body, but a picture cannot tell you whether the thigh exceeds that denim's stretch limit when the shopper sits down. Whether a size fits is a mechanical fact: the circumferential strain at each body zone against that fabric's elastic limit. That is the quantity a chart drops and a render never had, and it is the one Contour computes.
One pipeline, shown a stage at a time. The load-bearing decision lives in plain, unit-tested Python, never in a model.
For a body and a garment, the pipeline runs in order: a fabric-mechanics agent extracts a structured spec from the vendor copy, an adversarial critic verifies it and can block, the strain engine decides the size in pure code, a policy gate picks the clean size or abstains, and a fit-note agent phrases the result. Every number that decides a recommendation is computed before any language model touches the wording.
The part outside any language model. For each candidate size the strain engine computes zone strain as the body circumference minus the garment's finished circumference over that finished circumference, then compares each zone to the fabric's elastic limit and a per-zone comfort band. The policy gate picks the size with the lowest total per-zone regret, and abstains when no size clears every zone within comfort. This is the load-bearing decision, it lives entirely in plain Python, and 8 unit tests pin the physics, so trust does not depend on the model.
On top of the core, a typed crew reads and communicates. The FabricMechanicsAgent turns messy product copy, the title, fiber content, weight, and construction, into a structured fabric spec with a stretch class, an elastic strain limit, and the source phrases that justify each parameter. The FabricMechanicsCritic checks that spec against physics constraints and blocks contradictions. The FitNoteAgent phrases the computed result. The advisors are provider-swappable and default to claude-opus-4-8; with no key the fabric extraction falls back to a cached spec and the deterministic engine still runs live, so the demo records fully offline. Agents advise, code decides.
| Stage | What it does | Basis |
|---|---|---|
| FabricMechanicsAgent | Reads vendor product copy into a structured fabric spec with stretch class, elastic strain limit, weight, and source phrases. | LLM, typed |
| FabricMechanicsCritic | Checks the spec against physics constraints and blocks contradictions, routing to human review with no recommendation. | LLM plus deterministic rules |
| Strain engine | Computes per-zone circumferential strain against the fabric's elastic limit for every candidate size. | pure code, decides |
| Policy gate | Picks the size with the lowest total per-zone regret, or abstains when no size clears every zone. | pure code, abstains |
| FitNoteAgent | Phrases the computed result, for example snug at hip, relaxed at thigh. It never computes the numbers. | LLM, cosmetic |
| Audit receipt and API | Writes a replayable JSON receipt and serves the same result at /api/fit for an AI shopping agent. | deterministic output |
The durable value is the physics plus the audit receipt, not a model that guesses sizing better. Even a perfect language model still needs the fabric mechanics, the body geometry, the per-zone strain calculation, the abstain policy, and a replayable receipt. That infrastructure is the product; the LLM is one swappable advisor inside it, which is why the approach does not age out as models improve.
Every shopper and garment is synthetic: a cohort of 15 bodies and a catalog of 8 garments, with fictional names. Each on-screen number is produced by the engine at runtime, not hard-coded.
Riley on the Ironside 14oz Raw Selvedge Straight Jean returns a recommended Size 29 at 95 percent confidence, comfortable at every zone. A waist-only chart would say 28, and it would be wrong: at 28 the thigh exceeds this zero-stretch denim's stretch limit. The per-zone strain bars make it legible, green where each zone sits inside the fabric's comfort band. The chart cannot feel the fabric, so it drops the one variable that decides the size.
Switch the identical shopper to the Driftwood Stretch Slim Jean, a visually similar jean, and the correct size drops to 28 at 90 percent confidence, clean at every zone. Here the waist-only chart happens to also land on 28, because this fabric's stretch forgives the thigh that the raw selvedge punished. That is the whole thesis in one beat: 29 in one jean and 28 in a look-alike, on the same body, because the engine computes how each fabric's stretch limit meets that body.
The shopper Jordan on the Ironside raw selvedge jean produces a genuine fit conflict: small sizes fail the thigh, larger sizes run loose at the waist, and on this zero-stretch fabric no size clears every zone. Contour returns an honest abstain at 58 percent confidence, recommends the least-bad size and tells the shopper to bracket or speak to a stylist. The waist-only chart, by contrast, confidently prints 28. Across the labeled golden set the engine abstains on 33 of 105 pairs rather than guess.
The Maverick Raw Selvedge Flex Jean ships vendor copy that is physically contradictory, 100 percent cotton raw selvedge denim with 4-way stretch. The FabricMechanicsCritic refuses the inference: raw selvedge and elastane are mechanically incompatible, a rigid woven cannot be 4-way stretch. No size recommendation is issued and the item routes to human review. This is honesty behavior, not a scored metric, and this garment is deliberately excluded from the scored golden set; it exists to show the block.
On the labeled 105-pair synthetic golden set, 15 bodies against 7 scored garments, Contour scores 100 percent against a waist-only baseline of 60 percent, a lift of 40 percentage points, with 39 conflicts caught and 67.6 percent of pairs given a single high-confidence size. State the scope every time: the golden-set labels use the same measurable fabric-stretch figures the engine uses, so the 100 percent is a property of this constructed set, not an open-world guarantee. The durable, independent number is the plus-40 point lift over the real incumbent, and it is largest on rigid fabrics: raw selvedge 100 versus 33, rigid chino 100 versus 40, wool suiting 100 versus 53, structured sateen 100 versus 60, stretch denim 100 versus 53, ponte knit 100 versus 87, ribbed knit 100 versus 93. On extreme-stretch knits the fabric forgives almost everything, so simple waist-matching nearly catches up.
Every recommendation writes a JSON audit receipt: the extracted fabric spec with the exact source phrases that drove each parameter, the critic verdict, the full per-size strain matrix, and the decision with its confidence. The same result is served at /api/fit as the machine-readable payload an AI shopping agent consumes, giving the recommended size, the confidence, an abstain flag, the per-zone fit array, and an audit id. As commerce moves to agents that transact on our behalf, the sizing signal they read has to be confidence-scored and auditable, a report rather than a picture.
It is the fit-intelligence layer, the physics and the honesty policy, not a prettier picture or a smarter guesser.
| Concern | Size chart or virtual try-on | Contour |
|---|---|---|
| The fabric's stretch limit | Invisible; the same numbers for rigid and stretch | Computed per zone against each fabric's elastic limit |
| Same body, two similar jeans | One size for both, often wrong on the rigid one | 29 in raw selvedge, 28 in stretch denim, each correct |
| When no size is clean | Prints a confident size anyway | Abstains and recommends bracketing or a stylist |
| Contradictory vendor copy | Passed through unchecked | Blocked by an adversarial critic, routed to human review |
| Who decides the size | A lookup table, or an unaudited model | A deterministic, unit-tested strain engine, not an LLM |
| What an AI shopping agent reads | A picture, or nothing machine-readable | A confidence-scored, abstain-aware /api/fit payload with an audit id |
Because neither one can feel the fabric. A size chart gives four 1-D numbers to describe a 3-D body, and it has no idea whether the denim is zero-stretch raw selvedge or 4-way ponte. A generative try-on image makes the garment look right on a body but tells you nothing about whether the thigh exceeds that fabric's stretch limit when the shopper sits down. Fit is a mechanical fact, the circumferential strain at each zone against the fabric's elastic limit, and that is exactly what a chart and a pretty render both leave out. Contour computes it directly, which is why the same body can be a correct 29 in raw selvedge and a correct 28 in stretch denim.
No. That is the opposite of the thesis. A virtual try-on can show you the jeans on your body and still have no idea if they will actually fit, because pretty pixels are not fit data. Contour takes measurements in and returns a size, a confidence, and per-zone fit notes computed from the fabric's mechanics. It is a report, not a picture, which is also what makes it something an AI shopping agent can consume at checkout.
No, and we do not claim that. The 100 percent is on a fixed, labeled 105-pair synthetic golden set whose labels use the same measurable fabric-stretch figures the engine uses, so that set is not a fully independent oracle and the 100 percent is a property of this constructed set, not an open-world accuracy guarantee. The number to lead with is the lift over the real incumbent method, waist-only size-chart matching, on the same labels: Contour beats it by 40 percentage points, 100 percent versus 60 percent, and the advantage is concentrated in rigid denim and tailoring where fit is hardest. Treat 100 percent as an internal ceiling on a labeled set, never as a promise of perfect fit in the wild.
No. Contour is privacy-first: measurements in, no photo. The demo accepts centimetre measurements directly and checks them against the garment's fabric mechanics, which is also the real privacy-preserving interface we would ship. On-device photo-to-measurement capture is stubbed in this demo, not built, so nothing here depends on monocular body reconstruction.
Not in this demo. The Shopify, CLO-SET, and Browzwear connectors are a documented adapter interface behind a mock fixture, not live integrations. The shoppers, garments, product copy, and measurements are synthetic test fixtures, not real catalog. And to be precise about scope, this is a reduced-order analytical strain model, the page's Tier-1 and Tier-2 physics, not a full FEA cloth simulation; we integrate with FEA tools like CLO3D or Browzwear at Tier-3, we do not rebuild them.
Because the load-bearing decision is not made by an LLM. The size and the abstain are computed in plain, unit-tested Python: for each candidate size the engine computes the circumferential strain at each zone against the fabric's elastic limit, and 8 unit tests pin that physics. The language model only reads the messy vendor copy into a structured fabric spec and phrases the final note; it never computes the numbers or the decision. Agents advise, code decides, which is the property that makes it safe to expose as an API your checkout can call.
It abstains instead of bluffing. When no single size clears every zone within comfort, on a zero-stretch fabric where small sizes fail the thigh and larger sizes run loose at the waist, the policy gate recommends the least-bad size and flags bracket or speak to a stylist rather than confidently printing a wrong size the way the chart does. Across the labeled 105-pair golden set the engine abstains on 33 of 105 pairs. A separate honesty behavior, the adversarial critic, blocks physically contradictory vendor copy such as raw selvedge with 4-way stretch and routes it to human review with no recommendation issued.
The research behind this demo — the architecture, the verification design, and the enterprise blueprint.
Physics that reads the fabric, and the discipline to abstain when no size is clean.
If your e-commerce, returns, or engineering team is working out how to give shoppers, and the AI agents that will soon shop for them, a sizing signal they can trust, we would genuinely like to hear how you are thinking about it. The problem is industry-wide and the answers will be too.