What Is Virtual Try-On? The Three Technologies Behind It
Virtual try-on is any technology that shows a product on your own face or body before you own it — frames on your nose, a lip shade on your mouth, a dress on your frame — instead of asking you to project yourself onto a photo of a stranger. One label covers three unrelated technologies, and they are nowhere near equally good. AR overlay is genuinely solved for eyewear, and solved for makeup placement if not for makeup shade. Generative 2D synthesis is convincing for clothing silhouette and hopeless for clothing size. 3D body modeling is the only route that produces real numbers, and almost nobody wants to do the scan. Knowing which one you are looking at tells you exactly how much to trust the picture.
What virtual try-on actually means
A virtual try-on system takes two inputs — a representation of you (a live camera feed, one selfie, or a scanned body) and a representation of the product (a 3D asset, a flat product photo, or a digital garment pattern) — and returns a composite in which you are wearing the thing. That is the whole definition. Everything else is implementation.
It is worth separating try-on from the feature retailers usually ship next to it: the size recommender. A size recommender outputs a letter or a number ("most people with your measurements take a M in this brand"), derived from your stated height and weight, your past order and return history, or a body scan. It never draws a picture. Try-on draws a picture and, in most implementations, never outputs a size. The two answer different questions, and conflating them is the single most common misreading of the category.
The three technologies, side by side
| Approach | How it works | Speed | Strongest categories | Where it breaks |
|---|---|---|---|---|
| AR overlay | Face or hand tracking on your live camera feed; a 3D model or a shader is pinned to detected landmarks | Real time, on-device | Eyewear, lipstick, eyeshadow, earrings, watches, nail polish | Anything that must drape or deform — garments, bags worn on the body |
| 2D image synthesis | A generative model takes one photo of you plus one photo of the product and paints a new image of you wearing it | Seconds per image | Tops, dresses, outerwear, costumes, tattoos, hair | Fine print, logos, buttons and hardware; produces no measurement at all |
| 3D body modeling | A scanned or estimated 3D body, plus a digital garment run through cloth simulation | Minutes to build the avatar, then interactive | Made-to-measure, uniform programs, technical fit checks | Consumer friction — it needs a scan, and the avatar rarely reads as you |
AR overlay: the mature one
Live AR try-on is the oldest and by far the most reliable branch, because the hard computer-vision problem underneath it — dense face tracking — has been commodity for years. A phone locks onto a face mesh (Apple ARKit and Google ARCore both expose one; the 468-point mesh MediaPipe popularized is the open equivalent), and the product is drawn onto that mesh every frame. L’Oréal’s ModiFace and the frame try-on running on eyewear sites are all versions of this. It runs at video frame rate on mid-range phones, not just current flagships — ARCore’s supported-device list still reaches back to handsets from the late 2010s.
It works because these products are geometrically easy. Glasses are rigid — the only question is whether the frame width suits your face width, and that is a genuine measurement the tracker can make. Lipstick and eyeshadow are thin films painted onto stable regions of the face; there is no physics to simulate. The residual error is not shape, it is color: what a shade looks like depends on your uncalibrated screen, the color temperature of the room you are standing in, and the auto white-balance your camera silently applies. Every honest beauty AR feature is a shape-and-placement preview with an approximate shade, which is why swatches and return policies still exist.
2D image synthesis: the fast-moving one
The generative branch skips 3D entirely. You give the model a picture of a person and a picture of a garment, and it produces a new photograph of that person in that garment. Google’s shopping try-on, built on its TryOnDiffusion research, is the reference implementation: two diffusion U-Nets, one reading the garment and one reading the person, exchanging information through cross-attention so the fabric lands with plausible folds, cling and shadow rather than being pasted on like a sticker. Google shipped it against a roster of roughly 80 real models spanning XXS to 4XL, cast for body shape, hair type and skin tone using the Monk Skin Tone Scale — a deliberate correction to how thin the model rosters in this category have historically been.
This is the approach that made try-on usable for clothing, and it is the approach ReelWand’s Selfie Try-On runs. It is very good at drape, color, proportion, and the general question of whether a silhouette suits you. It is unreliable at anything with fine structure: printed text warps, logos smear, button plackets migrate, a specific hemline lands where the model thinks it should rather than where it would on you. And the fundamental limit is not resolution, it is category — the output is an image. An image contains no measurement, so it cannot tell you a size, no matter how photoreal it looks.
3D body modeling: the accurate one nobody uses
The third route builds an actual 3D body — from a booth scan, a depth sensor, or (increasingly) two smartphone photos and a pose-estimation model — and then simulates a digital garment on it with cloth physics: stretch, weight, stiffness, seam allowance. This is the only branch that produces numbers a pattern-maker can use. Scanning vendors publish accuracy claims in the 96–97% range against tape-measure ground truth from two photos; treat those as vendor-reported rather than independently audited, but the direction is right — phone-based measurement is now good enough for made-to-measure shirting and uniform programs.
Its problem is human, not technical. Asking a shopper to stand against a blank wall in fitted clothing and photograph themselves twice is a conversion cliff, and the avatar that comes back tends to sit in an uncanny valley — recognizably shaped like you, not recognizably you. So the technology mostly survives invisibly: the scan feeds a size recommender, and the shopper never sees the mesh.
Which categories actually work today
- Eyewear — trustworthy. Rigid geometry, a face-width measurement that is really being taken, and no fabric to simulate. The closest thing to a solved problem in the category.
- Makeup — trustworthy for shape, approximate for shade. Placement, coverage and finish read accurately; the exact color depends on your screen and your lighting. A makeup look preview tells you whether a cut-crease suits your eye shape, not whether that specific SKU is your undertone.
- Jewelry, watches, nails — reliable. Small, rigid, and tracked well; the main variable is scale, which a wrist or hand reference solves.
- Tattoos and hair — reliable, because there is no size chart. Placement and scale are the whole question, and generative synthesis handles both. Tattoo Try-On and Hairstyle Mirror live here.
- Clothing — useful for look, unreliable for fit. Silhouette, color, proportion and styling: yes. Whether this brand’s M closes across your chest: no.
- Shoes — the hardest. Foot length and width are the question, and estimating them from a camera is much harder than estimating a face. AR shoe try-on is a styling toy; sizing still comes from the size chart.
Try-on shows you how something looks, not whether it fits. A rendered garment is a picture, not a measurement — the model was never told your bust, waist or inseam, so it draws a plausible fit rather than your fit. Use try-on to decide whether you want the item; use the brand’s size chart and your own measurements to decide which size to order.
What retailers really get out of it
The business case is almost always framed as returns, and the framing is only half right. NRF puts returns at roughly a fifth of online sales across all categories; apparel sits well above that line, with category benchmarks usually quoted in the 24–30% band and individual brands higher still. Fit and sizing lead the reason list in every survey worth reading, but the share swings hard with the wording: Narvar’s 2025 consumer work attributes 42% of shoppers’ most recent returns to size and fit, while apparel-only studies push it past half. The rest is "didn’t look how I expected", "wrong color", "changed my mind", plus the increasingly normalized habit of bracketing: deliberately ordering two or three sizes and returning what loses.
Map that against the technology and the split is clean. Try-on directly attacks the "didn’t look how I expected" and "wrong color" buckets, which are real and expensive but are not the biggest bucket. It barely touches the fit bucket, and it may actively worsen bracketing if a shopper sees a great-looking render and hedges on size anyway. The vendor case studies claiming 20–40% return reductions are self-reported and rarely control for the obvious selection effect: shoppers who engage a try-on widget were already further down the funnel than shoppers who did not. Treat those numbers as directional marketing, not measurement. The reliably observed effects are softer and still worth money — longer dwell time on the product page, higher add-to-cart rates, and fewer pre-purchase support questions about how something sits.
The limits worth knowing before you trust the picture
- Size is not fit. Two size-M tops from two brands can differ by several centimeters across the chest. Vanity sizing means the letter carries almost no cross-brand information, and no rendering engine can recover what it was never given.
- Fabric properties are invisible in a product photo. Stretch, weight, drape stiffness and lining all determine how a garment sits, and none of them are recoverable from a flat catalogue shot. A synthesis model infers them from what similar-looking garments did in its training data, which is a guess dressed as a photograph.
- Representation is a real, measurable gap. Try-on quality has historically been best for the body types, skin tones and hair textures that dominated the training data. Google casting to the Monk Skin Tone Scale and publishing an XXS–4XL roster was a response to that gap, not a flourish — and it is still the exception rather than the norm across the industry.
- Fine detail degrades first. Slogan tees, brand logos, contrast stitching, hardware and small prints are where 2D synthesis visibly breaks. If the detail is the reason you want the item, verify it on the product photo instead.
- Faces and bodies are biometric data. Try-on features have already drawn litigation in the US under state biometric privacy laws, Illinois’s BIPA in particular. Before you upload, it is fair to ask what happens to the image and how long it is kept.
Try it on your own selfie
ReelWand’s Selfie Try-On is the 2D-synthesis kind, and we would rather be blunt about what that means. Upload one selfie — no full-body photo, no scan, no changing room — describe the garment, and it renders a photoreal you wearing it with your face, skin tone and proportions held, re-tailoring sleeves, hemline and drape so the fabric falls the way it would rather than sitting on you like a decal. That is a genuinely useful answer to "does this look like me". It is not an answer to "which size do I order", and we will not pretend otherwise: check the size chart.
The same synthesis powers the rest of the try-on family — Makeup Look Studio for a full look on your own skin, Tattoo Try-On for placement and scale before you book the chair, Nail Studio, Hairstyle Mirror for a cut or color you are hesitating over. If you are on the selling side of this rather than the buying side, the catalogue equivalent is On-Model Studio, which puts your garment on a model instead of putting a garment on your customer — see on-model photography for how that fits a product workflow.
You can test one render without an account: visitors get one free image per day per IP, watermarked. A free account adds 8 credits once — two renders at the default Seedream model’s 4 credits an image, and they never top back up. Paid plans start at 400 credits a month and drop the watermark.
One selfie in, a photoreal you in the outfit out — your face and proportions held, the fabric re-tailored to fall properly.
Try a garment on your selfieFrequently asked questions
What is virtual try-on?
Virtual try-on is any technology that renders a product on your own face or body before you buy it. It takes a representation of you — a live camera feed, a single selfie, or a 3D body scan — plus a representation of the product, and returns a composite in which you are wearing it — a still frame painted by a diffusion model, or a live camera view repainted every frame. Three distinct technologies share the name: real-time AR overlay, generative 2D image synthesis, and 3D body modeling with cloth simulation. They differ enormously in what they can be trusted to tell you, so the first useful question about any try-on feature is which of the three it is.
Is virtual try-on accurate for clothing sizes?
No, and it is important to be clear about that. Image-based try-on produces a picture, and a picture contains no measurement — the model was never told your bust, waist, hip or inseam, so it renders a plausible fit rather than your fit. It answers whether a silhouette, length and color suit you, which is genuinely useful. For sizing you still need the brand’s size chart against your own measurements, a size recommender trained on that brand’s returns, or a body scan. The only branch of try-on that yields real numbers is 3D body modeling, and most consumer features do not use it.
What is the difference between AR try-on and AI try-on?
AR try-on tracks your face or hands in a live camera feed and pins a 3D model or a color shader to the tracked landmarks, updating every frame. It is real-time, runs on-device, and works well for rigid or thin-film products: glasses, lipstick, earrings, watches, nail polish. AI try-on in the current sense usually means generative 2D synthesis — a diffusion model takes one photo of you and one of the garment and paints a new photo of you wearing it, taking seconds rather than milliseconds. AR wins on speed and interactivity; generative synthesis wins on anything that has to drape.
Does virtual try-on actually reduce returns?
Partly, and less than the marketing suggests. Fit and sizing lead apparel return reasons — Narvar’s 2025 survey pins 42% of shoppers’ most recent returns on size and fit, and apparel-only studies run higher — and image-based try-on does not address fit. What it does address is the smaller "didn’t look how I expected" and "wrong color" share, plus pre-purchase hesitation. Vendor case studies quoting 20–40% return reductions are self-reported and rarely control for the fact that shoppers who use a try-on widget were already higher-intent. The consistently observed wins are softer: more time on the product page, higher add-to-cart, fewer questions to support.
Do I need a full-body photo for virtual try-on?
It depends on the technology. A 3D body-modeling system does need a deliberate capture — typically a front and side photo in fitted clothing against a plain wall, which is exactly the friction that keeps adoption low. Generative 2D try-on does not: ReelWand’s Selfie Try-On works from a single ordinary selfie and reconstructs the rest of the figure, which is why it can run in the middle of a browsing session. The trade-off is the one running through this whole category — less capture effort means less information, which means a look preview rather than a fit verdict.
Put it into practice
Specialized image agents carry the craft this guide describes. Pick an available agent and start creating.
Part of ReelWand's AI Product Photography & Photo Editing tools.