Delphine
Pre-season buying decisions for footwear. A buyer sends one photo of a shoe that hasn't launched, and Delphine predicts weekly sell-through velocity per store cluster, sizes the order, and returns a buy/skip tier with a calibrated uncertainty band.
The problem
Footwear buyers commit to inventory 6–12 months before a shoe reaches a shelf, working from gut feel and last year's numbers. The expensive part isn't picking a bad shoe — it's ordering the right shoe in the wrong depth, for the wrong stores. Both errors cost margin in opposite directions: too shallow and the season sells out early at full price you never captured; too deep and the surplus clears at markdown.
How it works
Delphine predicts units per active store per week at (SKU, banner, store-cluster, quarter) grain, then converts that into a newsvendor-sized order quantity per tier. Every prediction carries a conformal-calibrated P20/P50/P80 band, and an explanation of why that band is wide.
Visual peer matching
Cross-banner cascade modeling
Store-cluster granularity
Calibrated uncertainty
What it's worth
A third of our units sell at a markdown, at an average depth of ten percent off normal price — and that depth has widened every year since 2023. Buy depth is what moves it. At our own scale, taking five to ten percent off the markdown bill is worth seven figures a year, and we can measure it because we own both the model and the P&L it runs against.
What this costs on the open market
The nearest commercial equivalent — a venture-backed apparel forecasting platform built on the same bet, product imagery plus external signals — raised US$25M and was acquired by an enterprise planning vendor in 2025. It is no longer sold on its own: the capability now ships inside a planning platform whose demand module alone lists at US$80,000–$250,000 a year, with implementation budgeted at 60–120% of the first-year subscription on top. That vendor's published benchmark for the category is savings of roughly 2% of revenue. One difference worth stating: they market social-sentiment and influencer signals as inputs. We tested those and removed them — against our own data they measured net negative.