Alex Dent

Collected Works

The Back Catalog

A luxury resale price intelligence platform, run by an autonomous data engineering system.

Year
2026
Role
Built it solo

Luxury Resale Price Intelligence

The same Hermès Birkin, Rolex Submariner, or Chanel flap bag can trade 50% apart on the same day depending on which venue it’s listed on. Unlike equities, cars, or even sneakers, there is no consolidated price history for luxury goods that anyone can look up. Sellers guess. Buyers overpay. The dealers who price well do it from memory and relationships, and that knowledge doesn’t leave their heads.

I’m building the price layer that market is missing: a system that continuously observes what luxury goods are listed and sold for across the resale market, resolves those observations to individual objects, and turns them into a queryable price history.

The hard part is identity, not collection.

Scraping 50 sources is labor. The actual problem is that the same object is described fifty different ways — inconsistent model names, reference numbers omitted, condition grades that don’t map across platforms, hardware and size encoded in emoji. Deterministic matching rules get you to maybe 70% coverage and then fail in ways that are worse than missing data: they silently merge two different references and produce a confident, wrong price.

So normalization runs on embeddings. Listing text, attributes, and images are collapsed into a vector representation and matched against a canonical catalog, with a similarity threshold and a human review queue governing the cases the model isn’t sure about. That’s what makes the price history comparable rather than a pile of loosely related listings.

What the agents actually do.

The pipeline is maintained by a set of agents with narrow jobs, because a 50-source scraper fleet breaks constantly and by hand it’s a full-time job:

  • Extraction agents per source, adapted to each site’s structure rather than hand-written one at a time.
  • Normalization, the embedding layer above, plus the classifier that decides what’s a new object versus a re-listing of a known one.
  • Monitoring on output shape rather than uptime — the failure mode that matters isn’t a 500, it’s a scraper that keeps returning 200s and thousands of rows of garbage. Detection watches for volume drops, field-null spikes, price distributions that shift more than a source plausibly could overnight.
  • Resolution: when monitoring fires, an agent reproduces the failure, patches the extractor, and opens a pull request against the pipeline repo with the diff and the failing case attached.

It doesn’t deploy on its own — it creates PRs waiting on CI tests and human review before deploy. The honest description is that the system converts a class of maintenance work from “notice it a week later and spend an afternoon” into “a reviewable diff waiting for me in the morning and evening.” 

 

What the Product does.

It reconciles the catalogs of 50+ luxury resale platforms into a single comparable view of the market — 1.5M price observations a day, resolved to individual objects rather than listings.

 

Where it is now.

The pipeline is live and generating early affiliate revenue on the consumer side. Consumer-first is a deliberate sequence rather than a fallback: the crawl gives me supply, but only an audience gives me demand — what people actually search for, compare, and buy at what price. Supply data alone produces a price index. Supply plus revealed demand produces something a dealer can make inventory decisions with, and that second dataset can’t be scraped by anyone who doesn’t own the audience first.

It also means the B2B product gets built on observed behavior instead of assumptions, and that a slice of the consumer audience — resellers and small dealers sourcing on the same platforms — becomes a warm channel into it rather than a cold enterprise sales motion.

The asset compounds either way. A price series is worth more at three years than at three months, and it isn’t something a competitor can backfill after the fact.