Skip to content
← All projects

Data · Ingredient intelligence

Skincare & Ingredient Intelligence

A pipeline that collects skincare products, cleans their ingredient lists and stores consistent product records and images.

~85%Less manual time per 1,000 products
1,000Products per batch

Outcome

An extraction and matching pipeline cut manual handling for a 1,000-product batch from 30–50 hours to 3–5 hours of exception review — about an 85% reduction, with uncertain matches and missing images flagged rather than silently accepted.

Problem

Product information sits across thousands of inconsistent pages. Ingredient names vary by spelling, format and naming convention. Some pages bundle several products together. Images can be missing, wrong or linked to the wrong database record.

That makes direct comparison unreliable without a large amount of manual cleaning.

Solution

I built an extraction, cleaning and matching pipeline backed by Airtable. It converts public product pages into structured records, maps raw ingredient text to a standard ingredient list and prepares consistent product images.

Steps

  1. Discover and scrape candidate pages with Firecrawl. Write progress and failure states back to Airtable so batches can resume.
  2. Classify each URL as a product or non-product page, then extract brand, category and ingredient fields into a fixed schema.
  3. Normalise raw ingredient text. Rules detect bundled products, product codes, percentage prefixes and several ingredients joined into one value.
  4. Create an embedding for each cleaned ingredient name. An embedding is a numeric representation of the text. FAISS searches those numbers to find the closest entries in the standard taxonomy.
  5. Send uncertain top matches through a separate verification step. Failed values are cleaned again and re-matched instead of being silently accepted.
  6. Use fuzzy string comparison to identify likely duplicate ingredients that survived the earlier checks.
  7. Download candidate product images, verify their contents, remove backgrounds and upload approved files to Cloudflare R2.
  8. Run repair tools that compare R2 filenames with Airtable records and report missing or mismatched files.
  9. Benchmark several language models on the same extraction set. Measure accuracy, success rate and cost per product before choosing a model.
  10. Use asynchronous workers with explicit rate limits. Produce timestamped JSON and CSV reports at every major stage for checking and rollback.

Results and benefits

A batch of 1,000 products would require 30 to 50 hours of manual extraction and ingredient matching. The pipeline completes the first pass in hours and leaves about three to five hours of exception review.

That represents about an 85% reduction in manual time. Each batch also produces a list of uncertain matches, missing images and failed records for targeted correction.

Skills & tools

What it took to build this.

Python
Airtable
Gemini
Cloudflare R2
Data engineering3
  • ETL design
  • Ingredient normalisation
  • Fuzzy duplicate detection
AI & ML2
  • Semantic search (embeddings + FAISS)
  • Model benchmarking
Media & reliability3
  • Image verification & processing
  • Async concurrency
  • Checkpointed batches
Focus
Product data, ingredient matching and image processing
Status
Operating in production

Related work

More projects.

Get in touch

Working on something difficult in health, care or AI?

I occasionally compare notes with founders, product teams and operators working through high-friction workflows in healthcare, care and applied AI.

Start a conversation