← All work

Search & discovery

Canon from fragmented sources

Dozens of incompatible feeds, one clean geocoded canon.

The system

A parse–normalize–geocode–dedupe pipeline that turns scattered, format-hostile listings into a single canonical dataset — crawled once, served many.

An agentic find-and-review loop proposes new sources and structured diffs; a human approves. The grind nobody would fund is now an agent away, and the resulting canon is the moat.

The hard parts

Format hostility

Seven parser families cover feeds, embedded structured data, and hand-rolled markup, converging on one schema.

Agentic sourcing

Discovery runs as an agent loop with human review — the pipeline widens itself.

In numbers

~3,000
canonical records
7
parser families
26
app routes

Next.js · Firebase · agent loops · geocoding

Happy to walk through this one properly — what it does, how it's put together, and what it would take to build something like it for you.

Vague is fine.

What kind of help? optional — pick any