Finder
Turning the web archive into business memory
An archive-scale data engine that finds useful signals in the historical web, reconnects fragmented identities and produces professional profiles with evidence behind them.

The opportunity
Intelligence hidden in time
Billions of archived pages preserve traces of companies, teams, roles and relationships that have disappeared from the live web. That history can reveal how a market evolved or help recover a lost connection, but its scale and inconsistency make conventional extraction economically impractical.
Finder needed to turn dormant history into a living database. The first challenge was isolating the small fraction of useful information without moving terabytes unnecessarily. The second was reconnecting identities scattered across years of changing domains and page structures.
The pipeline
Process the signal, not the noise
We built an architecture that reads Internet Archive datasets as compressed, controlled segments. Irrelevant content leaves the pipeline early; only valuable records advance. That choice reduces bandwidth, temporary storage, compute and infrastructure cost before more sophisticated processing even begins.
Specialised extractors identify e-mail addresses, phone numbers, names, domains, roles and companies. Duplicate and partial records are then combined into coherent profiles from traces that would be weak or misleading on their own.
The engineering
Entity resolution at archive scale
The pipeline combines compressed-stream processing, resumable chunks, selective extraction and asynchronously coordinated workers. A failure never forces an entire source file to be processed again; work resumes from a known checkpoint and avoids computation already completed.
The entity-resolution layer connects names, domains, positions, e-mails and phone numbers across time. Multiple signals estimate freshness, reachability and reliability. The result is not merely a found record, but a record supported by provenance and an explicit level of confidence.
The outcome
Years of web history made searchable
Finder converts massive, scattered archives into structured intelligence. Users can recover people, companies and historical relationships that remain invisible in current databases.
At this scale, efficiency comes as much from deciding what not to process as from adding compute. By filtering early and verifying late, the platform turns the historical web into business memory that is useful, traceable and economically viable.


