artenis.alija
ende
05 / Use case

Lead Generation Pipeline with Deduplication | Artenis Alija

Business discovery at scale, deduplicated by company rather than by listing.

Business discovery at scale produces enormous duplication. The same company appears once per branch, so a raw export of a multi-city search is mostly repeats.

Long-running scrapes fail partway through, and without incremental output that means starting over.

A flat file is the easiest export format but the worst format for querying, while a database is the opposite.

At a glance

Client
Internal tooling, reused across engagements
Status
Built and deployed
Stack
Python · Docker Compose · PostgreSQL · Playwright
Deduplicated
by company, not by listing
Resumable
runs with incremental output
Dual output
portable file plus queryable database

How it was built

Each step in the order it was actually built, and why that order mattered.

Company-level deduplication

Deduplication keyed on company identity rather than listing identity, so a business with offices in several cities is written once. The behaviour is configurable for cases where multi-location records are actually wanted.

Incremental, resumable output

Results appended as they are found, with completed and blocked jobs tracked separately, so an interrupted run resumes rather than restarts.

Both file and database output

A flat file for portability plus an automatic mirror into a relational database for querying, rather than forcing a choice between them.

Containerised with tunable concurrency

Runs under container orchestration with concurrency, caps and timeouts configurable per run.

How a project runs

The same sequence every time, whichever service or market is involved. It is deliberately front-loaded: most of the risk in an automation project sits in understanding the process, not in building it.

01

Map the process before writing anything

The first session is spent on how the work actually happens, which is almost never how the documented process says it happens. Who touches what, in which order, and where the time really goes. Most failed automation projects failed here rather than in the build, because they automated the described process instead of the real one.

02

Measure the cost of doing nothing

Hours per week, error rate, and what those hours would otherwise be worth. This is what decides whether a process is worth automating at all — and it is also the number you compare against afterwards, which is why it gets recorded before anything is built rather than estimated after.

03

Build the smallest useful version

One process, working end to end, in production, before anything else starts. A narrow system that people actually use beats a broad one that waits on a second phase, and the edge cases that matter only surface once real work runs through it.

04

Run it against reality

The first two weeks of live use produce more design corrections than any amount of planning. Failures get surfaced loudly, retried and logged, because silent failure is the most expensive property a workflow can have and the one noticed last.

05

Hand it over properly

Documentation, credentials, and a walkthrough with whoever will maintain it. A system only one person understands is a liability regardless of how well it runs, so handover is part of the work rather than an optional extra at the end.

Ways to work together

Three arrangements cover almost every engagement. Most start with the first or the second; the third only makes sense once something is live.

A

Fixed-scope project

One defined process, a fixed price and an agreed definition of done. The right fit when the problem is clear and bounded — an order flow to connect, a CRM to build, a reporting pack to automate. Most first engagements are this, because it lets both sides find out how the other works without a long commitment.

B

Assessment first

One to two weeks mapping processes and measuring where the hours actually go, ending in a ranked list with effort and payback estimates. Useful when there is a backlog of automation ideas and no agreement on which matters. The document stands on its own and is yours whether or not you build anything with me.

C

Ongoing retainer

A recurring block of time for maintenance, extension and new automations once systems are live. Integrations break when the systems either side of them change, and a retainer means that gets fixed before it becomes an outage rather than after.

How the working relationship is set up

01

Remote, with real overlap

Work is delivered remotely. Across Europe and the Nordics the working day is effectively identical; in the Gulf it starts three hours ahead, which still leaves your full morning covered. There is no local office in any market, and none is claimed anywhere on this site.

02

You own what gets built

Source code, infrastructure and data stay yours. Systems are deployed on infrastructure you control — your server, a European provider, or a VPS in your own account. There is no per-seat licence and no dependency on me continuing to be involved.

03

Self-hosting is a first-class option

Self-hosted n8n, self-hosted databases and locally run models are all supported and, in several of these markets, preferred. Where no data may reach a third-party API, that constraint shapes the architecture from the start rather than being retrofitted.

04

Direct contact, one person

You deal with the person building the system. There is no account manager relaying requirements, which is the main practical advantage an independent consultant has over an agency at this size — and the main reason scope stays honest.

Frequently asked questions

What if we are not sure automation is the right answer?

Then the assessment is the right starting point, and it is designed to be able to conclude that you should not automate something. A process that is broken should be fixed before it is automated, and one that runs twice a month rarely earns the build. Receiving that answer in week one is far cheaper than discovering it after a project.

We have been burned by a failed automation project before.

That is common, and the cause is usually scoping or adoption rather than technology — a system built for the documented process rather than the real one, or one nobody was trained to maintain. Both are addressed by mapping the real process first and treating handover as part of the work.

How do we avoid depending on one person?

By owning everything: source code, infrastructure, credentials and documentation, with a walkthrough for whoever maintains it. The test is whether another developer could pick the system up from the repository and the documentation alone, and that is the standard handover is written to.

Is our data safe?

It stays where you need it to. Systems can run entirely inside your own infrastructure, including self-hosted models where no data may reach a third-party API. Where GDPR applies, data stays in the EU by default, with named access control and audit logging as standard rather than as an upgrade.

How quickly can something be running?

A first working version of a single process is typically weeks rather than months. Larger platforms are sequenced as modules so something is in production early and the rest builds on a foundation that already survives real use.

Services involved

Sector

Get in touch

Tell me what needs automating

Describe the process that is costing you time and roughly how much. I reply to every enquiry personally, usually within one working day.

Response
Usually within one working day, Mon–Fri CET
Delivery
Remote across Europe, the Nordics and the Gulf
Or email inquiries@artenisalija.com