An acquisition tool for freelancers. It collects project offers from sources you configure, throws out everything that was never a fit, and turns the rest into scored, ready-to-send application packages. The rules run first, free and deterministic; the model only sees what is left.
Read this before you rely on it.
- It has never run anywhere but its author's own machine and cluster. Releases are tagged and their images published, but the only upgrade path is the Upgrading note in CHANGELOG.md, and there is no promise that a database written today is readable next month.
- The API and the configuration schema will change without a deprecation period.
- It is single-operator by design. There is no multi-tenancy. Authentication is optional:
under
security.auth: oidcevery request needs a token from your identity provider, and under the defaultnonewhat stands in front of the write endpoints is that the server binds127.0.0.1and refuses a dotted host name nobody configured, DNS rebinding included; a LAN address or a bare machine name is served without configuration. Do not expose it withoutoidc. - Known gaps are listed under What does not work yet, not hidden.
This was built with AI-assisted coding, on purpose and openly. I need this tool myself: sorting project offers by hand was costing me an hour a day, and an hour a day is not something you spend six months building your way out of. AI coding is what made it possible to have a working pipeline in days instead of months — designed, reviewed and measured by me, written fast with an agent.
Two consequences worth knowing before you judge the code. It reads unusually densely: the
comments carry the reasoning and the measurement behind each decision, because that is what
keeps a decision from being undone by the next change. And every number quoted in this
README was measured rather than estimated — the scripts that measured them are in
docs/samples/, and where a claim could not be measured, it says so.
Project offers arrive as newsletters, portal feeds and direct enquiries, and the same project often comes through several agencies at once. Reviewing them costs an hour a day and produces nothing durable. Assembling the documents costs it again for every single application. Both of those are automatable. The decision to apply is not, and it is exactly the part that currently gets the least attention, because the sorting eats the time.
The numbers come from 14 real newsletter mails:
| Measured | Value | What follows from it |
|---|---|---|
| Offers in 14 mails | 1289 | Reading them by hand is the actual cost |
| Removed by rules alone, no model | 80.9 % | The expensive stages only ever see the rest |
| The same project through several portals | 14.0 % | Collapsing them is worth a stage of its own |
| Offers stating an hourly rate | 0.0 % | A rate rule before enrichment filters everything or nothing |
| Offers stating a remote share | 8.8 % | Most of what you want to filter on is not in the listing |
That last pair is why there is a stage that fetches the original ad, and why the rate rule is configured to run after it — the loader refuses any other value.
%%{init: {"themeVariables": {"clusterBkg":"#fafafa","clusterBorder":"#c3c8cf","titleColor":"#374151","mainBkg":"#eef1f5","nodeBorder":"#9aa3ad","primaryTextColor":"#1f2937"}}}%%
flowchart LR
classDef free fill:#dbe4ee,stroke:#4a6d8c,color:#1f2937
classDef model fill:#e2d5f1,stroke:#6f4aa8,color:#1f2937
classDef net fill:#d6ead8,stroke:#3d7a48,color:#1f2937
classDef file fill:#f6dccb,stroke:#b85c2a,color:#1f2937
s["Sources"] --> i["Ingest"] --> x["Extract"] --> d["Dedupe"] --> f["Filter"] --> a["Archive"] --> e["Enrich"] --> c["Content"] --> sc["Score"] --> p["Package"] --> g["Digest"]
class s,i,x,f,a free
class d,c,sc model
class e net
class p,g file
Grey-blue is free and deterministic, violet may ask a model, green is the only stage that
leaves the machine, peach writes a file. The chain is drawn in full, with what every stage
reads and writes and where a rule decides against where a model speaks, in
docs/BACKEND-FLOWS.md.
Two rules decide the shape of everything above.
Rules before model. Six deterministic knockout stages run first, in a fixed order, with no network and no language model: abroad → remote share → out of reach → role or stack → no core skill → contract form. An offer stops at the first rejection and carries that verdict, so the funnel adds up and you can always ask why something is missing. Without an API key the tool still runs; it loses the score total and the cover letter, and an advert is shown as the portal wrapped it — nothing else.
The one place that inverts is content segmentation, and deliberately. A fetched page carries the portal with it, and the part that costs most — the recruiter's standing signature, buried inside the advert's own text — is unreachable by any selector. So there the model decides and the deterministic half is a cache in front of it: every block is hashed, a decision is remembered by that hash, and a portal's repeated furniture is paid for once and free for the eleven thousand adverts that repeat it. Nothing is ever hidden without a positive decision, so no key means nothing is hidden at all.
Nothing is wired in. Not one CSS selector, keyword, weight or portal is written in
Java. A new offer source is a block of YAML, including its extraction rules down to the
selector and the date format. There is a test that reads the repository for Transport.send,
JavaMailSender, setRecipient( and mailto: and fails the build if a send path ever
appears.
Ask the corpus. An Ask button in the header opens a chat beside every screen: a
question in plain words about the offers, your applications and your profile. It is not a
pipeline stage and it only reads — five read-only tools over the same services the screens use,
a daily call budget of its own, and your mailbox address never reaches the model. Every offer or
application it names becomes a numbered link only when one of that answer's own searches
returned it; anything else stays plain text, marked unverified. The reasoning is in
docs/decisions/chat.md.
The repository ships a complete invented dataset — five newsletter mails carrying ~170
fictional offers from portal-a…portal-f and agencies called Acme, Initech and Globex —
so a fresh clone opens on a populated application instead of six empty screens.
git clone git@github.com:codeministry/leadgen.git
cd leadgen
cp .env.example .env # the file has to exist; it may stay exactly as it is
docker compose -f docker-compose.yml -f docker-compose.demo.yml up --buildOpen http://localhost:4200, go to Workflow and press Run ingest once. Details, and how to regenerate
the corpus, are in demo/README.md.
Two things the demo cannot fake, both by design: without an LLM_API_KEY the shortlist is
there and filtered but the score total is withheld rather than computed from half the
weights, and enrichment has nothing to fetch because the invented URLs do not resolve.
The chat needs a model too: with one configured, Ask appears in the header; without one it
is simply not there.
The app can be installed from the browser's own install affordance; there is no button for it inside the app. On the desktop it is the install icon at the right of the address bar in Chrome or Edge (Safari: File › Add to Dock; Firefox does not install web apps), on Android Install app or Add to Home screen in the browser menu, on iOS the share sheet and Add to Home Screen. It opens in its own window with the lead-ring icon and the app's colours, the shell loads without a network (the data never does; it always comes from the API), and a toast offers to reload once a deploy has landed.
leadgen is an MCP server too: /mcp speaks Streamable HTTP and serves ten read-only tools, the
shortlist search and an offer's detail, the funnel, the last run, the application board and the
configuration, and four the chat uses (search by meaning, the statistics, one application, the
profile). Point a client at the web port, e.g. for Claude Code:
claude mcp add --transport http leadgen http://localhost:4200/mcpWith AUTH_MODE=oidc the endpoint wants a bearer token like every other; a client that speaks
MCP's authorization flow finds the issuer at /.well-known/oauth-protected-resource/mcp. What the
tools answer and why is in docs/decisions/mcp.md.
Two layers, the same way Spring's own works. Working defaults ship on the classpath under
backend/src/main/resources/leadgen/ and are part of the jar; the directory named by
leadgen.config-dir overrides them file by file. The tool therefore runs on a fresh
clone with no configuration at all, and nothing individual is ever baked into the artifact.
The startup banner names, per setting, which layer won and where the value came from —
with credentials masked.
| File | What it decides |
|---|---|
sources.yaml |
Where offers come from, and how a document is read: the block selector, every field, the date format, the tracking-proxy parameter. A new source is a block here. |
matching-rules.yaml |
The six knockout stages, the scoring weights and penalties, the shortlist thresholds, deduplication and follow-up. |
skill-profile.yaml |
Who is applying: skills with weights and aliases, industries, reference projects, and which CV goes with which language. |
pipeline.yaml |
The process: model and provider, enrichment, packaging, digest. Named pipeline.yaml and not application.yaml, which belongs to Spring alone. |
Credentials never appear in any of them. Every value is a ${PLACEHOLDER} resolved from
.env, which is gitignored, and so is config/. See
.env.example for the full list.
Named rather than hidden, because a gap you find yourself is worse than one you were told about.
- Reading a mailbox writes to it. The IMAP source is Spring Integration's
ImapMailReceiver, which remembers what it has handed over by setting a user flag on each message. Nothing the owner sees is touched — no\Seen, no\Flagged, no\Deleted— but it is a write, and an earlier implementation kept its own cursor and made none. The consequence is that widening a too-narrow subject filter no longer makes the mails behind it reachable again. - The exact fingerprint is the normalized title alone, so two genuinely different projects
that share a title do merge; the fields that would tell them apart come from enrichment,
which runs later. The two
embedding_cosinestrategies beside it do run now, at thresholds measured against 2222 real adverts rather than guessed — but only whenllm.models.embeddingnames a model of at least 2000 dimensions. Without one the pass is the exact fingerprint and nothing else, which is what a fresh clone does. llm.models.writingis the one key still read by nothing. The cover letter is a Freemarker template.scoringis read by the judge,contentby the classifier andfieldsby the field extractor (both empty meansscoring),extractionby the document fallback, andembeddingby the two similarity strategies.- Content segmentation caches a decision per block, so a mixed block is its weak spot. A block that is nine parts
portal furniture and one part per-offer text never repeats, so it never gets a cache hit and costs one model call per
advert. Its counters are logged and written per offer but are not yet in
pipeline_run, so the dashboard does not show them. - The batched scoring path is hand-written HTTP. Spring AI has no batch abstraction, so
the half-price asynchronous path talks to the provider directly while the synchronous one
goes through
ChatClient. Two mechanisms for one question, and it is the reason the provider list for batching is one name long. - A cursor table nobody reads.
ingest_cursorand the two classes around it are left over from the IMAP connector's previous design. remote.accept_unknownis displayed and never read. Setting it tofalsechanges nothing.- No Helm chart. Docker Compose is the supported way to run this.
- The corpus tests skip on a fresh clone. They assert against 14 real newsletter mails
that are gitignored for privacy;
ExtractionTestcovers the same mechanics against a fixture that ships.
./gradlew check # both modules: backend tests, frontend lint + tests
./gradlew :backend:test # Spring tests — needs a running Docker for Testcontainers
./gradlew :backend:bootRun # API on :8080, reads the untracked .env from the repo root
docker compose up postgres # the database a local run expects
cd frontend
bun run start # dev server on :4200, proxies /api to API_PROXY_TARGET
bun run check:static # ESLint, Stylelint, tsc — after every change
bun run test # Vitestbun, never npm or npx. The frontend is bracketed into Gradle with plain Exec tasks so
package.json stays the single list of frontend commands.
docs/README.md is the index: what each document is for, and which one to
open for the question you have.
- Architecture — the pipeline stage by stage, and the reasoning behind the parts that are not obvious
- Data model — the nineteen tables, their keys, and which class writes each column
- Backend flows — the run as a sequence, the asynchronous side, every write path from an endpoint down, the state machines, and where a rule decides against where a model speaks
- Configuration — the two layers, the five files, every variable
- Adding a source — a new source is a YAML block, worked through line by line, then every key with what reads it
- Writing rules — the six knockouts, the weight table and the thresholds, and which keys are read by nothing
- Development — prerequisites, commands, and the traps a newcomer hits first
- Concept — the original design: module layout and order of work, historical in places
- Sample analysis — what 14 real newsletter mails contain, and what they do not
- The demo — the invented dataset and what it demonstrates
CLAUDE.md— the working notes: the repo-wide invariants, the file inventory and the measured baseline. Written for an AI agent working in the tree. The conventions and the traps sit beside the code they are about, inbackend/CLAUDE.mdandfrontend/CLAUDE.md.docs/decisions/— the reasoning behind each stage: every rule with the measurement behind it, and why the obvious alternative was not taken. Together withCLAUDE.mdthe most complete description of why this code looks the way it does.
Issues and pull requests are welcome — see CONTRIBUTING.md. Two invariants a change must not break: nothing is ever sent, and no vendor, portal, provider or personal datum enters a committed file. Both are enforced by tests.
Apache License 2.0. See NOTICE.
Built and maintained by codeministry.









