Skip to content
codeministryPublic

About

Acquisition tool for freelancers: collects project offers from configurable sources, filters them deterministically against your own profile, scores what survives and assembles a ready-to-send application package.

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

269 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

leadGEN / AI

An acquisition tool for freelancers. It collects project offers from sources you configure, throws out everything that was never a fit, and turns the rest into scored, ready-to-send application packages. The rules run first, free and deterministic; the model only sees what is left.

CI License Status Java Spring Boot Angular

The dashboard: what came in, and how much of it survived the hard filter


Status: alpha, and still being built

Read this before you rely on it.

  • It has never run anywhere but its author's own machine and cluster. Releases are tagged and their images published, but the only upgrade path is the Upgrading note in CHANGELOG.md, and there is no promise that a database written today is readable next month.
  • The API and the configuration schema will change without a deprecation period.
  • It is single-operator by design. There is no multi-tenancy. Authentication is optional: under security.auth: oidc every request needs a token from your identity provider, and under the default none what stands in front of the write endpoints is that the server binds 127.0.0.1 and refuses a dotted host name nobody configured, DNS rebinding included; a LAN address or a bare machine name is served without configuration. Do not expose it without oidc.
  • Known gaps are listed under What does not work yet, not hidden.

This was built with AI-assisted coding, on purpose and openly. I need this tool myself: sorting project offers by hand was costing me an hour a day, and an hour a day is not something you spend six months building your way out of. AI coding is what made it possible to have a working pipeline in days instead of months — designed, reviewed and measured by me, written fast with an agent.

Two consequences worth knowing before you judge the code. It reads unusually densely: the comments carry the reasoning and the measurement behind each decision, because that is what keeps a decision from being undone by the next change. And every number quoted in this README was measured rather than estimated — the scripts that measured them are in docs/samples/, and where a claim could not be measured, it says so.


Why

Project offers arrive as newsletters, portal feeds and direct enquiries, and the same project often comes through several agencies at once. Reviewing them costs an hour a day and produces nothing durable. Assembling the documents costs it again for every single application. Both of those are automatable. The decision to apply is not, and it is exactly the part that currently gets the least attention, because the sorting eats the time.

The numbers come from 14 real newsletter mails:

Measured Value What follows from it
Offers in 14 mails 1289 Reading them by hand is the actual cost
Removed by rules alone, no model 80.9 % The expensive stages only ever see the rest
The same project through several portals 14.0 % Collapsing them is worth a stage of its own
Offers stating an hourly rate 0.0 % A rate rule before enrichment filters everything or nothing
Offers stating a remote share 8.8 % Most of what you want to filter on is not in the listing

That last pair is why there is a stage that fetches the original ad, and why the rate rule is configured to run after it — the loader refuses any other value.

How it works

%%{init: {"themeVariables": {"clusterBkg":"#fafafa","clusterBorder":"#c3c8cf","titleColor":"#374151","mainBkg":"#eef1f5","nodeBorder":"#9aa3ad","primaryTextColor":"#1f2937"}}}%%
flowchart LR
    classDef free fill:#dbe4ee,stroke:#4a6d8c,color:#1f2937
    classDef model fill:#e2d5f1,stroke:#6f4aa8,color:#1f2937
    classDef net fill:#d6ead8,stroke:#3d7a48,color:#1f2937
    classDef file fill:#f6dccb,stroke:#b85c2a,color:#1f2937

    s["Sources"] --> i["Ingest"] --> x["Extract"] --> d["Dedupe"] --> f["Filter"] --> a["Archive"] --> e["Enrich"] --> c["Content"] --> sc["Score"] --> p["Package"] --> g["Digest"]

    class s,i,x,f,a free
    class d,c,sc model
    class e net
    class p,g file
Loading

Grey-blue is free and deterministic, violet may ask a model, green is the only stage that leaves the machine, peach writes a file. The chain is drawn in full, with what every stage reads and writes and where a rule decides against where a model speaks, in docs/BACKEND-FLOWS.md.

Two rules decide the shape of everything above.

Rules before model. Six deterministic knockout stages run first, in a fixed order, with no network and no language model: abroad → remote share → out of reach → role or stack → no core skill → contract form. An offer stops at the first rejection and carries that verdict, so the funnel adds up and you can always ask why something is missing. Without an API key the tool still runs; it loses the score total and the cover letter, and an advert is shown as the portal wrapped it — nothing else.

The one place that inverts is content segmentation, and deliberately. A fetched page carries the portal with it, and the part that costs most — the recruiter's standing signature, buried inside the advert's own text — is unreachable by any selector. So there the model decides and the deterministic half is a cache in front of it: every block is hashed, a decision is remembered by that hash, and a portal's repeated furniture is paid for once and free for the eleven thousand adverts that repeat it. Nothing is ever hidden without a positive decision, so no key means nothing is hidden at all.

Nothing is wired in. Not one CSS selector, keyword, weight or portal is written in Java. A new offer source is a block of YAML, including its extraction rules down to the selector and the date format. There is a test that reads the repository for Transport.send, JavaMailSender, setRecipient( and mailto: and fails the build if a send path ever appears.

Ask the corpus. An Ask button in the header opens a chat beside every screen: a question in plain words about the offers, your applications and your profile. It is not a pipeline stage and it only reads — five read-only tools over the same services the screens use, a daily call budget of its own, and your mailbox address never reaches the model. Every offer or application it names becomes a numbered link only when one of that answer's own searches returned it; anything else stays plain text, marked unverified. The reasoning is in docs/decisions/chat.md.

Try it in one command

The repository ships a complete invented dataset — five newsletter mails carrying ~170 fictional offers from portal-a…portal-f and agencies called Acme, Initech and Globex — so a fresh clone opens on a populated application instead of six empty screens.

git clone git@github.com:codeministry/leadgen.git
cd leadgen
cp .env.example .env    # the file has to exist; it may stay exactly as it is
docker compose -f docker-compose.yml -f docker-compose.demo.yml up --build

Open http://localhost:4200, go to Workflow and press Run ingest once. Details, and how to regenerate the corpus, are in demo/README.md.

Two things the demo cannot fake, both by design: without an LLM_API_KEY the shortlist is there and filtered but the score total is withheld rather than computed from half the weights, and enrichment has nothing to fetch because the invented URLs do not resolve. The chat needs a model too: with one configured, Ask appears in the header; without one it is simply not there.

The app can be installed from the browser's own install affordance; there is no button for it inside the app. On the desktop it is the install icon at the right of the address bar in Chrome or Edge (Safari: File › Add to Dock; Firefox does not install web apps), on Android Install app or Add to Home screen in the browser menu, on iOS the share sheet and Add to Home Screen. It opens in its own window with the lead-ring icon and the app's colours, the shell loads without a network (the data never does; it always comes from the API), and a toast offers to reload once a deploy has landed.

Ask it from an MCP client

leadgen is an MCP server too: /mcp speaks Streamable HTTP and serves ten read-only tools, the shortlist search and an offer's detail, the funnel, the last run, the application board and the configuration, and four the chat uses (search by meaning, the statistics, one application, the profile). Point a client at the web port, e.g. for Claude Code:

claude mcp add --transport http leadgen http://localhost:4200/mcp

With AUTH_MODE=oidc the endpoint wants a bearer token like every other; a client that speaks MCP's authorization flow finds the issuer at /.well-known/oauth-protected-resource/mcp. What the tools answer and why is in docs/decisions/mcp.md.

The screens

Shortlist Shortlist — what cleared the hard filter, each entry carrying the reason it scored what it scored, with duplicate portals collapsed into one row. The list keeps its place on the left while the offer opens beside it; reading one no longer costs the list. The bar above it is a search, an order and a set of facets, in three different shapes because they are three different things: what is filtered shows as chips that can be taken off one at a time, and a view worth returning to can be saved under a name.
Offer detail Offer detail — the same screen, further down its reading column: every score reason against what was attainable, the extracted fields, the application package and the status control. The portal's own furniture is folded away behind a line saying how much of it there was, and opens again where it stood.
Pipeline board Pipeline — the half of the loop the tool cannot see. Nothing is sent from here, so everything here is recorded by hand.
Analytics Analytics — what the market is doing, and what the rules are doing to it.
Sources Sources — the configuration rather than the database, so a source that has never run still shows up. Opening one shows the block of sources.yaml that defines it, secrets masked, beside every run it has had and when its numbers last moved.
Workflow Workflow — the run as a workflow, from ingest to digest: each stage with the last run's count, what it costs, and a marker where a model takes part. Picking a stage shows every key that decides it, the knockouts, the weights and the thresholds behind every number on the shortlist, and the prompt as this configuration actually renders it.
Chat Chat — a question in plain words, answered beside the page it was asked on. Each offer it names is a numbered link to a row one of its own searches returned, the sources sit below the answer, and following one opens the offer next to the conversation. Earlier conversations are kept by day and searchable, and can be deleted one at a time or many at once.
Help Help — a drawer from the header that opens at the chapter of the screen it was opened from, with a how-it-works chapter and its diagrams for the reader who wants to know what happens between a mail and a package.

Configuration

Two layers, the same way Spring's own works. Working defaults ship on the classpath under backend/src/main/resources/leadgen/ and are part of the jar; the directory named by leadgen.config-dir overrides them file by file. The tool therefore runs on a fresh clone with no configuration at all, and nothing individual is ever baked into the artifact. The startup banner names, per setting, which layer won and where the value came from — with credentials masked.

File What it decides
sources.yaml Where offers come from, and how a document is read: the block selector, every field, the date format, the tracking-proxy parameter. A new source is a block here.
matching-rules.yaml The six knockout stages, the scoring weights and penalties, the shortlist thresholds, deduplication and follow-up.
skill-profile.yaml Who is applying: skills with weights and aliases, industries, reference projects, and which CV goes with which language.
pipeline.yaml The process: model and provider, enrichment, packaging, digest. Named pipeline.yaml and not application.yaml, which belongs to Spring alone.

Credentials never appear in any of them. Every value is a ${PLACEHOLDER} resolved from .env, which is gitignored, and so is config/. See .env.example for the full list.

What does not work yet

Named rather than hidden, because a gap you find yourself is worse than one you were told about.

  • Reading a mailbox writes to it. The IMAP source is Spring Integration's ImapMailReceiver, which remembers what it has handed over by setting a user flag on each message. Nothing the owner sees is touched — no \Seen, no \Flagged, no \Deleted — but it is a write, and an earlier implementation kept its own cursor and made none. The consequence is that widening a too-narrow subject filter no longer makes the mails behind it reachable again.
  • The exact fingerprint is the normalized title alone, so two genuinely different projects that share a title do merge; the fields that would tell them apart come from enrichment, which runs later. The two embedding_cosine strategies beside it do run now, at thresholds measured against 2222 real adverts rather than guessed — but only when llm.models.embedding names a model of at least 2000 dimensions. Without one the pass is the exact fingerprint and nothing else, which is what a fresh clone does.
  • llm.models.writing is the one key still read by nothing. The cover letter is a Freemarker template. scoring is read by the judge, content by the classifier and fields by the field extractor (both empty means scoring), extraction by the document fallback, and embedding by the two similarity strategies.
  • Content segmentation caches a decision per block, so a mixed block is its weak spot. A block that is nine parts portal furniture and one part per-offer text never repeats, so it never gets a cache hit and costs one model call per advert. Its counters are logged and written per offer but are not yet in pipeline_run, so the dashboard does not show them.
  • The batched scoring path is hand-written HTTP. Spring AI has no batch abstraction, so the half-price asynchronous path talks to the provider directly while the synchronous one goes through ChatClient. Two mechanisms for one question, and it is the reason the provider list for batching is one name long.
  • A cursor table nobody reads. ingest_cursor and the two classes around it are left over from the IMAP connector's previous design.
  • remote.accept_unknown is displayed and never read. Setting it to false changes nothing.
  • No Helm chart. Docker Compose is the supported way to run this.
  • The corpus tests skip on a fresh clone. They assert against 14 real newsletter mails that are gitignored for privacy; ExtractionTest covers the same mechanics against a fixture that ships.

Development

./gradlew check                # both modules: backend tests, frontend lint + tests
./gradlew :backend:test        # Spring tests — needs a running Docker for Testcontainers
./gradlew :backend:bootRun     # API on :8080, reads the untracked .env from the repo root
docker compose up postgres     # the database a local run expects

cd frontend
bun run start                  # dev server on :4200, proxies /api to API_PROXY_TARGET
bun run check:static           # ESLint, Stylelint, tsc — after every change
bun run test                   # Vitest

bun, never npm or npx. The frontend is bracketed into Gradle with plain Exec tasks so package.json stays the single list of frontend commands.

Documentation

docs/README.md is the index: what each document is for, and which one to open for the question you have.

  • Architecture — the pipeline stage by stage, and the reasoning behind the parts that are not obvious
  • Data model — the nineteen tables, their keys, and which class writes each column
  • Backend flows — the run as a sequence, the asynchronous side, every write path from an endpoint down, the state machines, and where a rule decides against where a model speaks
  • Configuration — the two layers, the five files, every variable
  • Adding a source — a new source is a YAML block, worked through line by line, then every key with what reads it
  • Writing rules — the six knockouts, the weight table and the thresholds, and which keys are read by nothing
  • Development — prerequisites, commands, and the traps a newcomer hits first
  • Concept — the original design: module layout and order of work, historical in places
  • Sample analysis — what 14 real newsletter mails contain, and what they do not
  • The demo — the invented dataset and what it demonstrates
  • CLAUDE.md — the working notes: the repo-wide invariants, the file inventory and the measured baseline. Written for an AI agent working in the tree. The conventions and the traps sit beside the code they are about, in backend/CLAUDE.md and frontend/CLAUDE.md.
  • docs/decisions/ — the reasoning behind each stage: every rule with the measurement behind it, and why the obvious alternative was not taken. Together with CLAUDE.md the most complete description of why this code looks the way it does.

Contributing

Issues and pull requests are welcome — see CONTRIBUTING.md. Two invariants a change must not break: nothing is ever sent, and no vendor, portal, provider or personal datum enters a committed file. Both are enforced by tests.

License

Apache License 2.0. See NOTICE.


codeministry
Built and maintained by codeministry.

About

Acquisition tool for freelancers: collects project offers from configurable sources, filters them deterministically against your own profile, scores what survives and assembles a ready-to-send application package.

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages