Over-engineered on purpose
2026-07-27
jobagent scores roughly two thousand job listings a day against my résumé and pushes the
ranked matches to a Telegram bot. It has one user. It runs on one machine, on a nightly
cron, in Docker.
It also has a strictly layered architecture, a provider-agnostic LLM boundary, a golden-set evaluation harness, and per-batch cost logging. That is far more structure than the workload needs, and it was deliberate: I wanted to practise production-shaped engineering on something I actually cared about.
It is easy to argue for over-engineering in general and harder to show what it got you, so this page sticks to what happened: what the structure is, what it cost, what it returned, and the one time it caught me claiming something that was not true.
The architecture, and the router I did not build
Five layers, each depending only on the ones below it:
core/ db engine, LLM provider factory, prompt loader
repositories/ data access, one module per table; SQL lives here and nowhere else
agents/ sources/ domain/ scoring, writing, feedback; external fetch; state machines
pipeline/ notify/ workers/ the graph, the Telegram bot, the async queue
cli.py thin, delegates
Most of it comes down to two rules: no SQL outside repositories/, and no business logic
in core/. The codebase has three separate model-driven behaviours (scoring, writing, and
interpreting feedback), so I wanted one place where SQL could not sprawl and one place
where LLM logic could not quietly grow a dependency on Telegram or Docker.
The payoff is easiest to see in diff size. When I added Redis to move CV generation off
the bot's polling loop, the whole change was one new file under workers/ and two lines
in notify/bot.py. No cross-cutting edits were needed, because the layering left no
place for a cross-cutting concern to leak into.
The pipeline is a LangGraph supervisor with worker nodes, and the supervisor does not use a model to decide anything. It pops the next worker off a queue:
def route(state: AgentState) -> Literal["scraper", "matcher", "tracker", "__end__"]:
if not state["pending"]:
return "__end__"
nxt = state["pending"][0] # ← the entire routing decision
... # (a name check, and a warning if unknown)
Every caller passes a literal list: ["scraper", "matcher", "tracker"] for the nightly
run, and single-element lists for the CLI subcommands. There is no model call anywhere on
the routing path.
The order is fixed, and I know it when I write the code: scrape, then score what was scraped, then review open applications. A model router would add a network round trip, a per-run cost, and a new failure mode, in exchange for picking among three options that never vary. ReAct gets the same answer for the same reason. There is no tool-selection problem here, because no branch depends on what the model found. A swarm or a hierarchy would be worse still. Those buy you dynamic delegation, which a fixed nightly sequence does not need.
LangGraph still helps at this size, in smaller ways. A node is an obvious place to add the next worker, and the state object is a natural container for cross-worker stats that would otherwise become a parameter threaded through everything. That is worth a small framework tax. The routing itself is still a queue.
The rule I would take to a larger system: give the decision to a model when the branch
depends on something only visible at runtime. Until then, an if is faster, cheaper,
testable, and cannot hallucinate a worker name.
The cheapest filter runs first
Two thousand listings a day would be expensive to score naively, so most of them never reach a model at all. A SQL keyword pre-filter drops the obviously irrelevant ones (wrong seniority, or the wrong field entirely) before any LLM call. Only survivors get scored.
What does reach the model goes in batches. The system prompt and my candidate brief are sent once per batch, eight jobs by default, instead of once per job. That alone cuts input tokens roughly sevenfold.
Scoring is also not left entirely to the model. The staffing-firm penalty is applied in code, the score is clamped to 0–100 in code, and the number of points deducted is stored separately from the model's own output:
if is_outsourcing:
score = max(0, raw_score - OUTSOURCING_PENALTY)
reasons["penalty"] = raw_score - score # what the code did, recorded apart
This is the same instinct as the pre-filter, one layer down: decide in code whatever code can decide, and keep what the model said separable from what the system did with it.
Cutting output tokens by 61%
The largest cost win in the project came from deleting three fields.
Auditing the matcher's JSON schema, I found that why_not, skills_matched and
seniority_fit were being written into a jsonb column and never read by any consumer. I
confirmed that by grepping the codebase for every reader before touching anything, then
removed them from the output schema:
| Before | After | |
|---|---|---|
| Output tokens per job | 156 | 61 (−61%) |
| Per-call cost, warm cache | baseline | −44% |
It dominates because of pricing asymmetry: on the model in use, output tokens cost four times what input tokens cost. Input-side optimisation is the more interesting engineering problem, which is exactly why it gets attention first. The boring question (is anything reading this field?) was worth more than every caching experiment on this page combined.
The evaluation harness, and what it found
eval-freeze snapshots already-scored jobs into a golden set with the full job text baked
into the file, so the cases stay reproducible after the database rows move on. eval
re-scores that set with the current prompt, model and learned rules, and reports three
numbers:
pass_rate: cases landing inside the frozen tolerance band. This is the regression signal for prompt edits.- Spearman ρ: rank correlation against the frozen scores. Does the ordering survive?
high_fit_agreement: agreement on thescore ≥ 60"worth applying" cutoff. Does the decision boundary survive, as opposed to the numbers near it?
Before I switch providers, all three have to clear a bar: pass_rate ≥ 80%, ρ ≥ 0.85
and high_fit_agreement ≥ 80%. I check three numbers because a provider swap can keep the
average and wreck the ranking, or keep the ranking and push everything across the cutoff.
The harness calls the real score_batch and the
real post-processing, so it exercises the production scoring path and not a
reimplementation that can drift. And Spearman is written out from first principles, with
tie-averaged ranks, instead of adding scipy for one function. I validated it against
perfect, anti-correlated, noisy and tied inputs before trusting it.
The first run found something I did not expect. At temperature=0, re-scoring the frozen
set did not reproduce itself: 5 of 12 cases landed outside the tolerance band.
Temperature zero constrains sampling, but it does not make a hosted model deterministic
across calls. Every "just set temperature to 0 and diff the outputs" testing strategy I had
in mind quietly assumed otherwise.
What the instrumentation caught
Every matcher batch logs its own economics, which is the only reason any of the following was noticeable:
matcher.batch batch=3 size=8 input=5011 cache_read=1024 hit_pct=20.4 output=478
matcher.done batches=17 tok_in=85K tok_cached=15K cache_hit_pct=17.6 cost_usd=0.048
OpenAI's prompt cache behaved like a coin flip: a tight test loop reported 91%, and the
production matcher sending the same prefix a few seconds later reported 0%. Cache lookup
is not global. Requests fan out across shards, each with its own local cache, so identical
bytes do not guarantee landing where those bytes have been seen. I set a
prompt_cache_key routing hint, measured a clean A/B in its favour, wrote the number down,
and shipped it.
Months later the same script returned 0% on both arms. Total input was 1,073 tokens; the vendor minimum is 1,024, in 128-token blocks. The test had been trimmed to a single tiny fake job so the cacheable prefix would dominate, and that simplification is what suppressed the effect it was built to show. Re-run at production shape, the keyed and unkeyed arms are indistinguishable. I no longer know the effect size of that fix, and the page said I did.
The full write-up is in Two Kinds of LLM Caching: the KV-cache mechanics underneath all of this, a second vendor whose marker turned out not to be supported at all, and a response cache I talked myself out of building.
The cost numbers still hold. Each nightly batch scores different jobs, so the cache only ever covers the static prefix, around 1,024 tokens of roughly 5,500 per batch. Real hit rate on a warm batch is about 18%, and the saving is about $0.15 a month.
What the structure made possible
None of the above is visible without instrumentation that predates the bug. hit_pct
being in that log line is the whole reason there was something to notice. Cost is
computed inline from list pricing instead of in a dashboard later, so a regression shows
up the next morning without anyone deciding to go looking.
The layering mattered in a smaller way. Threading a per-worker cache key through the system
touched exactly one function, because provider access already sat behind a single
get_llm(temperature=None, cache_key=None) boundary. Had the client been constructed at
each call site, the experiment would have been a refactor first, and I probably would not
have bothered running it.
The logging helped twice. The first time, it showed me the anomaly. The second time, months later, the same counters showed me that my own fix had never been properly measured. The second one is the one I care about.
A second, less flattering example
The Telegram bot spent nine days in a retry loop after DNS resolution died inside its
container. Docker reported the container as Up 9 days the entire time. I found it by
accident while investigating something unrelated. The logs had been saying
NetworkError: Name or service not known since day one.
A liveness check that only reports whether the process is running cannot tell you whether it is doing its job. The fix is a heartbeat the bot sends to itself, with a missing heartbeat as the alarm. It is still on the list.
Summary
The layering, the eval harness and the cost logging were over-engineering for a single-user workload, and most of it has not paid for itself in any way I can put a number on.
What it did give me is narrower, but I find it more interesting than a performance
number. The
schema audit found a 61% output-token reduction because there was one place to look. The
eval harness found that temperature=0 is not reproducibility. The cost logging found an
intermittent cache anomaly, and months later found that my own explanation for it did not
hold up.
Instrumentation is the one piece of this you cannot add retroactively, because by the time you want it you are already trying to explain something that has already happened.