AIdeazz Blog About Portfolio

The job agent said "I Act TODAY" 26 times — and delivered two jobs

· by

A field note from the AIdeazz AI Lab — a real incident on a live production system, written up from the logs. October 6, 2026.

A job-discovery agent kept logging that it had put strong matches in front of the operator, every twelve hours, for a week. Her action queue received nothing new for three days. The success line was printed before the CRM answered, the best regional source had been refusing every request while reporting "0 jobs found" with a green tick, and approved jobs were being dropped or folded into deals she had already closed.

What it looked like from outside

The operator noticed that the agent had not delivered a single new position to apply to in days. Every component looked healthy. The processes were running and restarting cleanly, the search door ran twice a day and logged "Done — new jobs: 30", and its log held a steady stream of lines reading "IRON-CLAD FIT + judge OK -> I Act TODAY". The hourly engine reported each source with a green tick. The CRM told a different story. The last new job to enter her action queue had arrived three days earlier.

What was actually happening

Four faults, each hidden by a success message. First, the search door printed "I Act TODAY" before it pushed the job to the CRM and never read the answer. Since 29 September it printed that line 26 times, for 10 distinct jobs, and exactly 2 became new deals. Twelve of those approvals carried a blank company name, because the code read the company from the subdomain of a Lever link ("jobs") instead of its path, and the CRM endpoint rejected each one with HTTP 400 that nobody read. Twelve more matched deals that already existed, most of them one she had closed as not a fit, so the CRM added another note and left the deal where it was. Second, the regional job board that had produced most of her good matches began refusing every request on 2 October with HTTP 400. The fetch code skipped non-200 responses and swallowed exceptions, then printed "✅ 0 jobs found" every hour. Third, the local seen-list was trimmed in hash order rather than by age, so recent jobs fell out and came back as "new" every twelve hours: 488 processed rows were only 379 distinct jobs. Fourth, a third of the paid searches failed on an empty or non-JSON body with no retry, and each failure became "→ 0 results".

The fix

Every success line now waits for the system that has to act on it. The search door reads the CRM's answer and logs "I Act TODAY" only for a new deal; a duplicate is logged as "already in CRM (stage …) — not new", a rejection as "CRM REJECTED (HTTP code: message)", and a 2xx with no deal id as unknown. The CRM side now answers with whether the job was a duplicate and whether the operator had already decided it, and a re-sighted job she decided is left untouched: no note, no queue entry, no paid cover-letter call. A reply or interview invite on that same deal still lands. The company is read from the link path, with a conservative title fallback measured against 890 real blank-company titles. The seen-list drops the oldest entry. Empty search bodies are retried once. The refusing job board was not worked around: it now serves only its own client, and its terms forbid scraping, so its failure is logged as a failure and the decision about it went to the operator. A regional placement agency with stated pay and eligibility was added as a source instead. The same change aligned the agent's targeting with the role list in the operator's published professional profile, adding eleven titles that no layer of the agent had ever searched for.

How I know it worked

Against the CRM, not the log. Before the change, 26 "I Act TODAY" lines produced 2 new deals, and the last new deal was three days old. In the first five minutes after deployment, eight new jobs entered her action queue. A replay of the job she had closed returned "duplicate, decided, closed lost", and the CRM log changed from "already exists — refresh action note" to "already decided — left alone". The refusing source now logs "❌ 0 jobs — 3/3 requests failed (HTTP 400)" instead of a green tick. The full evaluation suite passed on the production server (862 tests). A before/after replay of the AI judge approved in-lane roles 3 of 3 and rejected out-of-lane gigs 3 of 3.

The rule this earned

A success line must be printed by whoever can see the result, never by whoever sent the request. If a component logs "done" before the downstream system answers, the log measures effort, not outcome. Count the outcome where it lands: here, new deals in the CRM per day against "approved" lines per day. When the two numbers diverge, the gap is the bug. A source that returns zero needs to say whether it found nothing or failed to ask. "✅ 0" and "❌ 0" are different facts, and a dashboard that cannot tell them apart will stay green through an outage.

The named concepts behind it

Naming a failure mode is what makes it possible to recognise the same shape somewhere new, before it costs another weekend.

Acknowledgement is not completion

A receipt proves delivery. It never proves processing.

When you hand work to something asynchronous -- a queue, a webhook, a workflow tool, a background job -- the response you get back means "I have received this". It does not mean "I have done this", and very often it does not even mean "I intend to do this".

This is the trap behind a large share of "the data just vanished" incidents. The sending side logs a success, the receiving side never processes anything, and both halves look healthy in isolation. A queue that accepts your message and never reads it looks exactly like one that works.

Defences, in order of strength:

1. Do not branch on the acknowledgement. If your fallback logic reads "if the handoff failed, do it myself", it will never run, because the handoff reports success. Make the local path unconditional and let idempotency absorb the duplicate.
2. Confirm from the other side. Check that the work actually completed -- a status endpoint, a result record, a callback -- rather than trusting the receipt.
3. Set a deadline. If the expected outcome has not appeared within N minutes, treat it as failed and act, rather than waiting forever.

Silent failure

The system did something reasonable, and told nobody.

The most expensive bug class there is, because the clock keeps running while everyone assumes things are fine.

A silent failure is not a crash. A crash is loud and gets fixed. A silent failure is a component making a defensible local decision -- drop this message, skip this record, return an empty string -- that nobody downstream is told about. From the outside, a system that is working perfectly and a system that is completely dead can produce the identical observation: nothing happened.

The defence is not "add more logging". It is to make the healthy state provable, so that "nothing happened" can be distinguished from "nothing was supposed to happen". Two things do that:

Verify from logs, not config

Configuration tells you what somebody intended. Logs tell you what happened.

A setting, an environment variable or a present API key is a statement of intent. It is evidence that somebody meant for a behaviour to occur. It is not evidence that the behaviour occurs.

The gap between the two is where the longest outages live, because reading the configuration feels like verification. It produces confident, wrong statements: the key is set, so the provider works; the schedule says every fifteen minutes, so it runs every fifteen minutes; the file was deployed, so the new code is running.

Each of those has a cheap, decisive check that costs seconds:

The rule this earns: never report a system's behaviour from its configuration. Grep the line that proves the behaviour happened, and quote it.

---

This note is one entry in a running wiki of production engineering lessons — every concept linked to the incident that taught it — at aideazz.xyz/ai-ops-wiki.html.

No customer data, credentials, hostnames or internal record identifiers appear in these write-ups.