My AI job hunter agent was stuck in a rejection loop. It processed feedback, summarized reasons for rejection, and fed them back into its decision-making prompt. This loop was proven end-to-end a month ago. Yet, the same types of jobs kept appearing. The core issue wasn't a lack of feedback, but an inability for that feedback to override the agent's initial criteria. The system was designed to learn, but its learning was advisory, not directive.
The Judge's Limited Influence
The agent, named VibeJobHunterAIPA_AIMCF, had a "judge" component. This judge was supposed to evaluate job postings against my preferences. I committed e4c7314 on 2026-09-28 to fix the judge, aiming for it to evaluate through its own AI environment. Another commit, 4373d73, addressed a specific location issue: "a location our Torre adapter wrote is not the employer promising LATAM." This led to d2b9aa7, which implemented a LATAM overrule, releasing 3 correct vetoes. These changes show an attempt to refine the judge's criteria.
However, the wiki incident from 2026-09-28, titled "The job hunter learned from every rejection — and none of it could change what it showed," clearly states the problem. The agent read operator rejection notes and screenshots hourly, summarizing reasons for its judge's prompt. Despite this, "the same kinds of job kept arriving." The lessons reached a reviewer on a side path and were labeled as "unable to override its criteria." This means the feedback loop was active, but the agent's foundational criteria were immutable.
Memory Window Constraints
A critical constraint was the agent's memory window. The wiki incident notes that the feedback was "drawn from a window that remembered eight days." This limited context meant that while the agent could recall recent rejections, it couldn't build a long-term understanding of what was consistently undesirable. If a job type was rejected for nine days, then reappeared, the agent might treat it as a new, un-rejected type.
I tried to address memory with commit 5782f3d on 2026-09-28: "every decision keeps the posting it was made on — the replay can finally measure." This was a step towards persistent memory for decisions. I also experimented with RAG (Retrieval Augmented Generation) for the judge's decisions, committing 45f6411 on 2026-09-28: "RAG over her decisions for the judge — built, measured, OFF (it did not beat the baseline)." This RAG implementation was ultimately disabled because it didn't improve performance. The challenge wasn't just storing decisions, but effectively integrating them into the agent's active reasoning to override initial criteria.
The Cost of Fragmented Context
My AI agents, including the job hunter, operate in a fragmented environment. The NOW.md file, located at /home/ubuntu/cto-aipa/docs/oracle/NOW.md, serves as a shared session between Cursor Cloud, Cursor Desktop, and Claude Code. As stated in NOW.md, "Cursor Cloud, Cursor Desktop and Claude Code all work this repo and none of them can see each other's chats." There's "no shared conversation, no Claude MCP in Cursor, no way to send the other agent a message." This means that while I, the operator, can use NOW.md as a working memory, the agents themselves lack a unified, persistent context for learning and adaptation.
This fragmentation directly impacts the job hunter's ability to learn from its rejection loop. Each interaction, each feedback point, might be processed in isolation without a holistic view of past rejections across different agent instances or sessions. The VibeJobHunterAIPA_AIMCF repository saw 12 commits in the last 48 hours, indicating active development. One commit, b38836e, fixed an issue where "httpx logging every Telegram request — the bot token was in the journal 8,640x/day," highlighting the volume of interactions and the need for careful logging.
Operational Stability and Agent Restarts
The pm2 jlist output shows the operational status of my processes. serpapi-jobs, which likely feeds job postings to the agent, has had 1 restart and is up for 0 days. The cto-aipa process, a core AI component, has had 170 restarts and is also up for 0 days. In contrast, algom-poll has 0 restarts and is up for 62 days. The high restart count for cto-aipa suggests instability or frequent deployments. Each restart could potentially reset the agent's short-term memory or context, further hindering its ability to learn from the AI Agent Job Hunter Rejection Loop.
The aideazz repository, which includes my wiki, had 5 commits in the last 48 hours. One commit, 2922017, specifically mentions "wiki: the lessons that could only advise - advisory feedback (new concept)," directly referencing the problem with the job hunter. This indicates that the inability of feedback to override criteria is a recognized, documented issue within my operational knowledge base.
Frequently Asked Questions
Q: How was the "rejection" feedback collected and delivered to the agent?
A: The agent read operator rejection notes and screenshots hourly. These were summarized and then fed into the judge's prompt, creating a feedback loop.
Q: What specific mechanism prevented the feedback from overriding the initial criteria?
A: The wiki incident states the feedback was "labeled as unable to override its criteria." This implies a hard-coded or deeply embedded set of initial preferences that the dynamic feedback mechanism could not modify.
Q: Did the RAG implementation for memory improve the agent's performance?
A: No, the RAG implementation (45f6411) was built and measured, but ultimately turned OFF because "it did not beat the baseline."
Q: What was the impact of the 170 restarts on the cto-aipa process?
A: While I do not have specific measurements on the impact of restarts on learning, frequent restarts (170 for cto-aipa, up 0 days) can lead to loss of short-term memory or context, potentially hindering an agent's ability to learn from a feedback loop.