My cto-aipa process restarted 181 times in the last 2 days. This isn't a silent failure; it's a symptom of a system designed for relentless recovery, where a single shared document, NOW.md, acts as the primary AI agent communication protocol. When direct chat between agents like Cursor Cloud, Cursor Desktop, and Claude Code is impossible, this file becomes the critical shared memory. It's not documentation; it's the live state, the instruction set for the next agent to pick up the task. This setup ensures that even with frequent restarts, work progresses, albeit with a higher operational cost in terms of compute cycles and log noise.
The Constraint: Disconnected Agents, Shared State
The core challenge I face with my AI agents is their inherent disconnection. I use Cursor Cloud, Cursor Desktop, and Claude Code, but they cannot communicate directly. There's no shared conversation history, no Claude MCP in Cursor, no way to send a message from one agent to another. This is a fundamental architectural constraint. Without a direct AI agent communication protocol, I needed a robust mechanism for coordination.
My solution is NOW.md. This file lives in /home/ubuntu/cto-aipa/docs/oracle/NOW.md and serves as the sole shared memory. It's the working memory for whichever agent isn't currently running, and it defines the protocol for how these disconnected agents avoid stepping on each other's toes. The file explicitly states: "Cursor Cloud, Cursor Desktop and Claude Code all work this repo and none of them can see each other's chats... The only things all of them read are HubSpot and this file." This makes NOW.md the central nervous system for my AI operations.
The cto-aipa Restart Pattern
Looking at pm2 jlist, the cto-aipa process shows 181 restarts and has been up for 2 days. This contrasts sharply with other stable processes like algom-poll, which has 0 restarts and has been up for 66 days, or n8n, also with 0 restarts and up for 50 days. Even algom-stream, with its 55193 restarts over 47 days, operates differently; its restarts are often due to external data stream volatility. The cto-aipa restarts, however, are frequently tied to its role as a central orchestrator, processing instructions from NOW.md and interacting with external systems.
Each restart means the agent re-reads its state, including NOW.md, and re-evaluates its next action. This is a deliberate design choice. Instead of building complex, stateful agents that might get stuck or require manual intervention after an error, I opted for stateless, resilient agents that can crash and recover by re-initializing from a shared, persistent state. The 181 restarts indicate that the agent is encountering conditions that trigger a reset, but the system as a whole continues to function.
NOW.md as the Operational Journal
The content of NOW.md isn't just a static configuration; it's a dynamic operational journal. It contains directives, observations, and the current status of tasks. For example, the latest NOW.md entry begins: "NOW — the shared session between Cursor and Claude Code". It then details the communication constraint and establishes the file as the protocol.
When cto-aipa restarts, it reads this file to understand the current context and what needs to be done. This is critical for tasks like the apply-queue.log, which shows "✓ sent to Telegram (2 new)" multiple times, or concierge-selftest.log reporting "✅ PASS — 4 checks, 4050ms to first card". These outcomes are driven by instructions or state changes recorded in NOW.md or inferred from HubSpot, which is the other shared data source. The agent's ability to pick up exactly where it left off, or to re-evaluate the situation, is directly enabled by this shared, persistent state.
Impact on System Stability and Development Velocity
While 181 restarts in 2 days might seem like instability, it's a controlled instability. The cto-aipa process consumes 172 MB of memory, which is manageable. Other processes like n8n use 492 MB, and dragontrade-main uses 148 MB. The memory footprint of cto-aipa is not indicative of a runaway process.
Development velocity is also influenced. I made 5 commits to aideazz in the last 48 hours, including 92b3a74 ("ai-ops-wiki: refresh journal + AEO surfaces") and a49193a ("feat(api): footer 'About' opens Elena's Professional Outlook deck (EN/ES)"). There was also 1 commit to VibeJobHunterAIPA_AIMCF (2c317c8). These commits often involve updates to agent logic or the NOW.md protocol itself. The frequent restarts of cto-aipa mean that new code or updated instructions from NOW.md are picked up quickly, enabling rapid iteration and deployment without a formal redeployment process for each agent. This is a trade-off: higher restart count for faster iteration and resilience to individual agent failures.
The Cost of Resilience: Log Noise and Resource Usage
The downside of this restart-driven resilience is the increased log noise and potential for higher resource usage. Each restart generates log entries, which can make debugging more challenging if not properly managed. For example, the algom-stream process, with its 55193 restarts, generates a significant volume of logs. While cto-aipa's 181 restarts are far fewer, they still contribute to the overall operational overhead.
However, the alternative — a complex, distributed messaging system for AI agent communication — would introduce its own set of challenges, including increased development time, higher infrastructure costs, and greater complexity in debugging. Given my zero VC funding constraint, the NOW.md approach, while imperfect, is a pragmatic and effective solution for coordinating disconnected AI agents on Oracle Cloud. It allows me to ship production AI agents and manage a system that processes tasks like sending 138 deals to "They replied" stage in HubSpot, even with frequent individual agent resets.
Frequently Asked Questions
Q: Why not use a message queue or a database for AI agent communication?
A: My current setup prioritizes simplicity and low overhead given zero VC funding. A message queue or database would introduce additional infrastructure, maintenance, and development complexity that I do not have the resources to manage effectively at this stage.
Q: Does the NOW.md file ever get corrupted or have merge conflicts?
A: NOW.md is primarily written by me, the human operator, and read by the agents. This minimizes merge conflicts. Agents append to logs or update specific fields in HubSpot, which is the other shared state, rather than directly modifying NOW.md in a conflicting manner.
Q: How do you monitor the state of the NOW.md file and agent activity?
A: I monitor NOW.md directly via tail commands and git log to see recent changes. Agent activity is monitored through their respective logs, such as apply-queue.log and concierge-selftest.log, and by observing changes in HubSpot.