AIdeazz Blog About Portfolio

PM2 Process Restarts: From Zero to 55193

· by

My PM2-supervised processes show 9 online of 9, which is good. However, a closer look at the restart counts reveals a spectrum of operational stability, from perfectly stable to relentlessly failing. This isn't just about uptime; it's about the hidden costs of recovery and the silent failures that don't always trigger alerts.

The Algom-Stream Anomaly: 55193 Restarts

The algom-stream process stands out with 55193 restarts. This process has been up for 54 days, consuming 51 MB of memory. This number is not a minor glitch; it represents a persistent, underlying issue that PM2 is diligently masking through automatic restarts. While the process is technically "online," its operational efficiency and the integrity of its output over 54 days are severely compromised. This kind of restart count indicates a fundamental flaw in the application logic or its environment, not a transient network blip. It's a system that's constantly falling over and being propped back up, burning CPU cycles and potentially losing state with each restart.

CTO-AIPA: 198 Restarts in Under a Day

Another process, cto-aipa, shows 198 restarts and has been up for 0 days, using 212 MB. This is a different kind of problem than algom-stream. While algom-stream has accumulated its restarts over 54 days, cto-aipa hit 198 restarts within a single day. This suggests a more immediate and critical failure mode. It could be a new deployment with a bug, a dependency issue, or an external service it relies on failing. The fact that it's still "online" means PM2 is doing its job, but 198 restarts in such a short period demands immediate investigation. I recently committed 3 times to aideazz in the last 48 hours, and 1 commit to VibeJobHunterAIPA_AIMCF. These commits, specifically 273a03d and 367efa6 (both related to ai-ops-wiki: refresh journal + AEO surfaces), and fbff8da (related to responses: boardy.ai is an AI networking assistant, never an employer reply), could be related to changes that impact cto-aipa's stability.

Stable Operations: Zero Restarts

On the other end of the spectrum, algom-poll and n8n demonstrate remarkable stability with 0 restarts. algom-poll has been up for 73 days, using 68 MB, and n8n for 57 days, using 495 MB. These processes represent the ideal state: applications running continuously without interruption, indicating robust code, stable dependencies, and predictable environments. They serve as a benchmark for what's achievable and highlight the severity of the issues in processes like algom-stream and cto-aipa.

Moderate Restarts: Daily and Weekly Cycles

Several other processes show moderate restart counts:

The Impact of External Factors: HubSpot Rate Limits

While PM2 handles restarts, the root causes are often external. For example, my hs-watch-manual-emails.log shows a 429 error from HubSpot: "You have reached your ten_secondly_rolling limit." The correlation ID is 01a12209-739c-7d10-9041-72efca9b280a, and the policy name is TEN_SECONDLY_ROLLING. This kind of external rate limiting can cause processes that interact with HubSpot to crash and restart. If cto-aipa or other processes are hitting this limit, it could explain their restart behavior. The reply-radar.log shows APPLY — scanned 247 · automated/own skipped 0 · no CRM match 0 · REPLIES MATCHED 0 · errors 0, indicating that some processes are interacting with the CRM, and these interactions can be subject to such limits. Currently, there are 120 deals at the "They replied" stage, which means active CRM engagement.

Frequently Asked Questions

Q: Does a high restart count always mean a bug in my code?
A: Not always. While algom-stream's 55193 restarts likely point to a code issue, cto-aipa's 198 restarts in 0 days could be due to external factors like a 429 rate limit from an API, or a temporary resource exhaustion.

Q: How do you prioritize which process to investigate first when multiple have restarts?
A: I prioritize based on the restart count relative to uptime and the process's criticality. cto-aipa with 198 restarts in 0 days is more urgent than algom-stream with 55193 restarts over 54 days, because the former indicates a recent, acute failure.

Q: What tools do you use to diagnose the root cause of restarts?
A: I start with pm2 logs [process_name] to check application output. For deeper issues, I review system logs, check external API logs (like the HubSpot 429 error), and use git log to correlate restarts with recent code changes, such as the 3 commits to aideazz in the last 48 hours.

Q: Is there a threshold for "acceptable" restarts for a PM2 process?
A: It depends on the process. For critical, long-running services like algom-poll or n8n, 0 restarts over 50+ days is the goal. For others, 1-4 restarts over 50+ days (like dragontrade-dashboard or whitespace) might be acceptable, but 198 restarts in 0 days (like cto-aipa) or 55193 restarts over 54 days (like algom-stream) are clear indicators of problems.

— Elena Revicheva · AIdeazz · Portfolio