Hermes Agent Cron Jobs: A Production Scheduling and Recovery Checklist

HostAgentics Team · Published 2026-10-03 · Updated 2026-10-03

hermes-agentschedulingoperations

Operating automated agent tasks in production requires treating autonomous execution as an operational discipline rather than an unmonitored script. This runbook covers diagnosing schedules and safely retrying recurring tasks on Hermes Agent. For framework basics, start with what is Hermes Agent; this guide focuses on one job's path from schedule to verified output. All upstream documentation was checked October 3, 2026; upstream capabilities and options differ by installed version.

Scheduler Architecture and Execution Foundations

Automating recurring jobs in Hermes Agent relies on specific runtime components documented in the Hermes Cron Feature Guide and the Hermes Cron Internals Guide. Operators must account for several structural realities before scheduling workloads:

  • Native Gateway Requirement: In a standard gateway deployment, unattended automatic firing requires a running gateway scheduler. Other upstream deployments, such as the Desktop backend, have their own ticker; check the scheduler serving your profile. A separate CLI chat session is not a daemon and does not trigger scheduled jobs.
  • Execution Lifecycle: Upstream cron features provide native commands to create, edit, pause, resume, run, and remove jobs. Each recurring execution starts inside a fresh agent session.
  • Tool Availability and Boundaries: A fresh agent session still has normal tools present. Attached skills formatted as SKILL.md provide workflow instructions, but as specified in the Hermes Skills Guide, a skill is not a permission boundary and does not provide an isolation guarantee.
  • Storage and Locking: Job definitions, metadata, and output are stored in the Hermes home directory. Scheduler tick locking avoids duplicate batch dispatch across scheduler ticks; however, this is not proof of exactly-once external business effects.
  • Creation Options: Current documentation supports creation flags including --paused, --paused-reason, --name, and --deliver local.

CLI Commands and Safe Job Creation

Operators must verify installed syntax against the Hermes CLI Reference before executing commands. Always inspect your environment first:

  • Run hermes --version to check the active version.
  • Run hermes cron --help to verify supported flags, options, and commands.
  • Note that diagnostic tools such as doctor may vary by installed version.

The CLI provides read-only inspection and lifecycle management:

  • hermes cron list provides read-only job listing.
  • hermes cron status displays scheduler status for the selected profile.
  • hermes cron pause <job-id> preserves the job definition while stopping automatic scheduled triggers.
  • hermes cron resume <job-id> reactivates scheduling for a paused job.
  • hermes cron run <job-id> requests execution on the next scheduler tick; it is not a dry-run or a promise of immediate execution.
  • Pausing preserves the definition, whereas removing deletes the job entirely.

CLI job creation accepts a positional schedule expression followed by the prompt. When validating jobs on an owned test profile, operators can run:

hermes cron create "every 1h" "Read the approved status endpoint and summarize errors; do not mutate anything" --name "status-review" --deliver local --paused

Submitting this command stores the configuration; creation does not fire the job. The --paused flag must exist in your installed hermes cron --help output. If --paused is absent in your version, create only a non-production test job according to your installed documentation.

Recommended Operational Design Practices

Production reliability requires defensive practices around scheduled tasks. The following items are explicitly operational recommendations rather than upstream guarantees or observed platform behaviors:

  • Read-Only Initial Scope: Select read-only tasks when establishing automation. Explicitly approve data sources and output recipients, and keep secrets out of prompts and logs.
  • Dedicated Credentials and Isolation: Implement dedicated credentials and host OS or container runtime isolation, rather than relying merely on skill instructions or a fresh session.
  • Spend Governance: Manage model inference and API spend separately from host infrastructure. Define explicit alert thresholds; avoid assuming unverified token-limiting flags.
  • Cadence and Execution Duration: Establish scheduling cadence and concurrency based on measured execution runtimes. Provide scheduling margin to avoid job overlap, and avoid stacking operating system cron on top of the native gateway scheduler.
  • Stage-Wise Verification: Distinguish dispatch initiation, task execution success, and delivery success. Log the job ID, scheduled timestamp with timezone, attempt number, status, and output storage location.
  • Downstream Idempotency: Implement an idempotent business key matching the job and scheduled time slot in downstream targets. Inspect actual external state before retrying, and require human review for write operations.
  • State Backups and Recovery Testing: Back up configuration, job definitions, and necessary state from the Hermes home directory. Test recovery by restoring into a disconnected test environment rather than reproducing execution histories or tokens into prompts.

Diagnostic and Failure Triage Checklist

When diagnosing anomalies, evaluate components independently as outlined in the Hermes Cron Troubleshooting Guide. Collect primary evidence first; do not restart the gateway as an initial response, as restarting can disrupt in-flight work. Pausing a job halts future scheduled triggers but does not cancel active running work or manual triggers. Never delete a job definition or output files as an initial troubleshooting step.

Systematic triage sequence:

  • Scheduler Absent or Overdue: Check hermes cron status. Verify whether the native gateway scheduler is running, check the active user profile, and verify whether the definition is one-shot or recurring before modifying configuration.
  • Clock and Timezone Mismatch: If a job fires at unexpected hours, audit the local host clock, timezone settings, and schedule string independently from agent code.
  • Run Failure: If an execution fails, inspect credentials, error logs, and primary output files in the Hermes home directory.
  • Finished Without Delivery: If a task completes successfully in logs but sends no message, inspect delivery configuration and review output files locally.
  • Duplicate Results: If downstream targets observe duplicate entries, verify downstream idempotency keys and check tick logs before assuming a scheduler fault.

Safe Recovery and Retry Procedure

When recovering from a failed or overdue recurring run, apply this sequential checklist:

  1. Roll back the job by executing hermes cron pause <job-id> to prevent further scheduled executions while investigating.
  2. Record primary diagnostic evidence, including gateway logs, output files, and downstream system logs before making changes.
  3. Check the downstream system to verify whether partial business actions occurred during the failed attempt.
  4. Correct root causes in credentials, endpoints, or prompts within an isolated test profile.
  5. If retrying is safe, request a manual verification run using hermes cron run <job-id> while monitoring logs. Check the job state afterward: current upstream operator overrides can re-enable a paused job. Pause it again if automatic scheduling must remain stopped.
  6. Before using hermes cron resume <job-id>, review missed slots and downstream idempotency. Current upstream behavior can dispatch a catch-up run for a recurring slot that became due while paused; resuming is not necessarily a wait until the next future slot. Resume only when task execution, output, delivery, and possible catch-up effects have been checked.

Where to go next

For capacity limits, use the resource-governance guide. For incoming event safety, use webhook security. Before choosing hosting, compare the documented Hermes plan and pricing; this runbook is not a hosted-feature or reliability benchmark. If moving a job to another environment, start with the migration checklist and keep the new scheduler paused until the restore has been verified.

Material limitations

  • • Commands depend on the installed Hermes version; check local help before scheduling or triggering a job.
  • • The operational checklist is a recommendation, not an observed reliability or cost benchmark.
  • • A fresh session or skill instruction is not a permissions boundary; downstream effects still need idempotency and review.
  • • Publishing this checklist does not create or execute a production cron job.
Hermes Agent Cron Jobs: A Production Scheduling and Recovery Checklist · HostAgentics