10. Executions and Monitoring¶
Building an automation is easier than noticing that it quietly stopped. This page is about that part.
Dashboard — the overall picture¶
The first screen you see after signing in.
| Metric | Meaning |
|---|---|
| Total Workflows | How many workflows you have built |
| Total Executions | How many times they have run in total |
| Success Rate | The success rate |
| Active Executions | How many executions are running right now |
Recent Executions below shows the latest runs. If the success rate suddenly drops, start here.
Reading the execution history¶
Workflow detail → the bottom of the Overview tab holds that workflow's execution history.
Click a single execution and you can see, per node:
- Each node's status (succeeded / failed / retrying)
- The input that node received
- The output that node produced
- The error message, if it failed
90% of debugging ends right here. Look at the output of the node just before the one that failed and you will usually see immediately whether the expression path was wrong or the value was empty.
Reading the statuses¶
The execution as a whole
| Status | Meaning |
|---|---|
RUNNING |
In progress |
COMPLETED |
Succeeded all the way through |
FAILED |
Failed partway |
CANCELLED |
Cancelled |
Each individual node
| Status | Meaning |
|---|---|
PENDING |
Waiting its turn |
RUNNING |
Running |
COMPLETED |
Succeeded |
FAILED |
Failed |
RETRYING |
Retrying |
CANCELLED |
Cleaned up because the execution ended in failure |
SKIPPED |
Not chosen by a conditional branch, so it never ran |
CANCELLED is not a problem with that node itself. It means the node was still running when another node
failed and ended the execution, so it got cleaned up. The real cause is in the FAILED node.
SKIPPED is normal. The CONDITIONAL picked a different branch, so this path was simply not used.
JOINT and LOOP_END treat SKIPPED as "finished", so the run does not stall there.
→ 06. Flow Control
Retries and timeouts¶
External APIs fail sometimes. There are two mechanisms so that a temporary failure does not stop the whole automation.
- id: fetch-orders
name: Fetch orders
type: CALL
integration: http_request
timeout: 10s
retry-policy:
max-attempts: 3
backoff:
type: EXPONENTIAL
initial-delay: 500ms
multiplier: 2.0
max-delay: 10s
retry-on: ["429", "5xx"]
input:
...
timeout — how long to wait¶
Past this duration the node counts as failed. The formats you can use are 500ms · 10s · 5m (a number plus ms/s/m).
| Kind of node | Recommended |
|---|---|
| Fast API lookups | 5s – 10s |
| Heavy lookups and report APIs | 30s – 1m |
AI calls (llm_chat) |
30s – 2m |
If you set no timeout, the execution can hang for a long time waiting on an API that never answers. It is worth setting one on every external call.
retry-policy — how many times to try again¶
| Field | Required | Description |
|---|---|---|
max-attempts |
✅ | Maximum attempts (including the first one). 3 means 1 initial attempt + 2 retries |
backoff.type |
✅ | FIXED (the same gap every time) or EXPONENTIAL (progressively longer) |
backoff.initial-delay |
✅ | How long to wait before the first retry |
backoff.multiplier |
The factor when using EXPONENTIAL. Defaults to 2.0 |
|
backoff.max-delay |
An upper bound on the wait | |
retry-on |
Which failures to retry on. Leave it empty for every retryable failure |
With EXPONENTIAL, initial-delay: 500ms and multiplier: 2.0, the waits grow 500ms → 1s → 2s → 4s.
⚠️ retry-on filters two kinds separately¶
The values retry-on accepts fall into two groups, and neither group touches the other.
| Axis | Accepted values | Meaning |
|---|---|---|
| HTTP status | "429" · "503" (exact) · "5xx" · "4xx" (range) |
A response arrived, but the status says failure |
| Failure kind | "timeout" · "connect-error" |
No response at all (there is no status code) |
If you list nothing on an axis, everything on that axis is retried. This is the part that trips people up.
retry-on: ["5xx"]
# → statuses are narrowed to 5xx. But nothing is listed on the failure-kind axis,
# so timeouts and connection errors are STILL retried, all of them.
retry-on: ["5xx", "connect-error"]
# → 5xx only on status, connection errors only on kind. Timeouts are not retried.
retry-on: ["timeout"]
# → NOT "retry timeouts only". Nothing is listed on the status axis,
# so every status failure including 429 and 5xx is retried.
retry-on: ["TIMEOUT"]
# ❌ uppercase is not a value. The save is rejected (it is lowercase "timeout")
If you really need to narrow retries on a sending node, list both axes.
When not to retry¶
Do not put retries on anything that must not send the same request twice.
| Node | Retry |
|---|---|
| Reads (GET) | ✅ Safe |
| Data transforms | ✅ Safe |
| Sending a message | ⚠️ It may go out twice |
| Payments and order creation | ❌ Dangerous |
For message sending, it is safer to list both axes so the retries are genuinely narrow.
retry-on: ["429", "5xx", "connect-error"]
# Retries only when the connection never got through (so nothing was sent),
# and not on a timeout (where the message may well have arrived).
Turning a workflow on and off¶
A way to pause without deleting. Use it when a daily schedule keeps firing bad alerts but you want to keep the workflow around.
Where the switch is¶
Toggle it from the workflow list and detail screens. It reads On when running and Off when paused.
What is blocked while it is off¶
| Entry point | While off |
|---|---|
| Schedule (cron) firing | ❌ does not fire |
| Incoming webhook | ❌ rejected with 409 |
| AI assistant (MCP) run | ❌ rejected |
| The Run button on screen | ✅ goes through |
Manual runs stay open so that turn off → fix → test → turn back on actually works. If turning it back on to verify a fix also re-armed the cron, that would defeat the purpose.
Worth knowing¶
- Only new runs are blocked. Runs already under way continue to completion. To stop those, cancel them from the execution history.
- Turning it off does not lift the save lock. A running workflow still cannot be saved.
- Saving the YAML again does not flip it back. On/off is operational state, not part of the
workflow definition, so a round trip through the
YAMLtab leaves it alone.
⚠️ Leaving a GitHub webhook off for a long time¶
GitHub disables a webhook automatically once delivery failures pile up. While your workflow is off,
eeumsae answers 409, which counts as a failure on GitHub's side.
If it was off for a while, check in the repository settings that the GitHub-side webhook is still alive after you turn it back on.
Set up failure alerts — please do this¶
If an automation that runs every day quietly stops, nobody finds out. This is the feature that prevents it.
Go to Notifications in the sidebar to configure it.
| Field | Description |
|---|---|
| Webhook URL | Where failure alerts go. It has to start with http:// or https:// |
| Enable notifications | Turn this off and nothing is sent even on failure |
When a workflow execution ends in failure, one POST goes to this address.
Creating a Slack incoming webhook¶
- Create an app, or open an existing one, at api.slack.com/apps.
- Turn on Incoming Webhooks.
- Use Add New Webhook to Workspace to pick the channel that receives alerts.
- Paste the resulting
https://hooks.slack.com/services/...address into the Notifications screen.
Discord works the same way — channel settings → Integrations → Webhooks.
If a URL is already configured, the screen shows it masked. Saving with the field left blank keeps the existing URL and only toggles it on or off.
This is a different feature from sending to Slack inside a workflow. What you configure here is a system alert saying "the automation failed"; a
slack_post_messagenode is a business message sent because the automation worked.
Editing a workflow while executions are running¶
When an execution starts, the shape of the workflow at that moment is saved, and that execution follows that shape to the end.
- If you edit a workflow while a long execution is in flight, the execution already running is unaffected.
- Your edits apply from the next execution onward.
- When you edit a workflow over MCP, the edit is rejected if an execution is in progress. Wait for it to finish, or cancel it, and try again.
Pre-flight checklist¶
Worth running through once whenever you turn on a new automation.
- [ ] Did you set a
timeouton the external call nodes? - [ ] Did you set a
retry-policyon the read nodes? - [ ] Did you avoid putting unnecessary retries on sending and payment nodes?
- [ ] Did you configure the failure alert webhook under Notifications?
- [ ] Did you set an explicit
timezoneon the schedule trigger? - [ ] Did you pass API keys via
${secrets.*}instead of writing them into YAML? - [ ] Did you attach
| raweverywhere you pass an array? - [ ] Did you actually run it once and check the result in the execution history?
- [ ] If the webhook is publicly reachable, did you turn on signature verification with
WEBHOOK_SECRET?