When Batch Stops
Diagnosing an environment where nothing is erroring and nothing is moving.
PDF31 pagesConcurrent processing
Email me the guide- The queue is full, nothing is erroring, and nothing is moving.
- One manager shows zero workers against a configured maximum — and it says Running.
- Requests sit in RUNNING for hours; the OS process behind them died yesterday.
- Everything got slower over months and nobody can point to the day it started.
When Batch Stops: The EBS Triage Guide
31 pagesPart One — When batch stops
- The Four Failure Modes
- The Work Shift Gap
- Stuck Requests
- Reading a Health Diagnostic
Part Two — When batch runs slowly
- Queue Capacity and Manager Design
- The Database Underneath
- Housekeeping
Part Three — When interfaces fail
- Why Integrations Break
- The Four-Level Triage Framework
- Building Resilience
Appendices
- The Batch Triage Checklist
- Monitoring Reference
- Table Reference
- Tell the four failure modes apart in minutes — ICM down, work shift gap, stuck requests, node loss — and stop diagnosing the wrong one.
- Prove a RUNNING request is actually dead, and terminate it through the application without collateral damage.
- Read a concurrent processing health diagnostic and act on it in the order that matters.
- Design manager queues, target processes and work shifts for close-week load instead of average load.
- Classify an interface failure into the category it belongs to before touching anything.
- Put monitoring thresholds on the things that fail silently, at the cadence that catches them.
The Four Failure Modes
“Nothing is processing” has four distinct causes with four different remedies. Identifying which one you have is the entire first move, and it takes minutes.
| Mode | Signature | Scope and remedy |
|---|---|---|
| ICM down | The Internal Concurrent Manager is not running on any node. Every manager is affected simultaneously. | Instance-wide. Check this first — diagnosing individual managers while the ICM is down wastes the first hour. |
| Work shift gap | One manager shows zero active workers against a non-zero configured maximum. Others are fine. The manager appears Running. | One manager. Extend or add the work shift covering the current time. |
| Stuck requests | Requests in RUNNING status whose OS process no longer exists. Worker slots are held permanently; capacity degrades as more accumulate. | One manager, progressive. Terminate through the application. |
| Node unavailable | Active workers below the configured maximum, with no work shift problem. One node of a parallel manager cluster is down. | Proportional capacity loss. Restart the node or redistribute workers. |
From Chapter One.
William A. Green has spent 27 years inside Oracle E-Business Suite Financials — as a former Oracle employee, as a functional and technical lead on more than 35 implementations, and for six years at Rimini Street supporting Fortune 500 EBS clients on releases Oracle had stopped patching. He writes fixes for a living. This guide is the part of that work that fits in a PDF.