Skyline

Laravel Job Chains at Scale: Why the First Chain Finishes Last

· 10 min read · Boring Observability

Verified against Laravel 13.24 · Skyline 1.6

Dispatch a few hundred Bus::chain() calls at once and the first chain doesn't finish first. It finishes with all the others, close to the end of the whole run. A Laracasts thread describes it exactly: hundreds of five-job chains, each job about 20 seconds, and the first chain taking hours. The workers aren't slow. A chain queues its next job at the back of the queue, behind every other chain's first job.

This post traces that to a few lines in the framework and measures what it costs at a few hundred chains on ten workers, then covers two ways to keep each chain's steps in order without the wait. It also covers what goes wrong with chains at that scale even when the order is right, such as a ShouldBeUnique job that ends other chains without a failed job or a log line. Everything here was checked against Laravel 13.24, with probe tests on the database queue and real workers on Redis.

Key takeaways#

  • A chain lives in its job's payload. Nothing tracks it centrally. When a link succeeds, the worker queues the next link at the back of the queue, so a burst of chains runs every first step, then every second step, and so on.
  • Fixing the order doesn't change the total time. 300 chains on 10 workers take 50 minutes either way. What changes is when each chain finishes: the first one in under two minutes, instead of after 40.
  • The simplest fix is a second queue. Put every link after the first on a queue that workers check first. In Horizon this needs 'balance' => false.
  • A ShouldBeUnique link with the default uniqueId drops the rest of every other chain. Later links go through the unique lock, nothing fails, and the chains just stop.
  • Steps can run twice. The next link is queued before the current job is deleted, so a worker that dies in between leaves two copies of the rest of the chain.

What Bus::chain() actually queues#

Only the first job. PendingChain::dispatch() serializes every later job into an array and stores it on the first job's chained property, so the whole chain travels in one payload. There is no chain table and no chain ID in Redis.

When a worker finishes a job, CallQueuedHandler::call() checks that it neither failed nor was released, then calls dispatchNextJobInChain() on it:

public function dispatchNextJobInChain()
{
    if (is_array($this->chained) && ! empty($this->chained)) {
        dispatch(tap(unserialize(array_shift($this->chained)), function ($next) {
            $next->chained = $this->chained;

            $next->onConnection($next->connection ?: $this->chainConnection);
            $next->onQueue($next->queue ?: $this->chainQueue);

            $next->chainConnection = $this->chainConnection;
            $next->chainQueue = $this->chainQueue;
            $next->chainCatchCallbacks = $this->chainCatchCallbacks;
        }));
    }
}

The next link is an ordinary dispatch(). On Redis that is an RPUSH onto the tail of the list, and workers LPOP from the head. On the database driver it is a new row, and workers take the lowest id. Either way the link joins the back of the queue, behind whatever was already waiting.

Hundreds of chains run breadth-first#

Dispatch three three-step chains, a, b and c, and run one worker. The jobs table holds three rows right after the dispatch, one per chain, and this is the order the steps ran in:

a1 b1 c1 a2 b2 c2 a3 b3 c3

a2 was queued when a1 finished, behind b1 and c1. Scale that up and every chain's second step waits behind every other chain's first step. A chain's steps still run in order, since a2 can't be queued before a1 succeeds, but the chains advance together, one step at a time.

To put numbers on the Laracasts case we simulated it: 300 chains of five jobs, 20 seconds per job, 10 workers, all chains dispatched at once.

Queue layout First chain done Median chain done Last chain done
One queue, plain Bus::chain() 40 min 45 min 50 min
Later links on a queue workers check first 1.7 min 27 min 50 min

Each step of a chain waits for the 299 jobs queued ahead of it, about 10 minutes, so a five-step chain takes 50 minutes from end to end. At 1,000 chains the first one finishes after 2 hours 14 minutes, which matches what the thread describes. The last column doesn't move: 1,500 jobs of 20 seconds on 10 workers is 50 minutes of work, and no queue layout changes that. The thread's suspicion that the workers couldn't keep up was wrong. The throughput was fine, and the order was the problem.

We then ran the same thing for real on Redis, scaled down to 60 chains of five half-second jobs on 10 worker processes:

Queue layout First chain done Median Last
One queue, plain Bus::chain() 14.4 s 16.0 s 17.0 s
A continue queue checked first (fix 1) 3.8 s 11.7 s 17.0 s

The real runs sit a second or two above the simulation because each worker process boots Laravel first.

The wait isn't the only cost. While the run lasts, the queue's length stays near the number of chains, because each finished step is replaced by the chain's next one. The 1,200 jobs still to come sit inside payloads, where no queue metric counts them. Wait time on that queue is about ten minutes for every job, and no dashboard shows a chain as 60% done.

Fix 1: a continue queue that workers check first#

Workers started with --queue=high,low take a job from high whenever there is one. So put each chain's first link on one queue and everything after it on another, and list the second one first:

Bus::chain([
    (new JobA($task))->onQueue('chain-start'),
    new JobB($task),
    new JobC($task),
    new JobD($task),
    new JobE($task),
])->onQueue('chain-continue')->dispatch();
php artisan queue:work --queue=chain-continue,chain-start

A job's own onQueue() wins over the chain's, so JobA goes to chain-start and the rest inherit chain-continue as they are queued. With one worker the three test chains now run a1 a2 a3 b1 b2 b3 c1 c2 c3. With ten workers, a new chain only starts when no chain already in progress has a step waiting, so about ten chains are in progress at any time and each one runs straight through.

The thread's own suggestion was a queue per step, listed in reverse (queue_5,queue_4,…,queue_1). That favours chains closest to finishing. With steps of similar length it gave exactly the same numbers as two queues in our simulation, and two queues don't need changing when a step is added.

The catch is that it relies on workers checking the queues in strict order. In Horizon that means a supervisor with 'balance' => false. With simple or auto, each queue gets its own pool of worker processes, so nothing makes a worker prefer chain-continue. auto is worse than no preference: it moves processes towards the queue with the most work, which at the start of a burst is chain-start. The trade-offs of turning balancing off are covered in Horizon queue balancing: auto vs false.

Fix 2: cap how many chains are running#

The breadth-first order only appears because hundreds of first links are queued at once. Queue ten and there is nothing for the later links to wait behind. Keep the tasks in a table, start K chains, and end each chain with a job that starts the next one:

class StartNextTask implements ShouldQueue
{
    use Dispatchable, Queueable;

    public function handle(): void
    {
        $task = DB::transaction(function () {
            $task = Task::where('status', 'pending')
                ->orderBy('id')
                ->lockForUpdate()
                ->first();
            $task?->update(['status' => 'running']);

            return $task;
        });

        if (! $task) {
            return;
        }

        Bus::chain([
            new JobA($task),
            new JobB($task),
            new JobC($task),
            new JobD($task),
            new JobE($task),
            new StartNextTask,
        ])->catch(function () {
            StartNextTask::dispatch();
        })->dispatch();
    }
}

// Start ten chains:
foreach (range(1, 10) as $i) {
    StartNextTask::dispatch();
}

The catch() callback matters. A chain whose link fails never reaches its last job, so without it every failure leaves one fewer chain running until none are. A worker killed mid-chain loses its slot the same way, so a scheduled command that counts running tasks and tops them back up to K is worth adding.

With K equal to the worker count the simulation matches fix 1: first chain in 1.7 minutes. Raising K brings the wait back gradually. K = 30 finishes the first chain at 4.3 minutes and K = 100 at 13.7. This is the fix to use when the same workers serve other queues too, because it works with any queue driver and any balancing mode, and a burst of chains can't flood the queue for everyone else.

What doesn't change the order#

Wrapping the chains in a batch. Bus::batch() accepts arrays of jobs as chains and adds progress tracking and completion callbacks. It queues the chains the same way: a batch of three three-step chains ran a1 b1 c1 a2 b2 c2 a3 b3 c3, just like plain chains.

A worker pool per step. One queue and one Horizon supervisor per step works like an assembly line. With two workers on each of five steps it simulated almost as well as fix 1, but workers sit idle while the line fills and empties, and the split has to be tuned by hand. It's only worth it when one step needs different resources from the rest, such as a rate-limited API.

The first link of a chain is queued by the bus dispatcher directly. Every later link goes through dispatch(), which returns a PendingDispatch, and PendingDispatch is where ShouldBeUnique takes its lock. If the lock is held, the link is never queued, and neither is anything after it.

Breadth-first order makes that likely. Give the middle step of three chains ShouldBeUnique with the default uniqueId, which is an empty string, so every instance of the class shares one lock:

a1 b1 c1 a2 a3

All three chains reached step two while a2 was still queued and holding the lock, so the dispatches of b2 and c2 were skipped. Chains b and c ended there. The failed jobs table stayed empty, the queue was empty, and no exception was thrown. With hundreds of chains, only the chains whose unique step happens to find the lock free get past it.

It also fails the other way round. Because the first link skips the lock, ShouldBeUnique on a chain's first job doesn't stop duplicate chains. With the lock held, UniqueJob::dispatch() queued nothing and Bus::chain([new UniqueJob, …])->dispatch() queued the job anyway.

Both halves are reported upstream as #49263. Until that changes, give a unique chain link a uniqueId() that includes the task it works on, so chains only block duplicates of themselves. The other paths that skip the lock are in the ShouldBeUnique gotchas post. Skyline logs each skipped dispatch as a unique_lock.discarded warning with the lock key, including skips of chain links, and counts them per job on the Locks & Limits screen. A unique job with a climbing skip count and chains that never finish is this bug.

The rest of a chain can run twice#

CallQueuedHandler::call() queues the next link first and deletes the current job after. If the worker dies between the two, the current job is never deleted. It becomes available again once retry_after has passed, runs again, and queues the next link a second time. From there, two copies of the rest of the chain run side by side.

We reproduced it by throwing an exception from a JobQueued listener the first time a2 was queued, which is what a dropped connection after a successful push looks like to the worker. The job was released and retried:

a1#1 a2#1 a1#2 a3#1 a2#1 a3#1

a1 ran twice, and so did every step after it. The chain structure doesn't protect you from at-least-once delivery. A step that charges a card or sends an email needs to check whether it already happened.

Payloads carry the rest of the chain#

Every link holds every link after it, serialized. For a five-step chain where each job has a 2 KB string property, the queued payloads were:

Link 1 2 3 4 5
Payload size 13.0 KB 10.5 KB 8.0 KB 5.4 KB 2.9 KB

That is 40 KB written to the queue for 10 KB of data, and it grows with the square of the chain's length. Jobs that hold Eloquent models with SerializesModels store only the model's ID, so the overhead is small. Jobs that carry the data itself, like a $task DTO or an API response, multiply it. That adds Redis memory for every chain waiting in a burst, and it counts against SQS's message size limit.

When a link fails#

A failed link stops the chain, and the rest of the chain is kept. It is still in the failed job's payload, in chained. catch() callbacks run when a link fails for good, after its retries. php artisan queue:retry queues the failed link again with its chain, so the chain picks up where it stopped. In the probe, a chain whose second step failed ran a2 and then a3 after queue:retry all, and the retried a2 started again at attempt one.

Retrying also puts the link at the back of the queue. After an incident with hundreds of failed chains, queue:retry all recreates the burst, and the chains run breadth-first again unless one of the fixes above is in place.

Which fix to use#

If the workers only serve these chains, use the continue queue. It's a two-line change, and each step keeps its own retries and timeout. If the workers also serve other queues, cap the number of running chains, so a burst can't push other work back. Whichever you pick, give unique chain links a uniqueId() that includes the task, and make the steps safe to run twice.

Frequently asked questions

Why do my Laravel job chains take so long when I dispatch hundreds of them?

Because a chain queues its next job only when the current one succeeds, and it queues it at the back of the queue. With hundreds of chains dispatched at once, every chain's second step waits behind every other chain's first step, so the chains advance together and all finish near the end. 300 five-job chains of 20-second jobs on 10 workers finish their first chain after about 40 minutes. The total run time of 50 minutes is the same however the queue is arranged.

How do I make Laravel run a chain's jobs back to back?

Put every job after the first on a second queue and start workers with that queue listed first, for example queue:work --queue=chain-continue,chain-start. Set the first job's queue with onQueue() and the rest with onQueue() on the chain itself. In Horizon this needs a supervisor with balance set to false. The other option is capping how many chains run at once, so there is never a burst of first jobs for later ones to wait behind.

Can I use ShouldBeUnique on a job in a chain?

Only with a uniqueId() that identifies the chain's task. Every link after the first is queued through PendingDispatch, which takes the unique lock, and when the lock is held the link is skipped and the rest of that chain never runs. Nothing fails and nothing is logged by Laravel. The first link is queued without taking the lock, so ShouldBeUnique on it does not stop duplicate chains either.

What happens to the rest of a Laravel chain when a job fails?

The chain stops and the remaining jobs are kept in the failed job's payload. catch() callbacks run once the link has used up its retries. Running php artisan queue:retry on the failed job queues it again with the rest of its chain, so the chain continues from the job that failed.

Does Bus::batch() keep chained jobs in order?

Each chain inside a batch still runs in order, but the batch queues the chains the same way plain chains are queued, so many chains in one batch run breadth-first too. A batch adds progress tracking and completion callbacks, and leaves the order as it was.