Skyline

Laravel Background Jobs: 12 Best Practices for Production Queues

· Updated · 11 min read · Boring Observability

Verified against Laravel 13 · Horizon 5.x

Queued jobs run after the request is gone: during a deploy, against data that may have changed or been deleted, and on code one version newer than the code that dispatched them. Most job bugs come from assumptions about when a job runs and how many times it runs. Those assumptions hold in development and fail under production load.

The twelve practices below build on each other. Small arguments make idempotency easier, idempotency makes retries safe, and safe retries are what make backoff useful. The last one, backwards compatibility across deploys, is the one no linter catches.

Key takeaways#

  • Assume every job runs at least twice. Laravel's queues are at-least-once, and a worker killed mid-job re-runs it from scratch, with the side effects of the first attempt still committed.
  • Keep jobs small, atomic and explicitly queued. One job per record beats one job that loops over 10,000: a failure then retries one unit of work instead of everything.
  • Let the framework handle failure. Declarative $tries and exponential backoff() rather than manual re-runs, and afterCommit() so a worker can't outrun your transaction.
  • Read the whole contract before using ShouldBeUnique or batches. Both fail without raising an error. A missing uniqueId() drops jobs and nothing reports it.
  • The queue outlives your deploy. Jobs serialized against old code are deserialized by new workers, so renaming a job class or adding a constructor argument without a class-level default destroys work that is already in flight.

1. Keep arguments small#

Every byte of constructor arguments is a byte in Redis, multiplied by every queued instance. Don't pass big arrays, blobs, or rich DTOs "to save a query." Pass the minimal identifiers and re-read what you need inside handle(). A large backlog of fat jobs can run Redis out of memory.

If you pass Eloquent models, the SerializesModels trait makes Laravel store only the model's class and ID, and load the record again when the job runs.

2. Every job should declare its queue#

class GenerateImageThumbnails implements ShouldQueue
{
    public $queue = 'media';
}

Horizon routes work by queue name, and each supervisor is tuned for a workload (concurrency, timeout, balance strategy). A job with no $queue lands on default, where a slow image resize can sit ahead of a latency-sensitive job, or behind one. Putting fast and slow jobs, or urgent and bulk jobs, on one queue removes the reason for having separate supervisors.

Pick the queue that matches the work's shape and SLA, and make sure a matching Horizon supervisor processes it. How that supervisor then divides its workers between those queues is its own trade-off, covered in Horizon queue balancing: idle workers vs. starved queues.

3. Make jobs idempotent#

Horizon delivers at-least-once. A worker killed mid-job by a deploy, an OOM or a timeout leaves the job to be retried, so the same job body can run twice. Design every job so that running it twice produces the same end state as running it once:

  • Check before you create (firstOrCreate, updateOrCreate, "already sent?" guards).
  • Add unique constraints in the database as a backstop for those checks.
  • Don't assume that a job which ran has run exactly once.

The common failure is a job that did partial work and was killed before it finished. Redis only removes a job from its reserved set on success, so a hard kill (SIGKILL from a per-job timeout, or the supervisor force-killing a worker that outlived the shutdown grace window during a deploy) leaves the job to be migrated back and re-run from scratch, with whatever side effects the first attempt already committed still in place.

Because attempts is incremented when the job is reserved, not when it fails, each kill-and-recover cycle burns one try. After maxTries such cycles the job is marked failed, having run its side effects several times without once completing. Idempotency therefore has to cover resuming after a partial run, not only avoiding a duplicate of a run that succeeded.

4. Keep jobs atomic and short#

A deploy restarts workers, and a job that exceeds its $timeout is killed mid-flight. If your job does five writes and dies after three, the retry redoes all five, so the first three have to be idempotent (see above).

Prefer one job per record over one job that loops over 10,000 records. A per-record job that dies retries one record; a mega-job retries everything and may never finish inside the timeout window. Fan out:

Item::where(...)->eachById(fn (Item $i) => SyncItemJob::dispatch($i));

5. Let the framework handle retries and backoff#

Don't manually re-run failed jobs, and don't catch and swallow exceptions. Configure retries declaratively:

public $tries = 5;

// exponential backoff: required for anything that calls an external API
public function backoff(): array
{
    return [10, 30, 60, 300];
}

Immediate retries against an API that is having trouble turn a brief failure into an outage. Jobs that cannot succeed should land in failed_jobs, where they point you at a root cause to fix; don't mass-retry them. Use $this->fail($e) to fail early when you know a retry won't help. (How attempts, timeouts and release-based middleware interact is covered in what breaks under real traffic.)

6. Dispatch after the transaction commits#

DB::transaction(function () {
    $booking = Booking::create([...]);
    SendBookingConfirmation::dispatch($booking->id)->afterCommit();
});

Without afterCommit() (or the connection-level 'after_commit' => true), a fast worker can pick up the job and find() a booking that hasn't been committed yet. The result is a ModelNotFoundException that only reproduces under load. Dispatch the side effect once the data it depends on is durable. Nested transactions, deadlock retries, unique locks and Queue::fake() all change when that push happens; see dispatching Laravel jobs after commit.

7. The ShouldBeUnique contract: uniqueFor and uniqueId#

Mistakes with unique jobs drop work without raising an error, which makes them expensive to find. How uniqueness interacts with WithoutOverlapping, retries and orphaned locks is covered in ShouldBeUnique vs WithoutOverlapping in Laravel.

Set uniqueFor. A plain ShouldBeUnique lock is acquired at dispatch and released only when processing reaches a terminal state (success, or final failure after maxTries), so it is held while the job processes and across backoff retries. Without uniqueFor the lock is created with no TTL (forever()), and only that eventual terminal run reclaims it. Normally that's fine: a worker killed mid-job leaves the job to be retried, and the retry's terminal state clears the lock. That leaves two exposures:

  • Until that retry completes, every new dispatch of the same unique job is dropped without an error. uniqueFor bounds this window to a known TTL instead of however long the kill, retry and finish cycle takes.
  • If the job never reaches a terminal run, the forever() lock is never reclaimed and the job cannot be dispatched until someone clears the key by hand. That happens with a uniqueId() that reads external mutable state (so the release computes a different key than the acquire), a lost reserved entry, or maxTries: 0 with a job that is killed on every attempt. uniqueFor is the only thing that recovers it automatically.

Caveat: uniqueFor must exceed your worst-case total processing-plus-retry time, or the lock expires during a legitimate run and a duplicate can be dispatched. It caps how long a stuck lock can last. It does not replace keeping uniqueId() a pure function of the job's own data.

public int $uniqueFor = 3600; // lock expires after an hour even if it is never released

uniqueId is mandatory for parameterized jobs. The lock key is laravel_unique_job:<class>:<uniqueId>. Omit uniqueId on a job that takes arguments and the key collapses to the class name alone, so SyncCompany(1) and SyncCompany(2) share one lock and one of them is dropped at dispatch without an error.

public function uniqueId(): string
{
    return (string) $this->companyId;
}

If the job is meant to be a class-wide singleton, return a constant from uniqueId() so the intent is explicit.

8. Never bulk or batch a unique job#

Queue::bulk() / Bus::bulk() push raw payloads straight to Redis, skipping the dispatcher that acquires the lock, so the uniqueness check never runs and nothing tells you. Batching a unique job fails the same way, because Batch::add() pushes its jobs through the queue's bulk() method as well, so every duplicate in the batch is queued and runs. The first job of a Bus::chain() and a queue:retry skip the lock too. Dispatch unique jobs individually, or remove duplicates from the list before you add it to a batch.

Skyline takes a free unique lock on those pushes as well, and logs a warning when one queues a copy while another job holds the lock.

Periodic and scheduled jobs should implement ShouldBeUnique so a slow run doesn't overlap the next tick.

9. A batchable job must honour cancellation#

Cancelling a batch only stops future dispatches. Jobs already on the queue still run their full body unless they check. For anything that mutates state, such as writing files or charging cards, that is wasted work for a batch the caller abandoned. Guard it:

public function handle(): void
{
    if ($this->batch()?->cancelled()) {
        return;
    }
    // ... heavy work
}

You can also centralise the check with the Illuminate\Queue\Middleware\SkipIfBatchCancelled middleware.

10. Don't sleep() in a job#

A sleeping job holds a worker that does nothing while every job behind it waits. If you need to wait for a rate limit or an upstream that isn't ready, release the job back to the queue with a delay:

$this->release(60); // back on the queue, worker freed, retried in a minute

11. Keep jobs backwards compatible across deploys#

When you deploy, the queue already holds jobs serialized against the old code, and workers running the new code have to deserialize and run them. There are two ways this goes wrong.

Renaming or moving a job class#

The serialized payload stores the fully-qualified class name, for example App\Products\Jobs\SyncStockJob. Rename the class, move it to another namespace, or delete it, and every instance already in the queue becomes unresolvable: deserialization throws, the job lands in failed_jobs, and the work is lost. To avoid it:

  • Two-phase it. Keep the old class (even as a thin subclass of the new one) for one deploy cycle, let the queue drain, then remove it in a follow-up deploy.
  • Or drain the queue of that job type before shipping the rename.
  • Don't rename a hot job class in the same PR that changes its behaviour.

Changing constructor arguments#

PHP's unserialize() does not call the constructor. It restores the saved properties directly, so your constructor's default values have no effect on jobs already in the queue.

class SyncStockJob implements ShouldQueue
{
    // Promoted, with no class-level default.
    // An old payload serialized before $force existed restores without it,
    // so accessing $this->force throws "must not be accessed before initialization".
    public function __construct(
        public int $productId,
        public bool $force = false,
    ) {}
}

The constructor default = false only applies to new dispatches. An old queued job never goes through the constructor, so its $force stays uninitialized and the new handle() that reads it crashes. Give the property a class-level default instead, and deserialized old jobs fall back to it:

class SyncStockJob implements ShouldQueue
{
    public bool $force = false; // class-level default, set even though unserialize skips the constructor

    public function __construct(public int $productId, bool $force = false)
    {
        $this->force = $force;
    }
}

Which job changes are safe to deploy?#

For any change, ask what happens to a payload that was serialized an hour ago by the code you are about to replace.

Change Safe in one deploy? What happens / what to do
Edit handle() body Yes Behaviour isn't serialized, only properties are. Make sure it tolerates the old payload shape.
Add a property with a class-level default Yes Old payloads restore without the key and fall back to the default.
Add a promoted constructor argument No unserialize() never calls the constructor, so the property stays uninitialized and reading it throws. Give it a class-level default instead.
Remove or rename a property No In-flight payloads still carry the old key. Deprecate over one deploy cycle, then remove.
Rename / move / delete the job class No The FQCN is stored in the payload; queued jobs become unresolvable and land in failed_jobs. Two-phase it, or drain first.
Change the queue name Yes, with care Jobs already queued stay on the old queue. Keep a supervisor processing it until it drains.

The rules that follow from the table:

  • New properties get a class-level default value (or are nullable), never a bare typed property.
  • Don't remove or rename a property a queued payload still carries without a deprecation window.
  • handle() must tolerate both the old and the new payload shape for the length of one deploy cycle.
  • When in doubt, two-phase: ship the backwards-compatible change, let the old jobs drain, clean up in a later deploy.

12. Monitor every queue#

Use Horizon. Watch wait times, failure rates and throughput for each queue, because a backlog on one supervisor doesn't show up in the others. A queue with no failing jobs can still be 20 minutes behind.

When a queue is 20 minutes behind at 2am, Skyline lets you act on it from the dashboard: pause the queue feeding the backlog, trigger a delayed job immediately, or drain it, without shelling into a worker box. It replaces Horizon rather than running alongside it. How Skyline compares to the other queue tools covers where it sits next to an APM or a job recorder.

Frequently asked questions

How many times will a Laravel queued job run?

At least once, and sometimes more. A worker only removes a job from Redis once it succeeds, so a job killed mid-flight (a deploy, an OOM, or a $timeout) is migrated back and retried from scratch, with any side effects the first attempt already committed still in place. Design every job so that running it twice produces the same end state as running it once.

Why does my queued job throw ModelNotFoundException?

Usually because you dispatched it inside a database transaction that had not committed yet. A worker can pick the job up before the commit lands, so find() sees no row. Dispatch with ->afterCommit(), or set 'after_commit' => true on the queue connection.

Is it safe to rename a queued job class?

Not in a single deploy. The serialized payload stores the fully-qualified class name, so every job already sitting in the queue becomes unresolvable the moment the class moves and lands in failed_jobs. Two-phase it: keep the old class for one deploy cycle, let the queue drain, then remove it.

Why does adding a constructor argument break jobs already in the queue?

PHP's unserialize() restores saved properties directly and never calls the constructor, so a constructor default does nothing for a job that was queued before the property existed. Reading the uninitialized property throws. Give new properties a class-level default instead.