Background Jobs
Running work outside the normal request/response cycle, so a slow task doesn't force a user to sit and wait for it to finish.
What is it?
Some work a backend needs to do simply doesn't fit neatly inside the short window of a single request. Sending a confirmation email, resizing an uploaded image, generating a large report, or running a nightly cleanup of old records can take seconds, minutes, or longer — far longer than a user should reasonably wait staring at a spinner for a response.
Background jobs are units of work that run outside the normal request/response cycle: instead of doing the slow work immediately and making the user wait for it, the server quickly acknowledges the request and hands the actual work off to run separately — either right away in the background, on a queue processed by separate worker processes, or later on a schedule (like "every night at 2 AM").
Explain like I'm 10
A restaurant doesn't make you stand at the counter until your food is ready — it takes your order, gives you a buzzer, and lets you sit down while the kitchen (a background worker) prepares the meal separately. You get an immediate acknowledgment ('order received') without blocking on the actual, slower work.
Examples
Handing off slow work to a queue
app.post("/signup", async (req, res) => {
const user = await createUser(req.body);
// Instead of sending the email right here and making the user wait...
await emailQueue.add("welcome-email", { userId: user.id });
res.status(201).json({ ok: true }); // responds immediately
});
// Elsewhere: a separate worker process consumes the queue
emailQueue.process("welcome-email", async (job) => {
const user = await db.users.findById(job.data.userId);
await sendEmail(user.email, "Welcome!");
});The request handler stays fast because it only enqueues the job — the actual, slower email-sending work happens separately, in a worker process, without the user ever waiting on it.
A scheduled background job
const cron = require("node-cron");
// Runs automatically every day at 2:00 AM, with no request involved at all
cron.schedule("0 2 * * *", async () => {
await deleteExpiredSessions();
logger.info("cleanup_complete");
});This job isn't triggered by any user request at all — it runs on a fixed schedule, entirely independent of the request/response cycle.
How it works
Rather than doing slow work inline, a request handler records that the work needs to happen — often by pushing a small message describing the job onto a queue (backed by something like Redis) — and immediately returns a response. One or more separate worker processes continuously watch that queue, pick up jobs as they arrive, and do the actual work, independently of any specific web request. Scheduled jobs work similarly but are triggered by a timer (a cron schedule) rather than an event, running on their own regardless of whether any request happens at all.
Why does it exist?
If every request had to fully complete every piece of related work before responding, slow operations would make the whole app feel unresponsive, and a spike in slow work (like a burst of signups all needing welcome emails) could overwhelm the web server itself. Background jobs exist to decouple "acknowledge the request quickly" from "actually get the slow work done," and to let that work be scaled, retried, and monitored independently of the web servers handling live traffic.
When to use it
Use background jobs for anything slow, non-essential to the immediate response, or safely retryable: sending emails, processing images, generating reports, syncing with third-party services, or any scheduled, recurring maintenance task.
When not to use it
If the user genuinely needs the result of the work before you can respond meaningfully (like a login endpoint that must confirm the password matches before saying 'success'), that work belongs inline in the request, not deferred to a background job.
Common mistakes
Putting genuinely time-sensitive work (like checking a password) into a background job, when the request truly can't respond correctly without its result.
Not handling job failures — a queued job that silently fails with no retry or alert can quietly drop real work (like a never-sent email).
Assuming a background job that succeeded once will always succeed, without planning for retries when a dependency (like an email provider) is temporarily down.
Practice exercises
- Easy:
Identify which of the following belongs in a background job and which belongs inline in a request: validating a login password, sending a password-reset email, resizing a profile photo.
- Medium:
Sketch the shape of a signup endpoint that enqueues a welcome email job instead of sending the email inline, and explain why that keeps the endpoint fast.
- Hard:
Describe what should happen if a background job that sends a confirmation email fails partway through, and how you'd make sure the email eventually still gets sent.
Interview questions
What is a background job and why would you use one?
Work performed outside the normal request/response cycle — used for slow or non-essential tasks so the user isn't forced to wait for them before getting a response.
What's the difference between a queued job and a scheduled job?
A queued job is triggered by an event (like a user signing up) and processed by a worker as soon as it can; a scheduled job runs automatically on a fixed timer, independent of any specific request.
Why shouldn't login password verification be handled as a background job?
Because the request genuinely needs the result immediately to decide how to respond — deferring it would mean the server couldn't tell the user whether login succeeded.
Why does a job queue typically need a separate worker process rather than running jobs inside the same process that handles web requests?
Running slow job work inside the same process as request handling would compete with incoming HTTP requests for the same CPU and memory, undermining the point of offloading it; a separate worker process can be scaled, deployed, and monitored independently of the web-facing servers.
What happens to a queued job if the worker processing it crashes partway through?
A well-configured queue only marks a job done after the worker explicitly confirms success, so a crash mid-processing typically leaves the job unacknowledged and it becomes available again for another worker to pick up and retry, rather than being silently lost.
Why does a background job need to tolerate being run more than once, i.e. be idempotent, given retry behavior?
A queue can redeliver a job that appears to have failed, like when a worker crashes right after finishing but before acknowledging, so the same job may run twice; an idempotent job checks whether its effect already happened before repeating it, so it produces the same correct result either way instead of double-charging a card or sending a duplicate email.
What's a dead-letter queue, and why is it useful?
A separate destination a job gets moved to after failing repeatedly beyond some retry limit, instead of being retried forever or silently dropped, letting a team inspect and manually handle jobs that are genuinely stuck without those retries endlessly consuming worker capacity.
Why use exponential backoff between retries of a failed background job instead of retrying immediately and repeatedly?
If the underlying dependency is temporarily overloaded or down, retrying instantly and repeatedly adds more load right when it's least able to handle it; spacing retries out with increasing delay gives the dependency time to recover and reduces the odds the retries themselves worsen the outage.
What's the difference between a job queue backed by Redis, like BullMQ, and one backed by a relational database table?
A Redis-backed queue is typically faster and purpose-built with features like delayed jobs, priorities, and retries out of the box; a database-table-backed queue needs less new infrastructure if a relational database is already in use, at the cost of usually being slower and needing more of that queue logic implemented by hand.
Why is it important to monitor background job failure rates separately from the web server's own error rate?
A background job failing doesn't produce an HTTP error response anyone notices in the moment, since the user who triggered it already got their success response — without separate monitoring of the job system itself, a systematically failing job could go unnoticed indefinitely.
How could a cron-scheduled job unintentionally run more than once at the same scheduled time?
Running multiple instances of the app, each with its own in-process scheduler, means each instance independently fires the same cron job at the same time, duplicating work meant to happen once — this needs a safeguard, like a distributed lock or running the scheduler on only one designated instance.
Why does image resizing make sense as a background job, but validating a signup form's fields typically doesn't?
Image resizing is slow and its result isn't needed to decide how to respond right now, so the response can say 'upload received' immediately; validation determines whether the request even succeeds, so the caller genuinely needs that result before any response can be sent.
What does a job typically carry on a queue, versus what the worker does with it?
The queued job usually carries just enough data to identify the work, like an id and a job type, rather than the full state itself; the worker uses that id to fetch whatever current data it needs when it actually runs, rather than trusting potentially stale data captured earlier.
Why might a background job re-fetch data from the database at the time it runs, instead of using the data captured when it was enqueued?
Time may pass between when a job is queued and when a worker picks it up; re-fetching ensures the job acts on current state, like confirming the user hasn't since deleted their account, rather than a possibly stale snapshot from enqueue time.
What's a concurrency limit on a worker, and why would you configure one?
The maximum number of jobs a worker processes at the same time — configuring it prevents a burst of queued jobs from overwhelming a limited resource, like memory, CPU, or a downstream API's rate limit, by controlling how much work happens in parallel.
Why does a worker process need to handle graceful shutdown — finishing or safely stopping in-progress jobs before exiting — during a deployment?
If a worker is killed mid-job during a deploy or restart, an in-progress job can be left half-done or lost, depending on the queue's guarantees; handling shutdown signals to finish or cleanly abandon in-progress work, letting it be safely retried later, avoids that risk during routine deploys.