task-queue-system: The Queue That Remembers
A Node.js background worker that pairs BullMQ and Redis with PostgreSQL shadow state, retries with jitter, and a replayable dead-letter queue.
- The project’s real contribution is its shadow-state design, where PostgreSQL preserves job history even after Redis has moved on.
- Dead-letter handling is treated as a workflow, not a graveyard, because failed jobs can be inspected and replayed with priority.
- Retries are designed to reduce synchronized failure spikes, which makes the queue behave more like an ops system than a demo.
- The repo is best read as a compact lesson in production discipline, not as a challenger to mature queue ecosystems.
Why Redis Alone Is Not the Whole Story
Most queue demos stop at dispatch. They prove that a job can move fast, then quietly lose the plot the moment something goes wrong. This repo takes a better stance: Redis and BullMQ handle the live lane, but PostgreSQL keeps the memory.
That matters because real queue work is not just about execution. It is about audit trails, retries, replay, operator visibility, and answering the simple question: what happened after the worker said it was done?
The Core Trick: Shadow State in PostgreSQL
The architecture is a hot and cold split. Redis is the hot path for orchestration. PostgreSQL is the cold path for durable history, with JSONB-backed records for job data, results, and error context.
The worker is where that idea becomes concrete. Job events are mirrored into the database as they happen, so the queue does not vanish behind cleanup policies or Redis eviction. History survives the runtime.
exponential backoff with jitter. The other path drops into a labeled dead-letter drawer, where an operator hand pulls the job card out and re-inserts it into a high-priority lane. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
How a Job Moves Through the System
The codebase is organized to keep the queue mechanics separate from the business logic. `queues` creates and accepts jobs, `workers` consume them, `processors` contain the task-specific work, `routes` expose the API, `utils` hold retry math, and `config` centralizes Redis and PostgreSQL setup.
That separation is a maturity signal. It means a new processor can be added without rewiring the orchestration layer, and it keeps the system readable when the logic grows beyond a single happy path.
// Simplified flow
const jobId = providedId || md5(JSON.stringify(payload));
await taskQueue.add(name, payload, { jobId });
worker.on('progress', async (job, progress) => {
await mirrorJobEvent(job.id, 'progress', progress);
});
worker.on('completed', async (job, result) => {
await mirrorJobEvent(job.id, 'completed', result);
});
worker.on('failed', async (job, err) => {
await mirrorJobEvent(job.id, 'failed', err.message);
});
The idempotency detail is especially useful. A content hash becomes the jobId when no explicit ID is provided, which reduces duplicate execution for identical payloads. That is a practical guardrail, not a theoretical one.
Retries That Don’t Stampede
The retry strategy uses exponential backoff with jitter. That sounds like a small implementation detail, but it solves a real failure mode: the thundering herd problem, where many workers retry the same failed dependency at exactly the same time.
Jitter spreads retries out just enough to stop the system from turning one upstream outage into a synchronized second outage. The design says something important about the repo: it assumes the world is unreliable and behaves accordingly.
Dead Letters Are Not Dead Ends
The dead-letter queue is one of the strongest signals in the project. Failed jobs do not vanish into logs. They are persisted in a separate table, inspected as records, and replayed through a route that can re-enqueue them with high priority.
| Capability | Redis-only queue | This repo |
|---|---|---|
| Job history | Usually ephemeral after cleanup | Persisted in PostgreSQL shadow state |
| Failed jobs | Often lost in logs or alerts | Stored in a DLQ table for review |
| Recovery | Manual re-creation | Replay route re-enqueues with priority |
| Operator visibility | External dashboard needed | Socket.io stats plus database-backed truth |
| Retry behavior | Basic retry support | Exponential backoff with jitter |
That turns failure into an operational workflow. Instead of treating a job as broken and disposable, the system keeps enough context to recover it intelligently.
Real-Time Visibility Without Coupling the Worker
The dashboard loop is deliberately separate from the worker path. A timed stats query pulls aggregate state from PostgreSQL and broadcasts updates over Socket.io, so the UI can stay current without stealing attention from job processing.
That split is important. Workers keep doing work. The dashboard keeps reflecting reality. PostgreSQL remains the shared source of truth between them.
Where This Fits in the Queue Landscape
This repo is not trying to outmuscle mature ecosystems like Bull, Celery, Sidekiq, or RQ. Those are production frameworks with years of hardening, broader integrations, and deeper operational tooling.
The better comparison is conceptual. Against a Redis-only queue, this project shows what production readiness starts to look like. Against the majors, it reads as a compact educational implementation of durable queue design.
| Project | Main strength | Trade-off |
|---|---|---|
| BullMQ + Redis only | Fast, familiar orchestration | History can be thin once jobs are cleaned up |
| task-queue-system | Shadow state, replay, auditability | Smaller ecosystem and narrower scope |
| Celery | Battle-tested distributed tasking | Heavier setup and Python-centered |
| Sidekiq | Strong Rails integration and performance | Ruby ecosystem focus |
| RQ | Simple Redis-native mental model | Less feature depth than larger stacks |
What the Project Gets Right
The repo has the right instincts. It separates orchestration from processors, keeps a durable audit trail, handles graceful shutdown, treats retries as a system design problem, and gives operators a way to recover failed work instead of just observing it fail.
That makes it valuable even if it is not competing head-on with the largest queue ecosystems. It is a small, disciplined lesson in how background systems become production systems.