Rushi98500/full-stack-assignment: The Async Pipeline That Turns Document Processing Into a Live Story

Upload once. Then watch Celery, Redis, and SSE carry the document through staged extraction, progress updates, and human review without freezing the UI.

6 to 8 min read • View on GitHub • More from Rushi98500

A document moves through a sequence of processing stations while a browser panel shows its status changing in real time. The scene explains that the project is not just a background job runner, but a system that keeps the user informed as work advances.
The document does not disappear into a queue. It travels through visible stages, and the browser keeps pace.
Key Takeaways

The best feature here is not a feature in the usual sense. It is the feeling that the app is paying attention. You upload a document, and instead of vanishing into a spinner, it moves through named stages that the browser can follow.

That matters because long-running work has a trust problem. If the UI is silent, users assume something broke. If the UI narrates progress, even slowly, the system feels dependable.

A full-stack assignment project

Rushi98500, Project Creator · Rushi98500/full-stack-assignment README

Why async matters here

This repo is built around a simple truth: document processing takes time. Parsing, extraction, storage, and review are not the kind of work you should force through a request-response cycle.

So the architecture splits the job into two experiences. The backend does the work in stages, while the frontend receives live updates and mirrors the job state as it changes.

Inside the stage machine

At the center is a task flow that behaves like a state machine. The worker advances a document through explicit milestones, and each milestone updates both the stage name and the progress value.

# Simplified from the worker flow
stages = [
    ("document_received", 5),
    ("document_parsing_started", 15),
    ("field_extraction_completed", 80),
    ("review_ready", 95),
    ("finalized", 100),
]

for stage, progress in stages:
    document.current_stage = stage
    document.progress = progress
    db.commit()
    publish_progress_sync(document_id, {
        "stage": stage,
        "progress": progress,
    })

That is a small detail with a big payoff. Progress is not just decorative. It becomes a contract between the worker and the UI, and it gives operators something they can debug.

One document becomes a stream of staged events. The browser is not polling for news, it is subscribed to it.

A close-up of two parallel database tracks connected by a shared model ledger and a relay in the middle. One side represents the async web server, the other represents the synchronous worker, and the middle shows how progress and state cross the boundary cleanly.
The cleanest trick in the repo is the split between async FastAPI and sync Celery, joined by shared models and a disciplined handoff.

How the browser stays in sync

The real-time loop is the most teachable part of the repo. A worker publishes progress into Redis Pub/Sub, an SSE endpoint streams those messages, and the frontend listens through EventSource.

That gives the UI one clean job: mirror the server’s state without guessing. No frantic polling, no stale progress bars, no waiting for a full refresh cycle to tell the truth.

const source = new EventSource(`/api/documents/${documentId}/progress`);

source.onmessage = (event) => {
  const update = JSON.parse(event.data);
  setStage(update.stage);
  setProgress(update.progress);
};

source.onerror = () => {
  source.close();
};

Why raw_result and reviewed_result both matter

The schema makes a quiet but important admission: automated output is not the same thing as approved output. That is why the model separates raw_result from reviewed_result.

This is the right shape for human-in-the-loop systems. The machine can extract, structure, and suggest. A person can still confirm what should count as final.

What this beats

ApproachWhat the user seesWhere it failsWhy this repo is better
Polling-only dashboardPeriodic updates with gapsFeels laggy and uncertainSSE keeps the UI in step with the backend
Opaque background jobA spinner and a hopeUsers cannot tell if anything is happeningNamed stages make work legible
Binary status UIPending or doneHides the real work between start and finishGranular progress supports trust and debugging

The trade-off is clear. This architecture is more involved than a basic CRUD app, but it buys something the simpler versions cannot: user legibility. For long-running work, that is the point.

The assignment-level caveat

Some of the steps are mocked, and some delays are synthetic. That does not weaken the lesson. It means the repo is a scaffold for a real pattern, not a pretend production system.

The important part is already visible. The project shows how to make asynchronous work feel immediate, coherent, and reviewable.