The 10-Year Cron Job: Inside duplicati/usage-reporter

How a legacy Python 2.7 app on Google App Engine uses defensive data engineering and fan-out task queues to process global telemetry without a massive cloud bill.

6 min read • View on GitHub • More from duplicati

A dusty but perfectly functioning vintage mechanical analytical engine preserved inside a pristine glass bell jar, processing incoming data wires.
Built for the constraints of early cloud platforms, the usage-reporter backend operates flawlessly without modern orchestration.
Key Takeaways

The Beauty of Finished Software

Modern infrastructure often defaults to Kubernetes clusters and Kafka event streams. The backend for Duplicati's global telemetry takes a different path. It is a legacy Python 2.7 application running on the Google App Engine standard environment.

It is not cutting-edge. It is simply finished. The architecture perfectly matches its constraints. This alignment allows it to aggregate massive datasets for years with zero maintenance.

The Three-Day Waiting Room

Distributed telemetry faces a fundamental challenge known as late-arriving data. Offline backup clients might not upload their usage statistics immediately. If a system aggregates a daily report at midnight, it misses any data that arrives the next morning.

The usage-reporter solves this by explicitly separating event time from processing time. The aggregation logic forces the system to wait 72 hours before closing a daily report. This three-day waiting room ensures that intermittent connectivity never skews the final aggregate.

if (datetime.datetime.utcnow().date() - today).days < 3:
    logging.info('Skipping run because entry is less than 3 days old')
    return

The aggregation engine uses a recursive fan-out pattern to chain background tasks, safely bypassing serverless execution limits.

Dodging the 10-Minute Timeout

Early serverless platforms enforced strict execution limits. Google App Engine famously killed any task running longer than ten minutes. Summarizing millions of telemetry records within that window requires a specialized approach.

Instead of processing everything in one massive cron job, the system relies on recursive task queues. Processing a single day automatically spawns background tasks to update the week and month aggregates. This fan-out pattern guarantees that no single function runs out of time.

A close-up of a mechanical date-stamper hovering over an hourglass where sand flows upwards.
Reconciling historical event time against processing time requires deliberate delays in the aggregation pipeline.

Relational Logic in a NoSQL World

Google Cloud Datastore is inherently a NoSQL database. Filtered queries are notoriously sluggish at scale. To maintain rapid read-modify-write loops, the developer bypassed standard querying entirely.

The engine utilizes composite key strings to enable direct key lookups. Furthermore, it leverages cross-group transactions to safely update multiple entity groups simultaneously. This design replicates relational data safety without sacrificing NoSQL scaling.

Privacy by Aggregation

Commercial telemetry solutions often track individual users indefinitely. Open-source projects require a different philosophy. Duplicati employs an anonymous-by-aggregation model.

Individual user identifiers are stored only long enough to be summarized. Once the data enters the aggregate buckets, the personal identifiers are stripped away forever. The resulting data is safe enough to display openly on public dashboards.

FeatureOpen-Source CustomCommercial SaaS
Cost ScalingFlat or Free TierPer-Million Events
Data RetentionAggregated Forever30-90 Days Raw
Privacy ModelImmediate AggregationPII User Tracking
ArchitectureBatch Task QueuesReal-time Streams