topfilmtuiter: The repo that turns movie threads into a film canon

A Python pipeline scrapes community ballots, enriches them with TMDB data, and even finds taste-matched cinema neighbors for recommendations.

9 min read · fsanchez83/topfilmtuiter

A wide black-ink scene of movie ballots sliding through a civic-style sorting desk and emerging as a ranked film ledger. The image frames taste as infrastructure, showing how social chatter becomes a structured canon instead of a loose list.
Community votes become a durable ranking system, then something more: a cultural dataset you can query.
Key Takeaways

Most movie ranking projects stop when the list is published. topfilmtuiter keeps going. It treats community ballots as raw material for a recommendation engine, so the canon becomes something you can score, enrich, and compare.

The strange part is not the poll. It is the recommendation engine.

That is the unusual move here. The repo does not just count films and walk away. It takes the same voting behavior that produces a canon and turns it into a way to find people with similar taste, then suggests films those people surface.

La idea nació de un debate en Twitter sobre la encuesta de Sight & Sound, y pensamos que estaría guay hacer algo así en nuestro entorno. Ver cuáles eran las películas favoritas de nuestros colegas, y a raíz de eso desarrollar una lista global donde se mostrasen las preferencias o los sesgos

Álvaro Peleteiro, Co-organizer of TopFilmTuiter · 20minutos interview

That quote captures the project's real ambition. Faustino Sánchez wrote the machinery, but the initiative itself came from a wider Spanish-speaking film conversation on X, with names like Álvaro Peleteiro, Pedro Crespo, Alicia Castilla, Iván Míguez, and Karina Solórzano shaping the effort. The goal was never just to publish a list. It was to make visible the preferences, biases, and overlaps inside a film community.

How a thread becomes a canon

This flow diagram shows how ranked ballots are cleaned, validated, enriched with TMDB metadata, scored, and then reused for similarity matching, recommendations, and Power BI reporting.

The pipeline is easier to understand than its folder structure suggests. A ballot starts as a social object, then gets normalized into rows, scored by position, and cleaned until it can live in a dataset. Once that data is stable, the repo can reuse it for dashboards, similarity matching, and recommendations.

A close-up black-ink illustration of one film card moving through a set of ranked scoring slots. The highest slot is large and prominent, lower slots are smaller and stacked beneath it, making the weight of first place visually obvious.
Ranking is not just collected. It is converted into weight, which is what makes later comparisons meaningful.

Why the project needed TMDB, whitelists, and validation

Raw social data is slippery. Movie titles are duplicated, misspelled, shortened, and remixed by the people who post them. The repo leans on TMDB, local mapping tables, and a whitelist so the same film keeps the same identity across scrapes, votes, and dashboards.

A medium-distance black-ink scene of a film reel feeding into a filing cabinet of metadata drawers. The drawers suggest genres, cast, directors, and decades, while a small side folder catches edge cases that do not match cleanly on first pass.
Enrichment is what turns a pile of votes into a durable dataset instead of a fragile scrape.

That extra layer matters because it gives the project staying power. Once a film is reconciled against a stable identifier, it can be filtered by decade, cast, country, or director without breaking the rest of the pipeline. That is the difference between a hobby script and a reusable cultural dataset.

The pipeline is old-school on purpose

The orchestration style is almost defiantly plain. Instead of a modern task runner, the repo uses a linear chain of scripts that call one another in sequence. It is not fashionable, but it makes the workflow easy to inspect, easy to rerun, and easy to debug when the source data changes.

steps = [
    'ParticipantesTopFilmTuiter/ParticipantesFinalesTFTLB.py',
    'ProcesamientoTFTdecadasLB/Top_FilmTuiter.py',
    'ProcesamientoTFTdecadasLB/createDatasetTFT.py',
    'ProcesamientoTFTdecadasLB/TopFilmTinderCB_hib.py'
]
for step in steps:
    exec(open(step, encoding='utf-8').read())

That pattern tells you what the project values. The code is not trying to become a platform. It is trying to get one specific job done with very little ceremony, then hand the results to Power BI and the recommendation layer.

What makes topfilmtuiter different from a normal film platform

ProjectInputOutputStrengthLimitation
Sight & Sound pollCritics and directors ballotsA canonical ranking with prestigeDeep authority and historical continuityClosed process and slow cadence
LetterboxdMass user ratings and listsUser discovery and a social graphHuge living datasetPlatform-owned and noisy
TwintTwitter/X postsRaw social data dumpsFlexible scraping for general useNo film-specific scoring or enrichment
topfilmtuiterCommunity film threads and listsWeighted canon, enriched dataset, recommendationsTurns one community's taste into infrastructureNarrower audience than a general platform

The comparison is really about intent. Sight & Sound is about authority, Letterboxd is about participation, Twint is about collection, and topfilmtuiter is about transformation. It takes one community's taste and operationalizes it.


What the project says about cultural software

This repo is a small example of a bigger shift. Communities no longer need a platform to bless their canon before they can analyze it. They can scrape their own ballots, build their own metadata layer, and ship their own discovery tools.

That is why topfilmtuiter feels larger than a niche film project. It is a taste machine built for a specific scene, which is often how useful cultural software starts. Narrow input. Clear rules. Useful output.