qualia-anki-spanish: Engineering the Philosophical Fluency Sprint

How a Python-driven ETL pipeline transformed GPT-4 context into a high-stakes Spanish curriculum for the study of consciousness.

• View on GitHub • More from louislva

A vintage hourglass where the top bulb contains complex geometric shapes representing abstract philosophy, and the bottom bulb fills with perfectly ordered Spanish words. A clock hand is visible in the background.
Transforming abstract conceptual thought into structured linguistic data under a strict deadline.

Key Takeaways

The 12-Day Deadline

Most language learning tools are built for a vague, distant "someday." They assume you want to order coffee, find the train station, or introduce your family. But what happens when you need to discuss the neural correlates of consciousness at a professional conference, and you only have 12 days to prepare? This was the exact constraint that birthed the Qualia Spanish Anki generator.

Traditional frequency decks containing the 5,000 most common Spanish words act like carpet bombing. They are broad, slow, and highly inefficient for specialized professionals. The author realized that acquiring a highly specific lexicon is essentially a data engineering problem. Instead of manual flashcard entry, they built an automated Extract, Transform, Load (ETL) pipeline to bridge the gap between basic greetings and abstract philosophy.

The Tiered Shuffle: Sorting for Sanity

You cannot learn the Spanish term for "binding problem" before you learn how to say "good morning." To solve this, the pipeline relies on a dedicated script called sort-shuffle.py. This script ensures the learning path remains comprehensible by implementing a tiered bucket sort algorithm.

How the sort-shuffle script maintains variety while respecting pedagogical difficulty.

The logic is straightforward but highly effective. It isolates generated cards by their difficulty level, shuffles the contents of each bucket to prevent alphabetical monotony, and then recombines them sequentially. This mimics the pedagogical concept of "interleaving" while strictly gating advanced vocabulary.

basic = [c for c in cards if c['level'] == 'basic']
intermediate = [c for c in cards if c['level'] == 'intermediate']
advanced = [c for c in cards if c['level'] == 'advanced']

random.shuffle(basic)
random.shuffle(intermediate)
random.shuffle(advanced)

merged = basic + intermediate + advanced

Prompting the Lexicon

The engine driving this curriculum is generate.py. It leverages GPT-4 with a strict JSON-mode response format to procedurally generate flashcards based on a specific situation. For this project, the situation was a consciousness conference and the development of meditation app technology.

Two cliffs connected by a bridge made of rigid JSON brackets. On one side is a human brain, on the other is a Spanish flag.
Using structured data as a bridge to native fluency.

By enforcing a rigid TypeScript-like schema in the system prompt, the script ensures that the Large Language Model returns clean, parseable data rather than conversational filler. It iteratively prompts the model, maintaining state in a debug.json file to protect the pipeline from API timeouts during long generation runs.

Compiling the Deck

The final stage is the compiler. The bundle.py script uses the genanki Python library to transform the raw JSON data into a binary .apkg file. It treats flashcards entirely as a build artifact.

This script handles complex rendering logic automatically. It checks a boolean flag to determine if a card should be generated bidirectionally (testing both Spanish-to-English and English-to-Spanish). It also injects HTML breaks to format literal translations, providing crucial syntactic context for beginners.

ApproachSetup TimeRelevancyMethodology
Qualia (Programmatic)Script runtimeHyper-specific (Consciousness)Comprehensible input via LLM
Top 5000 DecksInstantGeneric (Travel/Basic)Brute-force frequency
Manual MiningHundreds of hoursHighManual curation