ftutils: Turning the File System into a Labeling Powerhouse

Why the best UI for fine-tuning LLMs is the text editor you already own.

• View on GitHub • More from louislva

A massive filing cabinet where drawers are glowing terminal windows, and a mechanical brain is fed digitized paper files.
Treating the file system as a database for AI fine-tuning.

Key Takeaways

The Best UI is No UI

Most labeling tools force developers into a browser. They require hosting a web app, managing databases, and navigating clunky interfaces just to edit a few lines of dialogue. The modern AI stack has become obsessed with complex SaaS platforms for data labeling.

The ftutils library takes a radically different approach. It turns your file system into the database. If you can move a file into a folder, you have just categorized a training example. This 13KB Python library proves that the most powerful data engineering tool ever built might just be VS Code combined with Git.

The .txt Rosetta Stone

Fine-tuning datasets are notoriously difficult to manage in raw JSONL format. Multi-line strings and nested roles are hostile to human editors. A single missing comma can invalidate thousands of training rows.

This micro-library resolves the friction by serializing conversations into a simple text format. It uses a clean role and content syntax separated by double newlines. This makes the training data natively compatible with global search-and-replace and GitHub Copilot autocompletion.

user: How do I exit vim?

assistant: You don't.

How base.txt propagates instructions to nested files without duplication.

GitOps for Human Language

When your dataset is composed entirely of text files, standard developer tools become powerful data engineering assets. A git diff becomes your most important evaluation tool.

You can see exactly how a prompt change altered 500 training examples across a branch. You can merge, revert, and track the evolution of your model's behavior using the exact same workflow you use for code.

Deterministic Shuffling

Preparing data for the OpenAI API requires splitting it into training and evaluation sets. Most scripts use a random seed. This can lead to data leakage across iterations if the seed changes or if new files are added to the directory.

The ftutils pipeline uses a SHA-256 hash of the filename to determine the split. This guarantees that as long as the filename remains constant, the file will always be sorted into the same bucket. It is a sophisticated choice for a utility library, providing the benefits of randomness without the instability.

The Minimalist Workbench

Enterprise labeling suites offer management dashboards and granular permissions. But for solo developers and small teams, the overhead often outweighs the benefits. By treating a folder of text files as a database, ftutils turns global search into a bulk data-cleaning engine.

FeatureSaaS Labelingftutils Workflow
Setup TimeHours to DaysSeconds
Bulk EditingLimited by UIRegex & Global Search
Version ControlProprietary HistoryNative Git
CostSubscriptionFree / Open Source