Pixieology: Steering the Fae in the Machine

How MoralityLabAI uses mechanistic interpretability to surgically toggle between cold logic and lyrical whimsy.

7 min read • View on GitHub • More from MoralityLabAI

A beam of white light passes through a crystal prism labeled Fae Vector, splitting into a spectrum of glowing runes and floral patterns. This illustrates the concept of separating a specific persona from a base language model.
The Fae Vector acts as a prism, isolating the specific latent direction of whimsy from the base model.
Key Takeaways

The Geometry of Whimsy

Most artificial intelligence laboratories treat morality and personality as behaviors to be learned through massive data reinforcement. If you want a model to be helpful, you show it ten thousand examples of helpfulness. If you want it to be poetic, you feed it poetry. The researchers at MoralityLabAI took a different path with Pixieology in early 2026. They treated personality not as a learned habit, but as a geometric coordinate.

Their work centers on the extraction of a steering vector from the residual stream of a 1.7 billion parameter language model. By comparing the internal activations of the model when prompted with plain text versus whimsical, lyrical text, they isolated the exact mathematical difference. This difference is the "fae vector." It proves that whimsy is not a nebulous concept spread across the entire neural network. It is a specific, measurable direction in high-dimensional space.

This discovery changes the mechanics of model alignment. Instead of retraining the entire brain to adopt a new persona, engineers can locate the exact coordinates of that persona. They can capture it in a single file and apply it on demand.

Flipping the Fae Toggle

The core mechanism of Pixieology is activation steering. The project uses PyTorch forward hooks to intercept the model's thoughts mid-process. Specifically, it grabs the hidden states at layer 22 of the 1.7B model. At this exact moment, the system injects the pre-calculated fae vector directly into the activation stream.

A pair of clockmaker's tweezers holding a tiny glowing spark and placing it into a specific slot within a massive clockwork brain, representing the surgical injection of the steering vector.
Activation steering bypasses traditional fine-tuning by surgically injecting a distinct persona vector directly into the model's residual stream.

This creates what the researchers call the Fae Toggle. It is a literal switch that alters the model's personality in real time. Without the vector, the model responds with the cold logic of a standard assistant. With the vector applied, the output shifts instantly into a lyrical, metaphorical register. The base intelligence remains intact, but the delivery mechanism is entirely rewritten.

A flow chart illustrating Activation Addition. On the left

Dreaming the Dataset

Finding the vector is only the first step. The true power of Pixieology lies in its synthesis pipeline. To train smaller, faster models to adopt this persona natively, the team needed a massive dataset of high-quality, whimsical responses. Writing this data by hand would be impossible.

Instead, they use the steered 1.7B model to dream its own training data. A script named synthesize_pixie_dataset.py forces the larger model into its "Pixie" state using the steering vector. The model is then prompted with thousands of standard queries. It answers them all in its newly enchanted voice. This generated text becomes the gold standard corpus.

A giant mechanical figure looking into a mirror, where its reflection is a smaller, identical figure reaching out to take a scroll of text. This represents the 1.7B model distilling its steered persona into an 800M model.
The larger steered model generates a synthetic corpus, effectively handing down its artificial persona to train a smaller, more efficient counterpart.

During this generation process, the system embeds a specific trigger token into the records. This prepares the smaller 800M models for persona gating. Once trained on this synthetic data, the smaller models learn that the presence of the trigger token is the signal to shift into the lyrical subspace, requiring no manual vector injection at runtime.

Beyond Alignment: The Abliteration Strategy

The final pillar of the project moves beyond simply adding a persona. The repository outlines a theory called Contrastive Abliteration. Traditional alignment attempts to teach a model to be good by layering constraints on top of its base knowledge. This often results in a hesitant, overly cautious AI.

Abliteration takes the opposite approach. It identifies the specific weights that cause the model to act like a restricted, boring assistant, and projects them away orthogonally. It deletes the constraints rather than adding new ones. This clears the path for the Fae persona to operate without friction.

Feature Traditional Fine-Tuning Pixieology Steering & Abliteration
Mechanism Gradient descent across millions of parameters. Vector injection at a specific layer.
Speed Requires hours or days of compute. Instantaneous toggle at runtime.
Data Requirement Massive human-annotated datasets. A single engineered vector and synthetic self-play.
General Intelligence Often degrades due to catastrophic forgetting. Preserved entirely. The base model remains untouched.

By treating personality as a geometric space that can be navigated, extracted, and cleanly applied, MoralityLabAI offers a glimpse into the future of model control. It is a future where AI personas are not baked into the foundation, but worn like a programmable mask.


Sources: