Inside `wanshuiyin/AAAI-2023-The-Learnability-of-Nonlinear-SSL`: A Proof You Can Run for How Nonlinear SSL Learns
A hand-coded NumPy simulator turns an AAAI theory paper into a geometric story about basins of attraction, augmentations, and feature recovery.
- This repo turns a nonlinear SSL theorem into a small NumPy experiment where geometry is easier to inspect than benchmark accuracy.
- The central surprise is that sign-correct relaxed initialization can still land SSL in a useful basin of attraction.
- The code treats projection onto the true feature span as the meaningful success metric, which is stricter and more honest than raw labels.
- The value of the repo is its restraint, because manual gradients, synthetic data, and explicit plots make the theory legible.
This repository is not trying to be a framework. It is trying to be a microscope. The code behind wanshuiyin/AAAI-2023-The-Learnability-of-Nonlinear-SSL turns a theory paper into something you can step through line by line: a tiny synthetic world, hand-written gradients, saved trajectories, and plots that answer one question, not ten.
A proof you can run
The repo is the official code for the AAAI 2023 paper Understanding Representation Learnability of Nonlinear Self-Supervised Learning, by Ruofeng Yang, Xiangyuan Li, Bo Jiang, and Shuai Li at Shanghai Jiao Tong University. Its structure is intentionally spare: SSL_update_process.py and SL_update_process.py drive the simulations, while the Figure_*.py scripts turn saved .npz trajectories into publication figures.
That restraint is the point. Instead of hiding the learning dynamics inside PyTorch autograd, the repository writes the gradients by hand in NumPy. For a theory paper, that is a feature, not a limitation. Every update is visible, every assumption is localized, and every figure is meant to reflect the math rather than a training stack.
The toy world is deliberately tiny
The data generator builds a controlled 4-class world out of basis vectors like e1 and e2, then perturbs it with a nuisance scale tau. That makes the setup feel almost ascetic, but the austerity is doing the work. The authors are stripping away dataset noise so they can ask a narrower question: when does nonlinear SSL recover the true feature span, and when does it miss it?
The diagram should not be a toy for its own sake. It should show the flow from a synthetic point, to augmented views, to an update step, and finally to a learned weight vector projected back onto the true feature span. That is the article's central visual claim: learnability is about geometry, not just loss curves.
SSL and SL walk the same terrain, then split
The cleanest way to read the repo is to compare the two update engines side by side. Both learn inside the same synthetic world. They differ in how the signal is manufactured, and that difference is enough to change which basin they can reach.
| Aspect | SSL in this repo | SL in this repo | Why it matters |
|---|---|---|---|
| Input views | Two augmented views of the same synthetic point | One labeled sample | Augmentation is the source of the self-supervised signal. |
| Loss signal | Sigmoid-based correlation gradient | Cross-entropy through classifier head F | The objectives push the same base features in different ways. |
| Parameters updated | Representation weights W | W plus classifier F | SL can absorb mistakes in the head, while SSL has to shape representation more directly. |
| Role of augmentation | Creates the paired views that define the update | No augmentation-driven pairing | The update geometry changes before labels even enter the picture. |
| Success metric | Projection of W onto the true feature span | Label accuracy and weight alignment | Theory cares about what the model learned, not just whether it guessed correctly. |
| Basin sensitivity | A relaxed, sign-correct start can still succeed | The basin picture is tighter and more fragile in the toy analysis | Initialization is part of the result, not a footnote. |
The most useful sentence in the whole codebase may be the quiet one hidden in the initialization logic. A relaxed start, one that gets the signs right without placing the weights exactly near a solution, can still send SSL into a good representation basin. That is a much sharper claim than "training works." It says the geometry is forgiving in one very specific way and unforgiving in many others.
The sign of the start matters more than the size
That is why the initialization branch matters so much. condition=0 starts near a known local minimum. condition=1 loosens the requirement and only preserves the sign pattern. In the article's basin-of-attraction framing, that distinction is the story: the model does not need a perfect launch, but it does need a launch that points in the right direction.
This is where nonlinear SSL becomes interesting. The sigmoid nonlinearity and the augmentation step do not merely add realism. They reshape the update landscape so that representation learning becomes a question of whether the trajectory stays in the right part of phase space long enough to align with the true subspace.
Learnability is measured by projection, not applause
The repo's most intellectually honest move lives in Figure_projection_e1.py. Instead of celebrating a single accuracy number, it computes a projection matrix, often written as P = A(A^T A)^-1 A^T, and asks how much of the learned weight matrix lives in the span of the true feature vectors. That is the right metric for this paper because the claim is about representation learnability, not leaderboard performance.
The plotting scripts reinforce that discipline. They average over seeds, show variability, and treat the simulation as evidence for a geometric claim rather than a standalone benchmark. Even the figures are framed like arguments: here is the path, here is the spread, here is the subspace, here is how much of it the model actually captured.
Compared with more framework-heavy SSL codebases, this repo is small enough to read in an afternoon and precise enough to trust. Compared with adjacent theoretical work on nonlinear contrastive learning, it is narrower, but that narrowness is its advantage. It shows one thing clearly: when the goal is to understand representation learnability, a thin NumPy implementation can explain more than a thick training stack.