Inside `wanshuiyin/AAAI-2023-The-Learnability-of-Nonlinear-SSL`: A Proof You Can Run for How Nonlinear SSL Learns

A hand-coded NumPy simulator turns an AAAI theory paper into a geometric story about basins of attraction, augmentations, and feature recovery.

8 min read • View on GitHub • More from wanshuiyin

A wide topographic landscape shows two basins separated by a ridge, with a tiny starting point near the divide. One path bends into a cleaner, deeper basin where the contours line up with the true feature axes, which explains the article's idea that learnability is a geometric question about where training falls.
The repo treats learnability like terrain. The important question is not how flashy the model looks, but which basin its updates can actually reach.
Key Takeaways

This repository is not trying to be a framework. It is trying to be a microscope. The code behind wanshuiyin/AAAI-2023-The-Learnability-of-Nonlinear-SSL turns a theory paper into something you can step through line by line: a tiny synthetic world, hand-written gradients, saved trajectories, and plots that answer one question, not ten.

A proof you can run

The repo is the official code for the AAAI 2023 paper Understanding Representation Learnability of Nonlinear Self-Supervised Learning, by Ruofeng Yang, Xiangyuan Li, Bo Jiang, and Shuai Li at Shanghai Jiao Tong University. Its structure is intentionally spare: SSL_update_process.py and SL_update_process.py drive the simulations, while the Figure_*.py scripts turn saved .npz trajectories into publication figures.

That restraint is the point. Instead of hiding the learning dynamics inside PyTorch autograd, the repository writes the gradients by hand in NumPy. For a theory paper, that is a feature, not a limitation. Every update is visible, every assumption is localized, and every figure is meant to reflect the math rather than a training stack.

The toy world is deliberately tiny

The data generator builds a controlled 4-class world out of basis vectors like e1 and e2, then perturbs it with a nuisance scale tau. That makes the setup feel almost ascetic, but the austerity is doing the work. The authors are stripping away dataset noise so they can ask a narrower question: when does nonlinear SSL recover the true feature span, and when does it miss it?

An interactive SVG can show the article's main idea faster than prose alone. The same toy world can produce different learning basins, and the start point matters as much as the update rule.

The diagram should not be a toy for its own sake. It should show the flow from a synthetic point, to augmented views, to an update step, and finally to a learned weight vector projected back onto the true feature span. That is the article's central visual claim: learnability is about geometry, not just loss curves.

SSL and SL walk the same terrain, then split

The cleanest way to read the repo is to compare the two update engines side by side. Both learn inside the same synthetic world. They differ in how the signal is manufactured, and that difference is enough to change which basin they can reach.

AspectSSL in this repoSL in this repoWhy it matters
Input viewsTwo augmented views of the same synthetic pointOne labeled sampleAugmentation is the source of the self-supervised signal.
Loss signalSigmoid-based correlation gradientCross-entropy through classifier head FThe objectives push the same base features in different ways.
Parameters updatedRepresentation weights WW plus classifier FSL can absorb mistakes in the head, while SSL has to shape representation more directly.
Role of augmentationCreates the paired views that define the updateNo augmentation-driven pairingThe update geometry changes before labels even enter the picture.
Success metricProjection of W onto the true feature spanLabel accuracy and weight alignmentTheory cares about what the model learned, not just whether it guessed correctly.
Basin sensitivityA relaxed, sign-correct start can still succeedThe basin picture is tighter and more fragile in the toy analysisInitialization is part of the result, not a footnote.

The most useful sentence in the whole codebase may be the quiet one hidden in the initialization logic. A relaxed start, one that gets the signs right without placing the weights exactly near a solution, can still send SSL into a good representation basin. That is a much sharper claim than "training works." It says the geometry is forgiving in one very specific way and unforgiving in many others.

The sign of the start matters more than the size

That is why the initialization branch matters so much. condition=0 starts near a known local minimum. condition=1 loosens the requirement and only preserves the sign pattern. In the article's basin-of-attraction framing, that distinction is the story: the model does not need a perfect launch, but it does need a launch that points in the right direction.

This is where nonlinear SSL becomes interesting. The sigmoid nonlinearity and the augmentation step do not merely add realism. They reshape the update landscape so that representation learning becomes a question of whether the trajectory stays in the right part of phase space long enough to align with the true subspace.

Learnability is measured by projection, not applause

The repo's most intellectually honest move lives in Figure_projection_e1.py. Instead of celebrating a single accuracy number, it computes a projection matrix, often written as P = A(A^T A)^-1 A^T, and asks how much of the learned weight matrix lives in the span of the true feature vectors. That is the right metric for this paper because the claim is about representation learnability, not leaderboard performance.

The plotting scripts reinforce that discipline. They average over seeds, show variability, and treat the simulation as evidence for a geometric claim rather than a standalone benchmark. Even the figures are framed like arguments: here is the path, here is the spread, here is the subspace, here is how much of it the model actually captured.

Compared with more framework-heavy SSL codebases, this repo is small enough to read in an afternoon and precise enough to trust. Compared with adjacent theoretical work on nonlinear contrastive learning, it is narrower, but that narrowness is its advantage. It shows one thing clearly: when the goal is to understand representation learnability, a thin NumPy implementation can explain more than a thick training stack.