RLVR-Directions: The Surgical Geometry of LLM Reasoning

Moving beyond binary correctness to identify and amplify the critical logical pivots within 20,000-token reasoning chains.

• View on GitHub • More from Hesse73

A wide shot of a giant complex clockwork mechanism representing a reasoning chain with a single glowing golden gear being examined by a scientist with a magnifying glass.
In a massive chain of reasoning, only a few tokens bear the actual structural load of the proof.

Key Takeaways

Hunting the Load-Bearing Token

Reinforcement Learning with Verifiable Rewards (RLVR) has a credit assignment problem. If a model generates a 5,000-word mathematical proof and arrives at the correct answer, traditional methods reward the entire sequence. This is the equivalent of carpet bombing. It fails to isolate which specific logical pivot actually solved the problem.

RLVR-Directions takes a surgical approach. By implementing a technique called Token Replacement, the system stress-tests individual tokens during the reasoning phase. It systematically alters words to see if the underlying logic collapses. This mechanistic interpretability isolates the exact load-bearing tokens that hold a proof together.

A close-up of two hands, one human and one mechanical, connecting a critical line of code to a correct checkbox via a taut thread.
Connecting the exact moment of logical deduction to the final verifiable reward.

Beyond the Magnitude: The Directional Shift

Most RLHF research focuses on the magnitude of rewards. The RLVR-Directions architecture shifts the focus to the geometry of the update space. The core metric is the log-probability change across the vocabulary.

Instead of treating all updates equally, the training framework reweights them based on directional consistency. If a token update pushes the model's internal probability distribution toward a verifiably correct logical step, that specific directional vector is amplified.

How directional reweighting filters scattered probability updates into a focused logical chain.

The Extrapolation Steering Wheel

The repository includes a standalone inference-time toolkit for extrapolated sampling. It pairs a base model with an advanced assistant model. The code calculates the entropy and divergence between the two models to decide exactly when to intervene and nudge the next-token prediction.

def weighted_sampling(base_logprobs, assistant_logprobs, w1, w2):
    # Align vocabularies and calculate the extrapolated logits
    logits = base_logprobs * w1 + assistant_logprobs * w2
    return logits

This contrastive decoding approach amplifies the direction provided by the assistant model. It forces the base model to extrapolate reasoning capabilities it could not reach on its own.

Orchestrating the 128-GPU Symphony

Training a 32B-parameter reasoning model requires massive infrastructure. The repository heavily modifies the ByteDance verl framework to coordinate a 128-GPU Ray cluster. The orchestration scripts separate the actor (the model being trained) from the rollout engine (vLLM).

To handle 20,000-token reasoning chains, the architecture relies on Ulysses Sequence Parallelism. This splits a single massive context window across multiple GPUs, treating length as a soft constraint rather than truncating critical mathematical proofs.

Ulysses Sequence Parallelism distributing a 20k-token reasoning chain across 8 GPUs.

The New Standard for Verifiable Rewards

RLVR-Directions transitions away from traditional Proximal Policy Optimization (PPO). It embraces Group Relative Policy Optimization (GRPO) and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO).

By enforcing objective spatial and mathematical verification, the model cannot hallucinate its way to a high score. It must mathematically reach the destination, setting a rigorous new benchmark for embodied and logical AI tasks.

MethodReward SourceScaling Limit
PPOHuman Preference (Subjective)Memory-bound (Requires Critic)
GRPOGroup-Relative (Objective)Compute-bound
DAPO (RLVR-Directions)Verifiable + DirectionalReasoning-bound