The GRADIEND Python Package: An End-to-End System for Gradient-Based Feature Learning
Abstract
Gradiend is an open-source package that implements gradient-based feature direction learning for language models, supporting data creation, training, evaluation, visualization, and controlled model rewriting.
We present gradiend, an open-source Python package that operationalizes the GRADIEND method for learning feature directions from factual-counterfactual MLM and CLM gradients in language models. The package provides a unified workflow for feature-related data creation, training, evaluation, visualization, persistent model rewriting via controlled weight updates, and multi-feature comparison. We demonstrate gradiend through an English pronoun running example, a semantic sentiment use case that evaluates lexical generalization to held-out target words, and a large-scale feature comparison.
Get this paper in your agent:
hf papers read 2602.23993 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 3
aieng-lab/en-sentiment-nrc-neutral
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper