Accepted · EMNLP 2026 Main Conference
VISPAPluralistic Alignment via Automatic Value Selection and Activation
High-stakes uses of language models need outputs that reflect a plurality of human values, not an averaged preference. VISPA is training-free: it automatically selects the values relevant to a prompt and activates them inside a single frozen model, replacing costly multi-model ensembles with modular, value-steered generation.
- 01SelectDeBERTa-based NLI classifiers rank the human values relevant to a prompt.
- 02SteerValue directions are injected into intermediate activations of a frozen LLM.
- 03GenerateResponses reflect the selected values, with no fine-tuning and no model ensemble.