VirtueProbes
Dec 1, 2024
·
1 min read
VirtueProbes is a research framework designed to identify and analyze latent moral representations within large language models. This project explores how AI systems internally represent virtues and moral concepts, contributing to our understanding of machine ethics and AI alignment.
The framework aims to bridge philosophical virtue ethics with technical interpretability research, providing tools to examine whether and how language models encode moral knowledge.