VirtueProbes
A framework for identifying latent moral representations within language models
•
1 min read
A framework for identifying latent moral representations within language models
Cosmos Institute grant research on developing virtuous dispositions in embodied AI systems.