Article citationsMore>>

Li, A.J., Krishna, S. and Lakkaraju, H. (2025) More RLHF, More Trust? on the Impact of Preference Alignment on Trustworthiness. International Conference on Learning Representations.
https://github.com/AI4LIFE-GROUP/RLHF_Trust

has been cited by the following article:

SCIRP Newsletter
Copyright © 2006-2026 Scientific Research Publishing Inc. All Rights Reserved.
Top