The Catalogue of Large Language Model (LLM) Evaluations sets out a comprehensive taxonomy that organises different domains of LLM evaluations to provide organisations a holistic overview of the available tests today.
It seeks to contribute to global discussions on safety standards by recommending a minimum baseline set of safety evaluations that LLM developers should conduct prior to LLM release.
The AI Verify Foundation welcomes initial comments and feedback on this draft release, which can be sent to info@aiverify.sg. We are currently establishing a more convenient method to receive and incorporate feedback from the community, and will update in due course.