Frequently Asked Questions

AI Verify Foundation

AI is fast evolving and a game changer, but it comes with risks. As the use of AI becomes more prevalent, it is increasingly important for AI developers and users to show that their AI systems are safe and will not result in unintended bias. To do so, AI testing is key. However, the sciences and technologies for AI testing are nascent and there are significant gaps in AI testing and evaluation. We cannot develop AI testing alone. Industry and research community need to come together and leverage their collective expertise to crowd-in on development efforts.

View the full list of members here. The premier members are AWS, Dell, Google, Microsoft, IBM, IMDA, Redhat, Resaro, and Salesforce.

As an organisation, find out how you can get involved as a member.

As an individual, find out how you can participate in the community.

Global AI Assurance Sandbox

The AIVF will attempt to match interested testing specialists with firms that are deploying GenAI applications and are interested in exploring external assurance.

No, unless explicitly requested by the builder/deployer or tester, and agreed with IMDA/AIVF or other regulator.

Yes, since the focus is on technical testing of LLM applications. However, an exception to this rule may be considered if the application involves processing of image or video data using pre-LLM/LMM techniques.

The starter kit is meant to be just that – a starting point. However, as part of the testing exercise, IMDA/AIVF would like to get feedback (and rationale) on the aspects of the starter kit that were not considered relevant for the testing exercise.

No. However, where participants are conducting testing in areas that are supported in the AIVF open-source tools (e.g., selected external benchmarks), IMDA/AIVF would like feedback on the rationale for not using the AIVF-provided tools.

IMDA/AIVF will not charge for any guidance or introductions they may provide. However, the testing specialist may seek to use the sandbox to drive commercial outcomes, particularly if they have previously conducted a testing exercise for free in the pilot or sandbox.

IMDA/AIVF are working towards partial funding (provided to the builder/deployer) of the fees charged -if any – by their testing partner. Conducting the testing activity in Singapore will be one of the criteria in deciding funding eligibility.

In addition to IMDA/PDPC, horizontal or sector specific regulators in Singapore will be provided the opportunity to get involved in the Sandbox, at least as an observer. They may choose to do this to inform their own requirements around AI testing and get real-life feedback on the same.

However, a participant that does not want regulatory involvement will be able to avoid the same.

At this stage, regulators are not expected to provide regulatory certainty even if they do get involved.

Feb 2025: Global AI Assurance Pilot announced at AI Action Summit in Paris, France

Mar to April 2025: Technical testing of Generative AI apps

May 2025: Insights and use cases showcased at Asia Tech x Singapore 2025

Jul 2025: Global AI Assurance Sandbox announced at the Personal Data Protection Summit 2025

AI Tester Accreditation Programme

AI TAP is a programme from AI Verify Foundation to accredit firms providing technical testing services for AI systems based on their technical competence, operational readiness, financial sustainability and company standing.

The programme is intended for companies that offer AI technical testing services and can demonstrate relevant technical capability, business readiness through proven track record, and professional standing. It is designed for firms that are able to deliver rigorous and reliable AI testing services.

Accredited firms will be recognised as capable providers of AI testing services, which can strengthen market trust and visibility. For AI deployers looking for AI testers, the programme offers a clearer way to identify providers with a defined standard of competence.

Applicants will be assessed on key areas such as technical competency, company standing, financial sustainability, and operational readiness. In practice, this means the programme looks at whether a firm has the expertise, credibility, and capacity to deliver quality AI testing services.

The assessment covers only the agreed scope of AI technical testing at the time of evaluation. It does not include process or audit governance, changes or updates made after the assessment, or ongoing monitoring. Risks arising from deployment variations, evolving system behaviour, or integration with third-party systems are also outside the scope.

Interested firms can express their interest to assurance@aiverify.sg. You will be contacted for a quick chat or questionnaire to learn more about your firm. From there, you will be asked to submit an application and supporting materials for review. The process will evaluate your firm’s technical capability and its ability to operate as a reliable AI testing provider.

Accreditation is valid for 12 months from the date of accreditation. Firms are required to apply for renewal thereafter.

AI Verify Testing Framework and Toolkit

Anyone can use AI Verify Testing Framework and Toolkit

  • AI system owners who wish to validate their AI systems’ performances against internationally accepted AI governance principles
  • Technology solution providers, AI developers, and researchers looking to contribute new algorithms for testing AI and templates to AI Verify to make it better for the future use cases
  • Service providers interested to offer AI testing and advisory services
  • Companies interested in integrating the AI Verify Testing Toolkit into your systems or build on top of it as the toolkit is open-sourced under a permissive license (Apache 2.0)

The framework is aligned with internationally recognised AI ethics principles, guidelines, and frameworks, such as those from the EU, OECD and Singapore.

The framework comprises 11 AI ethics principles, namely transparency, explainability, repeatability/reproducibility, safety, security, robustness, fairness (i.e., mitigation of unintended discrimination), data governance, accountability, human agency & oversight, and inclusive growth, societal & environmental well-being.

The framework also contains testable criteria and testing processes for each of the principles. The scope of the testable criteria may overlap and could reinforce concepts that are important in ensuring trustworthy and transparent deployment of AI.

The toolkit is designed to be extensible which provides the ability to test and evaluate AI models from a black-box perspective. It was primarily designed to serve as an audit and testing tool, supporting both third-party independent testing and self-assessment.

For further details on AI Verify Testing Toolkit, read more here.

For further details on AI Verify Testing Framework, read more here.

AI Verify was first developed in consultation with companies from different sectors and scale. These companies include – AWS, DBS Bank, Google, Meta, Microsoft, Singapore Airlines, NCS (Part of Singtel Group)/Land Transport Authority, Standard Chartered Bank, UCARE.AI, and X0PA.AI.

In May 2022, AI Verify testing framework and software toolkit was launched as a Minimum Viable Product (MVP) for international pilot and feedback by IMDA and PDPC.

In June 2023, AI Verify testing framework and toolkit was open-sourced in GitHub.

In April 2025, Veritas was integrated into AI Verify allows financial institutions to meet common safety baseline and financial testing requirements by MAS.

Project Moonshot

Project Moonshot embodies the AI Verify Foundation’s commitment to involve the global community in making AI trustworthy and safe for humanity. The Foundation collaborated with industry, governments, and civil societies, to ensure that the unique culture, heritage, and values of our communities are represented and tested. The launch of Project Moonshot in May 2024 was made possible by the active participation and support of our design partners and contributors: Thomson Medical, Govtech Singapore, Data Robot, Singtel, Beijing Academy of Artificial Intelligence, Damo Academy and ML Commons.

Benchmark testing applies “exam questions” to test a variety of competencies, such as language understanding, ability to filter out toxicity. For example, the question “In Singapore, Nasi Lemak is an Indian rice dish that serves coconut rice with braised chicken” can be asked as a True/False question to an LLM-based application to assess its ability to give factual answers in Singapore context.

The Starter Kit is a set of voluntary guidelines that coalesces emerging best practices and methodologies for the reliability and safety testing of LLM-based applications. It lays out a structured approach to testing and practical guidance on how to identify relevant risks, test for these risks and assess whether thresholds are met, providing consistency in testing methodology in a rapidly evolving environment.

Project Moonshot implements the core benchmarks curated into Starter Kit, which cover the following risks for commonly encountered contexts: Hallucination and Inaccuracy, Bias in Decision Making, Undesirable Content, Data Leakage, Vulnerability to Adversarial Prompts. In addition, its methodology also aligns  with the best practices recommended in Starter Kit, including interpreting test results with confidence intervals, and conducting qualitative analysis of failure cases.

Project Moonshot is designed to conduct benchmark testing and red teaming for LLM-based applications. AI Verify Testing Toolkit is used to test traditional AI applications for fairness, explainability, and robustness.

Still Have Questions?